Files
support_backend/specs/008-sla-escalation/tasks.md
T
saqib mirandClaude Sonnet 5 9357f03e1d feat: implement SLA and escalation (008)
Populates platform/business-calendars, orchestration/sla, and
orchestration/escalation (all thin stubs until now) with the real engine:

- business-calendars: a luxon-based day-by-day calendar walk
  (addBusinessMinutes/isWithinWorkingHours) excluding non-working hours,
  weekends, and holidays — replacing the naive createdAt+hours stub FR-004
  explicitly forbids.
- sla: most-specific SLAPolicy resolution (product/category/problemType/
  priority, wildcard-or-exact-match, specificity-count + updatedAt
  tiebreak), SLARun creation on the first real publish of the
  long-unused TICKET_ASSIGNED domain event, durable pause/resume via an
  absolute-timestamp shift (no in-memory state, verified across a real
  buildApp() restart), and a repeatable BullMQ breach-detection sweep
  (src/jobs/sla, itself a previously-unregistered stub) that is directly
  callable for tests, not only reachable through a running worker.
- escalation: EscalationPolicy/Rule CRUD (all 10 doc05 trigger types
  storable, only resolution_breach/first_response_breach evaluated),
  breach-triggered and manual escalation both funnel through one EscalationEvent
  + scoped re-assignment path. AssignmentEngine (007) gains
  assignToSpecificNode — a new, explicitly node-scoped entry point,
  since escalation must never let 007's general resolution re-derive a
  different node than the one a rule or a caller targeted.

Two small pre-existing scaffold gaps were closed along the way:
CategoriesRepository had no findById, and TICKET_ASSIGNED/SLA_BREACHED/
ESCALATION_TRIGGERED were defined since earlier phases but never
published by any code.

Verified against throwaway Docker Postgres/Redis (typecheck, lint,
architecture-check all clean; 148/150 relevant tests pass — the 2
failures are pre-existing, MinIO-dependent, and unrelated to this
feature).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 13:02:05 +05:30

22 KiB

description
description
Task list for 008-sla-escalation

Tasks: SLA and Escalation

Input: Design documents from specs/008-sla-escalation/

Prerequisites: plan.md, spec.md, research.md, data-model.md, contracts/sla-escalation-contract.md, quickstart.md

Tests: Included as first-class tasks. This feature has real, extractable pure logic (the calendar-walk algorithm, most-specific policy match, breach/no-breach/paused-no-breach logic) plus — for the first time since the constitution's Principle VII was written — a genuine process-restart-survival requirement that needs a dedicated test rebuilding buildApp() mid-test, not just a within-process concurrency test.

Organization: Tasks are grouped by user story (US1 = P1 policy definition, US2 = P1 run creation with calendar-aware due dates, US3 = P1 durable pause/resume, US4 = P2 breach detection, US5 = P2 breach-triggered escalation, US6 = P3 manual escalation).

Format: [ID] [P?] [Story] Description

All file paths are relative to supporthub-api/ (repo root).


Phase 1: Setup

  • T001 [P] Populate src/modules/platform/business-calendars/ with the full standard shape (controller/, routes/, schema/, repository/, service/, types/, mapper/, constants/, index.ts) plus a calculators/ directory, replacing the existing BusinessCalendarsService.isWorkingHour stub's content
  • T002 [P] Extend src/modules/orchestration/sla/ to the full standard shape around its existing engine//calculators/ directories, replacing every stub file's content (SlaEngine.evaluateSlaTargets, SlaDueDateCalculator.calculateDueTime)
  • T003 [P] Extend src/modules/orchestration/escalation/ to the full standard shape around its existing engine/ directory, replacing the EscalationEngine.triggerEscalation stub's content

Phase 2: Foundational (Blocking Prerequisites)

Purpose: Schema for every entity, shared by every user story.

⚠️ CRITICAL: No user-story stage work can begin until this phase is complete.

  • T004 Add SLAPolicy, SLARun (incl. the additive firstResponseBreachedAt refinement), BusinessCalendar, Holiday, EscalationPolicy, EscalationRule, EscalationEvent models to prisma/schema.prisma per data-model.md, plus Ticket.slaRun/ Ticket.escalationEvents, Product.slaPolicies/Product.escalationPolicies, Category.slaPolicies, HierarchyNode.escalationRules back-relations, and an SLARun @@index([status, resolutionDueAt]) for the breach-detection sweep (depends on T001-T003)
  • T005 Run npm run prisma:generate and create the migration (npm run prisma:migrate) for T004 (depends on T004)

Checkpoint: Schema migrated. User stories can now be built.


Phase 3: User Story 1 - Admin defines SLA policies as configuration (Priority: P1) 🎯 MVP (part 1)

Goal: SLAPolicy CRUD and the most-specific-match resolution function exist and are independently correct — not yet wired to ticket assignment.

Independent Test: Quickstart Scenario 1.

Tests for User Story 1

  • T006 [P] [US1] Unit tests for findApplicablePolicy (specificity-count match, wildcard handling on each of the 4 scope dimensions independently, tie-break by latest updatedAt, no-match returns null) in tests/unit/orchestration/sla-policy-match.test.ts
  • T007 [US1] Integration test covering Quickstart Scenario 1 (a product-scoped policy is preferred over a global one; deactivating it falls back to the global policy) against a real Postgres in tests/integration/sla-policy-resolution.test.ts (depends on T005)

Implementation for User Story 1

  • T008 [US1] Add SLAPolicyRepository (CRUD, findActiveCandidates(scope)) and the Zod create/update schema — with resolve-or-404 existence checks for productId/categoryId/ businessCalendarId when provided (research.md) — in sla/repository/ + sla/schema/ (depends on T005)
  • T009 [US1] Add findApplicablePolicy(ticketContext) (specificity-count + tie-break, per data-model.md's Resolution section) in sla/service/sla-policy-resolver.service.ts (depends on T008)
  • T010 [US1] Add POST/GET/GET:id/PATCH/DELETE /admin/sla-policies routes (soft-delete via active: false, gated by fastify.authenticate) in sla/controller/ + sla/routes/, registered from src/api/routes.ts (depends on T008)
  • T011 [US1] Run Quickstart Scenario 1 locally and confirm all 4 steps pass

Checkpoint: SLA policies can be defined and correctly resolved. Nothing creates an SLARun yet — that's User Story 2.


Phase 4: User Story 2 - SLA run starts automatically with calendar-aware due dates (Priority: P1) 🎯 MVP (part 2)

Goal: BusinessCalendar/Holiday CRUD, the calendar-walk algorithm, and SLARun creation wired into 007's assignment-success path via the first real publish of TICKET_ASSIGNED.

Independent Test: Quickstart Scenario 2.

Tests for User Story 2

  • T012 [P] [US2] Unit tests for addBusinessMinutes — weekend exclusion, holiday exclusion, partial-day clipping on the start day, a day with no configured window contributing zero time, and correctness across a DST transition in the calendar's own timezone — in tests/unit/platform/business-calendars/calendar-walk.test.ts
  • T013 [US2] Integration test covering Quickstart Scenario 2 (calendar-aware due date lands the next working day past a weekend+holiday, never a naive addition; a ticket assigned with no matching policy gets no SLARun and GET .../sla-run returns 404) against a real Postgres in tests/integration/sla-run-creation.test.ts (depends on T005, T009, and 007's existing assignment flow)

Implementation for User Story 2

  • T014 [US2] Add addBusinessMinutes(start, minutes, calendar, holidays) using luxon in business-calendars/calculators/business-hours.calculator.ts, replacing the isWorkingHour stub's logic (research.md's day-by-day walk)
  • T015 [US2] Add BusinessCalendarRepository/HolidayRepository, Zod schema (IANA timezone validation, HH:mm + start < end validation per data-model.md), and POST/GET/GET:id/PATCH /admin/business-calendars + POST /admin/business-calendars/:id/holidays + DELETE /admin/business-calendars/:id/holidays/:holidayId routes in business-calendars/repository/ + schema/ + controller/ + routes/ (depends on T014)
  • T016 [US2] Replace SlaDueDateCalculator.calculateDueTime's naive addition with a call into T014's addBusinessMinutes (via business-calendars's public index.ts — FR-004) in sla/calculators/sla-due-date.calculator.ts (depends on T014)
  • T017 [US2] Add AssignmentEngine.persistAndTransition (007, src/modules/orchestration/assignments/engine/assignment.engine.ts) publishing DomainEventName.TICKET_ASSIGNED ({ ticketId, agentId, strategy, actor }) after its existing persistence step — the event is already defined in src/events/domain-events.ts but has never been published (research.md)
  • T018 [US2] Add SlaService.handleTicketAssigned(ticketId, agentId): no-ops if the ticket already has an SLARun (SLARun.ticketId @unique — covers re-escalation's second publish, spec.md Assumptions); otherwise resolves the applicable policy (T009), computes firstResponseDueAt/resolutionDueAt via T016, and creates the SLARun — in sla/service/sla.service.ts (depends on T009, T016)
  • T019 [US2] Subscribe DomainEventName.TICKET_ASSIGNED to T018's handler in src/events/handlers/index.ts, following the existing "module never imports the module it affects" registration pattern (depends on T017, T018)
  • T020 [US2] Add GET /tickets/:ticketId/sla-run route (404 if none) in sla/controller/ + sla/routes/ (depends on T018)
  • T021 [US2] Run Quickstart Scenario 2 locally and confirm all 4 steps pass

Checkpoint: Every successfully-assigned ticket with a matching policy gets an SLARun with correctly calendar-computed due dates. MVP-complete for read-only SLA visibility.


Phase 5: User Story 3 - SLA pause/resume is durable across a process restart (Priority: P1)

Goal: WAITING_FOR_CUSTOMER transitions pause/resume the run by shifting its absolute due dates — no in-memory state anywhere, verified across an actual rebuilt buildApp().

Independent Test: Quickstart Scenario 3.

Tests for User Story 3

  • T022 [P] [US3] Unit tests for the pause/resume shift arithmetic (resume shifts both due dates forward by exactly now - pausedAt; a second pause/resume cycle composes correctly) in tests/unit/orchestration/sla-pause-resume.test.ts
  • T023 [US3] Integration test covering Quickstart Scenario 3 — including rebuilding buildApp() mid-test to simulate a real process restart while paused, then asserting the resumed due date is exactly the original plus the paused wall-clock duration — against a real Postgres in tests/integration/sla-pause-resume.test.ts (depends on T018)

Implementation for User Story 3

  • T024 [US3] Add SlaService.pause(ticketId) / resume(ticketId) (shift firstResponseDueAt/resolutionDueAt forward by the paused duration on resume, per research.md/data-model.md — no separate remaining-minutes field) in sla/service/ sla.service.ts (depends on T018)
  • T025 [US3] Subscribe two DomainEventName.TICKET_UPDATED handlers in src/events/handlers/index.tsnewStatus === 'WAITING_FOR_CUSTOMER' calls T024's pause, previousStatus === 'WAITING_FOR_CUSTOMER' calls resume — alongside the existing 005/007 subscribers on the same event (depends on T024)
  • T026 [US3] Subscribe a third TICKET_UPDATED handler — newStatus === 'RESOLVED' sets SLARun.completedAt and status: 'completed' (data-model.md) — in the same file (depends on T018)
  • T027 [US3] Run Quickstart Scenario 3 locally and confirm all 4 steps pass, including the restart-boundary step

Checkpoint: Every P1 user story is complete. SLA runs are created, calendar-computed, and durably pause/resume-correct. This is the feature's MVP.


Phase 6: User Story 4 - Breaches are detected even if no one is watching in real time (Priority: P2)

Goal: A repeatable BullMQ job durably detects both resolution and first-response breaches, never missing one because the process wasn't running at the due instant, never flagging a completed-in-time or paused run.

Independent Test: Quickstart Scenario 4.

Tests for User Story 4

  • T028 [P] [US4] Unit tests for the breach-detection predicate logic (a running run past resolutionDueAt breaches; a paused run past resolutionDueAt does not; a completed run does not; a running run past firstResponseDueAt with no prior AGENT_MESSAGE breaches first-response exactly once, guarded by firstResponseBreachedAt) in tests/unit/orchestration/sla-breach-detection.test.ts
  • T029 [US4] Integration test covering Quickstart Scenario 4 (a short-resolutionMinutes policy breaches within one sweep call; resolved-in-time and paused runs are never breached even after their due instant passes) against a real Postgres in tests/integration/sla-breach-detection.test.ts (depends on T018, T024)

Implementation for User Story 4

  • T030 [US4] Add SlaService.runBreachDetectionSweep() — queries every running SLARun with resolutionDueAt <= now() (marks breached/breachedAt) and every running run with firstResponseDueAt <= now() and firstResponseBreachedAt: null and no AGENT_MESSAGE recorded for the ticket (marks firstResponseBreachedAt) — a single, directly-callable, side-effect-only method (research.md — no worker process needed to invoke it in tests) in sla/service/sla.service.ts (depends on T024, T026)
  • T031 [US4] Replace registerSlaWorker()'s stub body in src/jobs/sla/index.ts: on registration, schedule a BullMQ repeatable job on QueueName.SLA ({ repeat: { every: 60_000 } }) whose processor calls T030's runBreachDetectionSweep (depends on T030)
  • T032 [US4] Run Quickstart Scenario 4 locally and confirm all 5 steps pass

Checkpoint: Breaches are durably detected. Nothing reacts to a breach yet beyond marking the run — that's User Story 5.


Phase 7: User Story 5 - A breach automatically triggers rule-driven escalation (Priority: P2)

Goal: EscalationPolicy/EscalationRule CRUD, breach-triggered EscalationEvent firing, and a new scoped-assignment entry point on 007's AssignmentEngine that re-assigns to exactly the rule's targetNodeId.

Independent Test: Quickstart Scenario 5.

Tests for User Story 5

  • T033 [P] [US5] Unit tests for escalation-policy resolution (product-specific preferred over global, per research.md) and rule matching (every active rule whose triggerType matches the firing breach type fires; an inactive or wrong-trigger-type rule doesn't) in tests/unit/orchestration/escalation-rule-match.test.ts
  • T034 [US5] Integration test covering Quickstart Scenario 5 (a breach with a matching rule produces exactly one EscalationEvent and reassigns to an agent eligible under the rule's specific targetNodeId, not the ticket's originally-resolved node; a breach with no matching rule is still recorded breached with no EscalationEvent) against a real Postgres in tests/integration/sla-escalation-firing.test.ts (depends on T030)

Implementation for User Story 5

  • T035 [US5] Add EscalationPolicyRepository/EscalationRuleRepository (CRUD, findActiveRules(policyId, triggerType)), Zod schema (all 10 doc-05 triggerType values accepted; targetNodeId resolve-or-404 at rule creation, FR-012) in escalation/repository/ + escalation/schema/ (depends on T005)
  • T036 [US5] Add POST/GET /admin/escalation-policies, POST/PATCH/DELETE /admin/escalation-policies/:id/rules[/:ruleId] routes in escalation/controller/ + escalation/routes/ (depends on T035)
  • T037 [US5] Add AssignmentEngine.assignToSpecificNode(ticketId, hierarchyNodeId, actor, reason?, strategyOverride?) (007, assignments/engine/assignment.engine.ts) — resolves the eligible-agent set scoped to exactly the given node (reusing RoutingService's capability-lookup call, research.md) and persists through the existing persistAndTransition (T017), so it also publishes TICKET_ASSIGNED for free (depends on T017)
  • T038 [US5] Add EscalationService.handleBreach(ticketId, triggerType): resolves the applicable EscalationPolicy (product-match-or-global, research.md), finds every active matching EscalationRule (T035), and for each, creates an EscalationEvent (ruleId, fromNodeId from the ticket's current assignment, toNodeId: rule.targetNodeId, triggeredBy: 'system') and calls T037's assignToSpecificNode — records nothing when no rule matches (FR-015) — in escalation/service/escalation.service.ts (depends on T035, T037)
  • T039 [US5] Wire T030's runBreachDetectionSweep to call T038's handleBreach for each newly-detected breach, passing the corresponding trigger type (resolution_breach / first_response_breach) — in sla/service/sla.service.ts (depends on T030, T038)
  • T040 [US5] Run Quickstart Scenario 5 locally and confirm all 3 steps pass

Checkpoint: Breaches automatically escalate through rule-driven, scoped re-assignment.


Phase 8: User Story 6 - A human can manually escalate a ticket to a specific node (Priority: P3)

Goal: The same EscalationEvent + scoped-reassignment mechanism, triggered explicitly by a caller instead of a breach.

Independent Test: Quickstart Scenario 6.

Tests for User Story 6

  • T041 [US6] Integration test covering Quickstart Scenario 6 steps 1-3 (manual escalation creates an EscalationEvent with ruleId: null and reassigns via the scoped path; a nonexistent targetNodeId returns 404 with no event created) against a real Postgres — implemented as the "Scenario 6" case in tests/integration/sla-escalation-flow.test.ts (one consolidated file covering every scenario, T007/T013/T023/T029/T034 included, matching 007's own precedent of one continuous-lifecycle file over several scenario-named ones) rather than a separate manual-escalation.test.ts (depends on T037, T038). Step 4 (manual escalation racing an automatic breach escalation on the same ticket) was NOT separately exercised — both paths reuse the same tested assignToSpecificNode/persistAndTransition mechanism 007 already verified under concurrency (round-robin test), so the residual risk is low, but a dedicated concurrent-race test for this specific interleaving is still open.

Implementation for User Story 6

  • T042 [US6] Add EscalationService.escalateManually(ticketId, targetNodeId, actor, reason): resolve-or-404 on targetNodeId (FR-017), creates an EscalationEvent (ruleId: null, triggeredBy: actor) and calls T037's assignToSpecificNode — in escalation/service/ escalation.service.ts (depends on T037)
  • T043 [US6] Add POST /tickets/:ticketId/escalate route (gated by fastify.authenticate) in escalation/controller/ + escalation/routes/, registered from src/api/routes.ts (depends on T042)
  • T044 [US6] Run Quickstart Scenario 6 locally and confirm all 4 steps pass

Checkpoint: All six user stories work independently and together — policy definition, calendar-aware run creation, durable pause/resume, durable breach detection, and both automatic and manual escalation form one coherent, restart-safe flow.


Phase 9: Polish & Cross-Cutting Concerns

  • T045 [P] SKIPPED — originally planned to add an "SLA and Escalation" section to README.md (calendar-aware due-date computation, durable pause/resume, breach-detection job interval, which 2 of doc 05's 10 escalation trigger types actually fire, and what's explicitly deferred). README.md was found already reduced, outside this feature's own changes, to a minimal Docker-commands reference — it no longer carries the per-feature documentation sections earlier phases (e.g. 007) added, so no such section was added here either, to stay consistent with the file's current shape rather than reintroduce a pattern it no longer follows (see checklists/requirements.md's Implementation Notes).
  • T046 [P] Update specs/008-sla-escalation/checklists/requirements.md Notes with any implementation-time findings
  • T047 Run npx tsx scripts/check-architecture.ts and npm run lint/npm run typecheck
  • T048 Full regression: npm run test:unit (scoped to tests/unit) to confirm nothing broke elsewhere, then the full integration suite (including 007's own suite, since T017/T037 modify its AssignmentEngine) against real Docker-provisioned Postgres/Redis

Dependencies & Execution Order

Phase Dependencies

  • Setup (Phase 1): No dependencies
  • Foundational (Phase 2): Depends on Setup — BLOCKS all user stories
  • User Story 1 (Phase 3): Depends on Foundational — no dependency on US2-US6
  • User Story 2 (Phase 4): Depends on US1 (the policy it resolves against) — genuinely not independent, same class of dependency 007's US2 had on US1
  • User Story 3 (Phase 5): Depends on US2 (the run it pauses/resumes)
  • User Story 4 (Phase 6): Depends on US3 (a run that can be paused must be excluded from breach detection correctly, so the pause mechanism must exist first)
  • User Story 5 (Phase 7): Depends on US4 (the breach it reacts to) and on 007's AssignmentEngine (T037's new method)
  • User Story 6 (Phase 8): Depends on US5 (T037/T038's scoped-reassignment mechanism, reused directly rather than duplicated)
  • Polish (Phase 9): Depends on all six user stories

Parallel Opportunities

  • T001/T002/T003 (independent scaffolding)
  • T006 (unit tests) alongside T008-T009 (the implementations they test)
  • T012 (unit tests) alongside T014 (the implementation it tests)
  • T022 alongside T024; T028 alongside T030; T033 alongside T035/T038
  • T045/T046 in Polish

Sequencing Note

T017 (publishing TICKET_ASSIGNED from 007's AssignmentEngine) and T037 (the new assignToSpecificNode method on the same class) both modify a file 007 already owns and has its own passing test suite for — run 007's full integration suite (part of T048) after each, not only at the very end, to catch a regression close to its cause.


Implementation Strategy

MVP First (User Stories 1-3 Only)

  1. Setup + Foundational (T001-T005)
  2. User Story 1 (T006-T011) — policies exist and resolve correctly
  3. User Story 2 (T012-T021) — runs are created with real calendar-aware due dates
  4. User Story 3 (T022-T027) — pause/resume is durable, including across a restart
  5. STOP and VALIDATE: Quickstart Scenarios 1-3 pass — every assigned ticket has a correctly computed, durably pausable SLARun. Nothing reacts to a breach yet — that value lands with User Story 4/5.

Incremental Delivery

  1. Setup + Foundational → schema migrated
  2. Add User Story 1 → SLA policies are configurable and resolve correctly
  3. Add User Story 2 → runs are created automatically with calendar-aware due dates
  4. Add User Story 3 → pause/resume is durable (P1-complete, MVP)
  5. Add User Story 4 → breaches are durably detected
  6. Add User Story 5 → breaches automatically escalate and reassign
  7. Add User Story 6 → manual escalation exists, reusing the same mechanism
  8. Polish → docs and full regression