Files
support_backend/specs/008-sla-escalation/tasks.md
T

388 lines
21 KiB
Markdown
Raw Normal View History

---
description: "Task list for 008-sla-escalation"
---
# Tasks: SLA and Escalation
**Input**: Design documents from `specs/008-sla-escalation/`
**Prerequisites**: [plan.md](./plan.md), [spec.md](./spec.md), [research.md](./research.md),
[data-model.md](./data-model.md),
[contracts/sla-escalation-contract.md](./contracts/sla-escalation-contract.md),
[quickstart.md](./quickstart.md)
**Tests**: Included as first-class tasks. This feature has real, extractable pure logic (the
calendar-walk algorithm, most-specific policy match, breach/no-breach/paused-no-breach logic)
plus — for the first time since the constitution's Principle VII was written — a genuine
process-restart-survival requirement that needs a dedicated test rebuilding `buildApp()`
mid-test, not just a within-process concurrency test.
**Organization**: Tasks are grouped by user story (US1 = P1 policy definition, US2 = P1 run
creation with calendar-aware due dates, US3 = P1 durable pause/resume, US4 = P2 breach detection,
US5 = P2 breach-triggered escalation, US6 = P3 manual escalation).
## Format: `[ID] [P?] [Story] Description`
All file paths are relative to `supporthub-api/` (repo root).
---
## Phase 1: Setup
- [ ] T001 [P] Populate `src/modules/platform/business-calendars/` with the full standard shape
(`controller/`, `routes/`, `schema/`, `repository/`, `service/`, `types/`, `mapper/`,
`constants/`, `index.ts`) plus a `calculators/` directory, replacing the existing
`BusinessCalendarsService.isWorkingHour` stub's content
- [ ] T002 [P] Extend `src/modules/orchestration/sla/` to the full standard shape around its
existing `engine/`/`calculators/` directories, replacing every stub file's content
(`SlaEngine.evaluateSlaTargets`, `SlaDueDateCalculator.calculateDueTime`)
- [ ] T003 [P] Extend `src/modules/orchestration/escalation/` to the full standard shape around
its existing `engine/` directory, replacing the `EscalationEngine.triggerEscalation` stub's
content
---
## Phase 2: Foundational (Blocking Prerequisites)
**Purpose**: Schema for every entity, shared by every user story.
**⚠️ CRITICAL**: No user-story stage work can begin until this phase is complete.
- [ ] T004 Add `SLAPolicy`, `SLARun` (incl. the additive `firstResponseBreachedAt` refinement),
`BusinessCalendar`, `Holiday`, `EscalationPolicy`, `EscalationRule`, `EscalationEvent`
models to `prisma/schema.prisma` per data-model.md, plus `Ticket.slaRun`/
`Ticket.escalationEvents`, `Product.slaPolicies`/`Product.escalationPolicies`,
`Category.slaPolicies`, `HierarchyNode.escalationRules` back-relations, and an
`SLARun @@index([status, resolutionDueAt])` for the breach-detection sweep (depends on
T001-T003)
- [ ] T005 Run `npm run prisma:generate` and create the migration (`npm run prisma:migrate`) for
T004 (depends on T004)
**Checkpoint**: Schema migrated. User stories can now be built.
---
## Phase 3: User Story 1 - Admin defines SLA policies as configuration (Priority: P1) 🎯 MVP (part 1)
**Goal**: `SLAPolicy` CRUD and the most-specific-match resolution function exist and are
independently correct — not yet wired to ticket assignment.
**Independent Test**: Quickstart Scenario 1.
### Tests for User Story 1
- [ ] T006 [P] [US1] Unit tests for `findApplicablePolicy` (specificity-count match, wildcard
handling on each of the 4 scope dimensions independently, tie-break by latest `updatedAt`,
no-match returns `null`) in `tests/unit/orchestration/sla-policy-match.test.ts`
- [ ] T007 [US1] Integration test covering Quickstart Scenario 1 (a product-scoped policy is
preferred over a global one; deactivating it falls back to the global policy) against a
real Postgres in `tests/integration/sla-policy-resolution.test.ts` (depends on T005)
### Implementation for User Story 1
- [ ] T008 [US1] Add `SLAPolicyRepository` (CRUD, `findActiveCandidates(scope)`) and the Zod
create/update schema — with resolve-or-404 existence checks for `productId`/`categoryId`/
`businessCalendarId` when provided (research.md) — in `sla/repository/` + `sla/schema/`
(depends on T005)
- [ ] T009 [US1] Add `findApplicablePolicy(ticketContext)` (specificity-count + tie-break, per
data-model.md's Resolution section) in `sla/service/sla-policy-resolver.service.ts`
(depends on T008)
- [ ] T010 [US1] Add `POST/GET/GET:id/PATCH/DELETE /admin/sla-policies` routes (soft-delete via
`active: false`, gated by `fastify.authenticate`) in `sla/controller/` + `sla/routes/`,
registered from `src/api/routes.ts` (depends on T008)
- [ ] T011 [US1] Run Quickstart Scenario 1 locally and confirm all 4 steps pass
**Checkpoint**: SLA policies can be defined and correctly resolved. Nothing creates an `SLARun`
yet — that's User Story 2.
---
## Phase 4: User Story 2 - SLA run starts automatically with calendar-aware due dates (Priority: P1) 🎯 MVP (part 2)
**Goal**: `BusinessCalendar`/`Holiday` CRUD, the calendar-walk algorithm, and `SLARun` creation
wired into 007's assignment-success path via the first real publish of `TICKET_ASSIGNED`.
**Independent Test**: Quickstart Scenario 2.
### Tests for User Story 2
- [ ] T012 [P] [US2] Unit tests for `addBusinessMinutes` — weekend exclusion, holiday exclusion,
partial-day clipping on the start day, a day with no configured window contributing zero
time, and correctness across a DST transition in the calendar's own timezone — in
`tests/unit/platform/business-calendars/calendar-walk.test.ts`
- [ ] T013 [US2] Integration test covering Quickstart Scenario 2 (calendar-aware due date lands
the next working day past a weekend+holiday, never a naive addition; a ticket assigned with
no matching policy gets no `SLARun` and `GET .../sla-run` returns `404`) against a real
Postgres in `tests/integration/sla-run-creation.test.ts` (depends on T005, T009, and 007's
existing assignment flow)
### Implementation for User Story 2
- [ ] T014 [US2] Add `addBusinessMinutes(start, minutes, calendar, holidays)` using `luxon` in
`business-calendars/calculators/business-hours.calculator.ts`, replacing the
`isWorkingHour` stub's logic (research.md's day-by-day walk)
- [ ] T015 [US2] Add `BusinessCalendarRepository`/`HolidayRepository`, Zod schema (IANA timezone
validation, `HH:mm` + `start < end` validation per data-model.md), and
`POST/GET/GET:id/PATCH /admin/business-calendars` +
`POST /admin/business-calendars/:id/holidays` +
`DELETE /admin/business-calendars/:id/holidays/:holidayId` routes in
`business-calendars/repository/` + `schema/` + `controller/` + `routes/` (depends on T014)
- [ ] T016 [US2] Replace `SlaDueDateCalculator.calculateDueTime`'s naive addition with a call
into T014's `addBusinessMinutes` (via `business-calendars`'s public `index.ts` — FR-004) in
`sla/calculators/sla-due-date.calculator.ts` (depends on T014)
- [ ] T017 [US2] Add `AssignmentEngine.persistAndTransition` (007,
`src/modules/orchestration/assignments/engine/assignment.engine.ts`) publishing
`DomainEventName.TICKET_ASSIGNED` (`{ ticketId, agentId, strategy, actor }`) after its
existing persistence step — the event is already defined in `src/events/domain-events.ts`
but has never been published (research.md)
- [ ] T018 [US2] Add `SlaService.handleTicketAssigned(ticketId, agentId)`: no-ops if the ticket
already has an `SLARun` (`SLARun.ticketId @unique` — covers re-escalation's second publish,
spec.md Assumptions); otherwise resolves the applicable policy (T009), computes
`firstResponseDueAt`/`resolutionDueAt` via T016, and creates the `SLARun` — in
`sla/service/sla.service.ts` (depends on T009, T016)
- [ ] T019 [US2] Subscribe `DomainEventName.TICKET_ASSIGNED` to T018's handler in
`src/events/handlers/index.ts`, following the existing "module never imports the module it
affects" registration pattern (depends on T017, T018)
- [ ] T020 [US2] Add `GET /tickets/:ticketId/sla-run` route (`404` if none) in `sla/controller/` +
`sla/routes/` (depends on T018)
- [ ] T021 [US2] Run Quickstart Scenario 2 locally and confirm all 4 steps pass
**Checkpoint**: Every successfully-assigned ticket with a matching policy gets an `SLARun` with
correctly calendar-computed due dates. MVP-complete for read-only SLA visibility.
---
## Phase 5: User Story 3 - SLA pause/resume is durable across a process restart (Priority: P1)
**Goal**: `WAITING_FOR_CUSTOMER` transitions pause/resume the run by shifting its absolute due
dates — no in-memory state anywhere, verified across an actual rebuilt `buildApp()`.
**Independent Test**: Quickstart Scenario 3.
### Tests for User Story 3
- [ ] T022 [P] [US3] Unit tests for the pause/resume shift arithmetic (resume shifts both due
dates forward by exactly `now - pausedAt`; a second pause/resume cycle composes correctly)
in `tests/unit/orchestration/sla-pause-resume.test.ts`
- [ ] T023 [US3] Integration test covering Quickstart Scenario 3 — including rebuilding
`buildApp()` mid-test to simulate a real process restart while paused, then asserting the
resumed due date is exactly the original plus the paused wall-clock duration — against a
real Postgres in `tests/integration/sla-pause-resume.test.ts` (depends on T018)
### Implementation for User Story 3
- [ ] T024 [US3] Add `SlaService.pause(ticketId)` / `resume(ticketId)` (shift
`firstResponseDueAt`/`resolutionDueAt` forward by the paused duration on resume, per
research.md/data-model.md — no separate remaining-minutes field) in `sla/service/
sla.service.ts` (depends on T018)
- [ ] T025 [US3] Subscribe two `DomainEventName.TICKET_UPDATED` handlers in
`src/events/handlers/index.ts``newStatus === 'WAITING_FOR_CUSTOMER'` calls T024's
`pause`, `previousStatus === 'WAITING_FOR_CUSTOMER'` calls `resume` — alongside the existing
005/007 subscribers on the same event (depends on T024)
- [ ] T026 [US3] Subscribe a third `TICKET_UPDATED` handler — `newStatus === 'RESOLVED'` sets
`SLARun.completedAt` and `status: 'completed'` (data-model.md) — in the same file (depends
on T018)
- [ ] T027 [US3] Run Quickstart Scenario 3 locally and confirm all 4 steps pass, including the
restart-boundary step
**Checkpoint**: Every P1 user story is complete. SLA runs are created, calendar-computed, and
durably pause/resume-correct. This is the feature's MVP.
---
## Phase 6: User Story 4 - Breaches are detected even if no one is watching in real time (Priority: P2)
**Goal**: A repeatable BullMQ job durably detects both resolution and first-response breaches,
never missing one because the process wasn't running at the due instant, never flagging a
completed-in-time or paused run.
**Independent Test**: Quickstart Scenario 4.
### Tests for User Story 4
- [ ] T028 [P] [US4] Unit tests for the breach-detection predicate logic (a `running` run past
`resolutionDueAt` breaches; a `paused` run past `resolutionDueAt` does not; a `completed`
run does not; a `running` run past `firstResponseDueAt` with no prior `AGENT_MESSAGE`
breaches first-response exactly once, guarded by `firstResponseBreachedAt`) in
`tests/unit/orchestration/sla-breach-detection.test.ts`
- [ ] T029 [US4] Integration test covering Quickstart Scenario 4 (a short-`resolutionMinutes`
policy breaches within one sweep call; resolved-in-time and paused runs are never breached
even after their due instant passes) against a real Postgres in
`tests/integration/sla-breach-detection.test.ts` (depends on T018, T024)
### Implementation for User Story 4
- [ ] T030 [US4] Add `SlaService.runBreachDetectionSweep()` — queries every `running` `SLARun`
with `resolutionDueAt <= now()` (marks `breached`/`breachedAt`) and every `running` run with
`firstResponseDueAt <= now()` and `firstResponseBreachedAt: null` and no `AGENT_MESSAGE`
recorded for the ticket (marks `firstResponseBreachedAt`) — a single, directly-callable,
side-effect-only method (research.md — no worker process needed to invoke it in tests) in
`sla/service/sla.service.ts` (depends on T024, T026)
- [ ] T031 [US4] Replace `registerSlaWorker()`'s stub body in `src/jobs/sla/index.ts`: on
registration, schedule a BullMQ repeatable job on `QueueName.SLA` (`{ repeat: { every:
60_000 } }`) whose processor calls T030's `runBreachDetectionSweep` (depends on T030)
- [ ] T032 [US4] Run Quickstart Scenario 4 locally and confirm all 5 steps pass
**Checkpoint**: Breaches are durably detected. Nothing reacts to a breach yet beyond marking the
run — that's User Story 5.
---
## Phase 7: User Story 5 - A breach automatically triggers rule-driven escalation (Priority: P2)
**Goal**: `EscalationPolicy`/`EscalationRule` CRUD, breach-triggered `EscalationEvent` firing, and
a new scoped-assignment entry point on 007's `AssignmentEngine` that re-assigns to exactly the
rule's `targetNodeId`.
**Independent Test**: Quickstart Scenario 5.
### Tests for User Story 5
- [ ] T033 [P] [US5] Unit tests for escalation-policy resolution (product-specific preferred over
global, per research.md) and rule matching (every active rule whose `triggerType` matches
the firing breach type fires; an inactive or wrong-trigger-type rule doesn't) in
`tests/unit/orchestration/escalation-rule-match.test.ts`
- [ ] T034 [US5] Integration test covering Quickstart Scenario 5 (a breach with a matching rule
produces exactly one `EscalationEvent` and reassigns to an agent eligible under the rule's
specific `targetNodeId`, not the ticket's originally-resolved node; a breach with no
matching rule is still recorded breached with no `EscalationEvent`) against a real Postgres
in `tests/integration/sla-escalation-firing.test.ts` (depends on T030)
### Implementation for User Story 5
- [ ] T035 [US5] Add `EscalationPolicyRepository`/`EscalationRuleRepository` (CRUD,
`findActiveRules(policyId, triggerType)`), Zod schema (all 10 doc-05 `triggerType` values
accepted; `targetNodeId` resolve-or-404 at rule creation, FR-012) in
`escalation/repository/` + `escalation/schema/` (depends on T005)
- [ ] T036 [US5] Add `POST/GET /admin/escalation-policies`,
`POST/PATCH/DELETE /admin/escalation-policies/:id/rules[/:ruleId]` routes in
`escalation/controller/` + `escalation/routes/` (depends on T035)
- [ ] T037 [US5] Add `AssignmentEngine.assignToSpecificNode(ticketId, hierarchyNodeId, actor,
reason?, strategyOverride?)` (007, `assignments/engine/assignment.engine.ts`) — resolves
the eligible-agent set scoped to exactly the given node (reusing `RoutingService`'s
capability-lookup call, research.md) and persists through the existing
`persistAndTransition` (T017), so it also publishes `TICKET_ASSIGNED` for free (depends on
T017)
- [ ] T038 [US5] Add `EscalationService.handleBreach(ticketId, triggerType)`: resolves the
applicable `EscalationPolicy` (product-match-or-global, research.md), finds every active
matching `EscalationRule` (T035), and for each, creates an `EscalationEvent`
(`ruleId`, `fromNodeId` from the ticket's current assignment, `toNodeId: rule.targetNodeId`,
`triggeredBy: 'system'`) and calls T037's `assignToSpecificNode` — records nothing when no
rule matches (FR-015) — in `escalation/service/escalation.service.ts` (depends on T035,
T037)
- [ ] T039 [US5] Wire T030's `runBreachDetectionSweep` to call T038's `handleBreach` for each
newly-detected breach, passing the corresponding trigger type (`resolution_breach` /
`first_response_breach`) — in `sla/service/sla.service.ts` (depends on T030, T038)
- [ ] T040 [US5] Run Quickstart Scenario 5 locally and confirm all 3 steps pass
**Checkpoint**: Breaches automatically escalate through rule-driven, scoped re-assignment.
---
## Phase 8: User Story 6 - A human can manually escalate a ticket to a specific node (Priority: P3)
**Goal**: The same `EscalationEvent` + scoped-reassignment mechanism, triggered explicitly by a
caller instead of a breach.
**Independent Test**: Quickstart Scenario 6.
### Tests for User Story 6
- [ ] T041 [US6] Integration test covering Quickstart Scenario 6 (manual escalation creates an
`EscalationEvent` with `ruleId: null` and reassigns via the scoped path; a nonexistent
`targetNodeId` returns `404` with no event created; a manual escalation racing an automatic
breach escalation on the same ticket records both events without a corrupted final
assignment) against a real Postgres in `tests/integration/manual-escalation.test.ts`
(depends on T037, T038)
### Implementation for User Story 6
- [ ] T042 [US6] Add `EscalationService.escalateManually(ticketId, targetNodeId, actor, reason)`:
resolve-or-404 on `targetNodeId` (FR-017), creates an `EscalationEvent` (`ruleId: null`,
`triggeredBy: actor`) and calls T037's `assignToSpecificNode` — in `escalation/service/
escalation.service.ts` (depends on T037)
- [ ] T043 [US6] Add `POST /tickets/:ticketId/escalate` route (gated by `fastify.authenticate`)
in `escalation/controller/` + `escalation/routes/`, registered from `src/api/routes.ts`
(depends on T042)
- [ ] T044 [US6] Run Quickstart Scenario 6 locally and confirm all 4 steps pass
**Checkpoint**: All six user stories work independently and together — policy definition,
calendar-aware run creation, durable pause/resume, durable breach detection, and both automatic
and manual escalation form one coherent, restart-safe flow.
---
## Phase 9: Polish & Cross-Cutting Concerns
- [ ] T045 [P] Add an "SLA and Escalation" section to `README.md` describing the calendar-aware
due-date computation, the durable pause/resume mechanism, the breach-detection job interval,
which 2 of doc 05's 10 escalation trigger types actually fire, and what's explicitly
deferred (notification delivery, investigation/customer-response timers, reopen-cycle SLA
restart)
- [ ] T046 [P] Update `specs/008-sla-escalation/checklists/requirements.md` Notes with any
implementation-time findings
- [ ] T047 Run `npx tsx scripts/check-architecture.ts` and `npm run lint`/`npm run typecheck`
- [ ] T048 Full regression: `npm run test:unit` (scoped to `tests/unit`) to confirm nothing broke
elsewhere, then the full integration suite (including 007's own suite, since T017/T037
modify its `AssignmentEngine`) against real Docker-provisioned Postgres/Redis
---
## Dependencies & Execution Order
### Phase Dependencies
- **Setup (Phase 1)**: No dependencies
- **Foundational (Phase 2)**: Depends on Setup — BLOCKS all user stories
- **User Story 1 (Phase 3)**: Depends on Foundational — no dependency on US2-US6
- **User Story 2 (Phase 4)**: Depends on US1 (the policy it resolves against) — genuinely not
independent, same class of dependency 007's US2 had on US1
- **User Story 3 (Phase 5)**: Depends on US2 (the run it pauses/resumes)
- **User Story 4 (Phase 6)**: Depends on US3 (a run that can be paused must be excluded from
breach detection correctly, so the pause mechanism must exist first)
- **User Story 5 (Phase 7)**: Depends on US4 (the breach it reacts to) and on 007's
`AssignmentEngine` (T037's new method)
- **User Story 6 (Phase 8)**: Depends on US5 (T037/T038's scoped-reassignment mechanism, reused
directly rather than duplicated)
- **Polish (Phase 9)**: Depends on all six user stories
### Parallel Opportunities
- T001/T002/T003 (independent scaffolding)
- T006 (unit tests) alongside T008-T009 (the implementations they test)
- T012 (unit tests) alongside T014 (the implementation it tests)
- T022 alongside T024; T028 alongside T030; T033 alongside T035/T038
- T045/T046 in Polish
### Sequencing Note
T017 (publishing `TICKET_ASSIGNED` from 007's `AssignmentEngine`) and T037 (the new
`assignToSpecificNode` method on the same class) both modify a file 007 already owns and has its
own passing test suite for — run 007's full integration suite (part of T048) after each, not only
at the very end, to catch a regression close to its cause.
---
## Implementation Strategy
### MVP First (User Stories 1-3 Only)
1. Setup + Foundational (T001-T005)
2. User Story 1 (T006-T011) — policies exist and resolve correctly
3. User Story 2 (T012-T021) — runs are created with real calendar-aware due dates
4. User Story 3 (T022-T027) — pause/resume is durable, including across a restart
5. **STOP and VALIDATE**: Quickstart Scenarios 1-3 pass — every assigned ticket has a correctly
computed, durably pausable `SLARun`. Nothing reacts to a breach yet — that value lands with
User Story 4/5.
### Incremental Delivery
1. Setup + Foundational → schema migrated
2. Add User Story 1 → SLA policies are configurable and resolve correctly
3. Add User Story 2 → runs are created automatically with calendar-aware due dates
4. Add User Story 3 → pause/resume is durable (P1-complete, MVP)
5. Add User Story 4 → breaches are durably detected
6. Add User Story 5 → breaches automatically escalate and reassign
7. Add User Story 6 → manual escalation exists, reusing the same mechanism
8. Polish → docs and full regression