Files
support_backend/docs/10-implementation-roadmap.md
T
saqib mirandClaude Sonnet 5 f475a55a53 docs: add product and engineering specification
Full system blueprint (docs 01-10): product vision, integration &
security, AI support architecture, ticketing & problem management,
orchestration/SLA/escalation, database schema, backend/frontend
architecture, testing/observability/CI-CD, and the implementation
roadmap. This is the pre-implementation design reference the codebase
is being built against.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 16:22:21 +05:30

5.4 KiB
Raw Blame History

10 — Implementation Roadmap

Implement incrementally. Do not implement all business logic in one step. Each phase below must be typed, tested, documented, integrated, observable, and production-safe before moving on — a feature isn't "done" until it clears all seven of those, not just "coded."

Phased plan

Phase Focus Primary deliverables
1 Engineering foundation Repo scaffolding, Fastify modular-monolith skeleton, Next.js app skeleton, Prisma schema baseline, CI pipeline skeleton, env validation, health endpoints
2 SaaS integration Product/ProductIntegration models, credential validation, service-to-service auth (signed tokens/OAuth2/mTLS), inbound request contract, rate limiting
3 Product knowledge Knowledge/KnownIssue/ErrorCode/Runbook models, admin CRUD, versioning + publish state, retrieval (RAG) layer
4 AI support AI session/diagnosis/interaction models, classification, RAG-backed reasoning, confidence thresholds (configurable), tool system with permission/risk gating, runbook engine, verification logic
5 Ticketing Ticket + Problem models (kept separate), message types, attachment pipeline (object storage, scanning, expiring URLs), ticket lifecycle state machine
6 Support organization Team/Agent/AgentSkill/AgentAvailability models, dynamic HierarchyNode configuration, admin hierarchy editor
7 Orchestration and assignment Orchestration engine, capability matching, pluggable assignment strategies, concurrency-safe round robin, assignment history
8 SLA and escalation SLA policy engine, business calendar/holiday support, durable pause/resume via BullMQ, rule-driven escalation engine, escalation event audit
9 Problem resolution Investigation/RootCause/Solution/SolutionImplementation/SolutionVerification/Resolution models and workflows, customer confirmation + reopen flow
10 Agent/Admin UI Agent workspace (continues from AI context), admin configuration surfaces for every configurable subsystem above
11 Analytics, hardening, security, production deployment Reporting dashboards, full observability, security hardening pass, load/concurrency testing, production deployment pipeline

Note the dependency direction: Phase 4 (AI) and Phase 9 (resolution stages) both plug into Phase 5's ticket, so ticketing's core data model should be stable before AI or resolution logic is built against it — even though ticketing is listed after AI support here, expect to iterate the Ticket/Problem shape lightly across phases 49 rather than treating phase 5 as strictly sequential.

Success criteria checklist

Use this as the actual go/no-go list, not phase names — a phase can be "complete" on paper while missing several of these.

  • A SaaS product can securely integrate with SupportHub
  • The SaaS user can enter support from within the product
  • SupportHub receives trusted product/tenant/user context
  • A problem creates a durable support case immediately
  • AI understands the problem
  • AI retrieves correct product knowledge
  • AI can diagnose known issues
  • AI can guide the customer through supported troubleshooting
  • The system verifies successful resolution with evidence
  • AI can resolve supported issues automatically
  • Unresolved issues are escalated automatically
  • AI context is preserved in the ticket for the human agent
  • Orchestration chooses the correct support path
  • Capability/skill matching works correctly
  • Assignment respects availability/workload
  • SLA is applied correctly (calendar-aware, durable)
  • Escalation occurs according to configured policy
  • Human agents receive full context, don't restart diagnosis
  • Agents can investigate and resolve the problem
  • Resolution is verified and recorded
  • Customer sees the final result
  • Complete audit history exists
  • All critical operations are observable
  • CI/CD can validate and deploy the system safely

Open business decisions (do not invent final values)

Anything the business hasn't finalized yet must be explicitly marked in code, config schema, and documentation as one of:

CONFIGURABLE
OPEN BUSINESS DECISION
REQUIRES BUSINESS CONFIRMATION

Known candidates for this list at spec time:

  • Actual SLA minute values per product/priority/category
  • Actual escalation rule conditions and target nodes per product
  • Whether/when premium support tiers override "support enabled by default"
  • Confidence-threshold cut points for high/medium/low AI bands, per product
  • Assignment strategy choice per hierarchy node
  • Business calendar definitions (hours, holidays, timezones) per team

Never hardcode a placeholder value for any of the above and ship it as if it were final — mark it and surface it for confirmation instead.

Cross-reference map

If you're building... Read
The overall model and boundaries 01, 02
The AI agent 03
The ticket/problem data model 04, 06
Orchestration/SLA/escalation 05
Backend module layout 07
Any UI surface 08
Tests, CI, dashboards 09