Files
support_backend/README.md
T
saqib mirandClaude Sonnet 5 82d02bcdcd feat: implement AI support agent (005) — diagnosis, tools, runbooks, verification
Real Anthropic Claude integration per explicit product decision: a
ticket's AI session diagnoses the problem via a structured-output call,
applies a DB-configurable confidence-band policy (FR-005), and on
"proceed" reasons and acts through a small permission/risk-gated tool
system (FR-011/FR-012), optionally walking a matching runbook step by
step with the application — never the model — owning the step index
(FR-015/FR-016). Resolution requires real tool evidence, never customer
claims alone (FR-018) — verifyProductResolution is a documented
fail-closed placeholder mirroring the existing malware-scanner precedent,
since no real per-product operational signal exists yet.

AISupportSession.status mirrors onto Ticket.status through 003-ticketing's
existing AI_ANALYZING/AI_TROUBLESHOOTING/AI_VERIFYING/AI_RESOLVED/
HUMAN_ESCALATION state machine, discovered during planning to have been
built anticipating this exact feature. Two circular module dependencies
(escalation<->sessions, tools<->sessions) were designed around rather than
found as bugs: escalation is a pure summary formatter with no state
dependencies of its own, and tools stays a clean leaf module with zero
dependency on ai-support/sessions. Ticket creation enqueues the first
diagnosis turn via the existing queue infrastructure (off the hot path of
the inbound SaaS integration endpoint); a human actor changing ticket
status ends the AI session via the event-bus scaffold that existed in
this codebase but had never been wired to anything.

A real Prisma limitation was found and fixed before it reached tests:
compound-unique upsert rejects null for a nullable key column, so
AIConfidencePolicy uses find-then-update/create instead, same fix class
004 already used for the same underlying limitation.

Adds 9 unit tests (confidence-band, tool-policy-gate, runbook-step-
advance) and 6 integration test files, including the two constitution-
required standing E2E scenarios. AI-independent tests were run against
real Postgres/Redis/MinIO (88 passed, 0 failed across the full suite,
including every pre-existing 002/003/004 test). The AI-dependent tests
compile and skip cleanly via describe.skipIf but were not run against a
live model — no ANTHROPIC_API_KEY was available in this session; a real
key must be supplied before this feature can actually run.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 17:44:52 +05:30

8.9 KiB

Development

docker compose --env-file .env.development -f docker-compose.development.yml up -d --build

Test

docker compose --env-file .env.test -f docker-compose.test.yml up --build

Production

docker compose --env-file .env.prod -f docker-compose.prod.yml up --build -d

Stop

docker compose -f docker-compose.prod.yml down

Local environment setup

.env.development, .env.test, and .env.prod are gitignored (they hold real credentials) — copy .env.example to the one you need and fill in real values before running any command above.

CI/CD

Every push/PR triggers the Jenkins pipeline defined in Jenkinsfile. Stage order: checkout → install → environment validation → typecheck → lint → format check → unit test → integration test → E2E test → build → Docker build → publish → deploy. Publish/deploy only run on branches with a configured deploy target (main → prod, develop/test → test); other branches validate and build only. Pipeline run status and per-stage logs are visible in the Jenkins UI for the relevant job — see specs/001-ci-pipeline/quickstart.md for how to validate the pipeline itself, and specs/001-ci-pipeline/contracts/pipeline-stage-contract.md for the guarantees each stage makes.

Required Jenkins credentials (see the header comment in Jenkinsfile for exact IDs): per target environment (test, prod) a Postgres password, Redis password, JWT secret, and AWS access key/secret, plus one shared Docker registry username/password. None of these are ever read from a file in this repository.

SaaS Integration

POST /v1/support/requests is the trust boundary a registered SaaS product calls through — every request must carry a Authorization: Bearer <signed-token> header (HMAC-SHA256, signed with the integration's own secret) and a body matching the inbound contract. See specs/002-saas-integration/contracts/inbound-request-contract.md for the full validation order and error codes, and specs/002-saas-integration/quickstart.md for runnable scenarios.

Admins manage integrations under /admin/products/:externalProductId/integration (register) and /admin/integrations/:integrationId/{rotate,revoke,status,audit-trail}. These admin routes are not yet actually access-controlledfastify.authenticate is a stub pending the identity/auth module; don't expose them outside a trusted network until that's implemented.

Rate limits (rateLimitPerMinute, rateLimitPerUserPerMinute) are set per integration at registration time and enforced via a Redis-backed fixed-window counter, independent of the global @fastify/rate-limit floor already applied to every route.

Ticketing

A validated inbound request (see "SaaS Integration" above) creates a Ticket and Problem immediately — before any diagnosis. See specs/003-ticketing/contracts/ticket-lifecycle-contract.md for the full lifecycle state machine, message-visibility rules, and attachment pipeline, and specs/003-ticketing/quickstart.md for runnable scenarios.

  • Status transitions: PATCH /tickets/:ticketId/status requires expectedVersion (optimistic concurrency — a stale version is rejected with 409, never silently overwritten) and only accepts transitions defined in the state machine (400 INVALID_TRANSITION otherwise).
  • Messages: POST/GET /tickets/:ticketId/messages (customer-scoped — internal note types are never returned) and GET /agent/tickets/:ticketId/messages (agent-scoped — everything). A message's customer-visibility is always derived from its type, never caller-supplied.
  • Attachments: presigned-PUT upload (POST .../attachments/upload-urlPOST .../attachments/confirm) against MinIO/S3 — file bytes never transit this API. Nothing is downloadable yet (GET .../attachments/:attachmentId/download-url always returns 409): the malware scanner is a placeholder that fails closed until a real one (src/modules/ticketing/attachments/mapper/malware-scanner.ts) replaces it.
  • Local/test object storage is MinIO — see the minio service in docker-compose.development.yml / docker-compose.test.yml and the AWS_S3_ENDPOINT value in the corresponding .env.* file.

Product Knowledge

Admin CRUD for KnowledgeEntry/ErrorCode/KnownIssue/Runbook, plus GET /knowledge/retrieve — a filtered (not semantic/vector) query the future AI-support feature will call. See specs/004-product-knowledge/contracts/knowledge-contract.md for the full route list and specs/004-product-knowledge/quickstart.md for runnable scenarios.

  • Versioning: editing a published KnowledgeEntry or Runbook never overwrites it in place — it creates a new row (version incremented, isCurrentVersion: true), and the prior version stays queryable (GET /admin/knowledge/:code/versions). Requires expectedVersion; a stale value is rejected with 409, same concurrency pattern as ticket status updates.
  • Retrieval (GET /knowledge/retrieve?productId=&feature=&category=) only ever returns published, currently-effective, current-version entries scoped to the given product — validated entries are ranked ahead of unvalidated ones. An unregistered productId returns an empty array, not an error.
  • Full semantic/embedding-based retrieval is intentionally not implemented here — see specs/004-product-knowledge/spec.md Assumptions.

AI Support

A ticket's AI session diagnoses the problem via a real Anthropic Claude call, applies a configurable confidence-band policy, and — on "proceed" — reasons and acts through a small, permission/risk-gated tool system, optionally walking a matching runbook step by step. See specs/005-ai-support/contracts/ai-support-contract.md for the full route list and specs/005-ai-support/quickstart.md for runnable scenarios.

  • Requires a real ANTHROPIC_API_KEY (per explicit product decision — not a mock or pluggable-interface phase). The app boots and every non-AI test still passes without one; an AI session simply fails closed (escalates) if a reasoning call is attempted with none configured. AI_SUPPORT_MODEL/AI_SUPPORT_EFFORT and the system-wide confidence-threshold/question-budget defaults are all environment-configurable, never hardcoded (doc 11 §B2).
  • Session lifecycle: a session starts automatically (queued, off the hot path of POST /v1/support/requests) when a ticket is created, and its status mirrors onto Ticket.status through 003-ticketing's existing state machine (AI_ANALYZINGAI_TROUBLESHOOTINGAI_VERIFYINGAI_RESOLVED, or HUMAN_ESCALATION from any point) — see specs/005-ai-support/research.md "AISupportSession.status drives Ticket.status". A customer reply is POST /tickets/:ticketId/ai-session/messages; the current session (with its diagnosis and conversation) is GET /tickets/:ticketId/ai-session.
  • Confidence policy: PUT/GET /admin/products/:externalProductId/ai-policy sets per-product (optionally per-category) highThreshold/lowThreshold/maxClarifyingQuestions, falling back to env-configured system defaults when nothing is configured — applies to the very next diagnosis, no deploy required.
  • Tool system: every tool call the AI proposes is evaluated by a deterministic policy gate (src/modules/ai-support/tools/service/policy-gate.ts) before anything executes — the gate never reads the AI's own proposal/justification text, only the tool's declared risk level and product scope. Low-risk tools (getTicketSnapshot, searchProductKnowledge, verifyProductResolution, escalateToHuman) auto-execute; overrideTicketPriority is high-risk and always stays pending_approval — there's no human-approval UI yet (Phase 10), so it never actually runs, by design, not by oversight. Every proposal, decision, and result is recorded and auditable via GET /tickets/:ticketId/ai-session/actions.
  • verifyProductResolution is a documented, fail-closed placeholder (same pattern as ticketing/attachments's malware scanner) — it always returns confirmed: false, since there's no real per-product operational signal to check yet (doc 11 §A2). A ticket is only ever marked AI_RESOLVED on a passing result from this tool, never from what the customer says alone — so in practice, genuinely automatic AI resolution won't happen until a real verification signal replaces this placeholder.
  • Runbook engine: when a diagnosis's problemType matches a runbook's key for the product (004-product-knowledge), the application — never the model — tracks which step is current (currentStepIndex) and advances it by exactly one at a time; exhausting every step without resolving escalates with the full attempted sequence attached.
  • Semantic/vector retrieval, product-signal webhook verification, model routing/fallback, cost dashboards, and localization are intentionally out of scope here — see specs/005-ai-support/spec.md Assumptions.