Files
support_backend/specs/005-ai-support/research.md
T
saqib mirandClaude Sonnet 5 eceb00632d docs: correct 005-ai-support design to reuse the existing ticket state machine
003-ticketing's ticket-state-machine.ts already defines AI_ANALYZING/
AI_TROUBLESHOOTING/AI_VERIFYING/AI_RESOLVED/HUMAN_ESCALATION ticket
statuses, clearly authored anticipating this feature. Corrects the plan
before implementation: AISupportSession.status now drives Ticket.status
through the existing ticketsService.updateStatus (reusing its optimistic
concurrency), instead of an isolated status field the rest of the system
never sees. Also clarifies that knowledge retrieval is an in-process
service call through knowledge's index.ts, not an HTTP loopback.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 16:55:01 +05:30

23 KiB

Phase 0 Research: AI Support Agent

Decision: LLM provider integration — Anthropic SDK, manual loop, two calls per reasoning turn

  • Decision: Use @anthropic-ai/sdk directly (new Anthropic(), credential from ANTHROPIC_API_KEY). Each session turn is two model calls, not one, plus a deterministic step between them:
    1. Classify+diagnoseclient.messages.parse() with output_config.format (Zod schema via zodOutputFormat) against the conversation so far. No tools. Output: the structured AIDiagnosis shape (product/feature/problemType/severity/confidence/possibleCauses). Structured output and tool use are not combined in the same call — keeping diagnosis as a pure structured-output call means it can never emit a stray tool proposal, and keeps the confidence score honest (it's the model's stated belief about the classification, not entangled with whatever it also did with tools that turn).
    2. Deterministically (no model call): call knowledgeService.retrieve(...) (004, imported through ai-support/knowledge's public index.ts — Principle III; an in-process call, not an HTTP loopback to this same service's own GET /knowledge/retrieve route) scoped to the diagnosis's product/feature and the ticket's category, and apply the confidence-band policy (FR-004) to the diagnosis. This decides the branch: ask / proceed / escalate — this decision is application code, never delegated to the model.
    3. Reasoning/response — only on the "proceed" branch (or to word a clarifying question on the "ask" branch): a regular client.messages.create() call, given the diagnosis, the retrieved knowledge (as context, explicitly framed as data), the conversation so far, and — only on "proceed" — the tool registry's currently-enabled tools for this product. The model may return tool_use blocks here; a manual loop (not the SDK's beta tool runner) executes them, because policy evaluation has to happen before execution and has to be a first- class, auditable step of its own — see the tool-gating decision below — which the tool runner's run()-function-level gating pattern would bury inside each tool rather than express as a shared, visible gate.
  • Rationale: Directly implements doc 03 §1's four-step flow (Classification → Knowledge Retrieval → Reasoning → Next Action) as written, rather than collapsing it into one prompt that both classifies and acts — which would make confidence-band policy (a MUST per FR-004) something the model influences by how it phrases one big answer, instead of something applied deterministically to a discrete classification output.
  • Alternatives considered: One combined call producing diagnosis + response + tool calls together — rejected; makes it impossible to apply the confidence-band policy before the model has already committed to a response/tool calls, which is backwards from FR-004's "confidence decides what happens next." The SDK's beta tool runner — rejected for the reasoning call specifically (not for tool definition, which still uses the same Anthropic.Tool shape); the runner's per-tool run() gating pattern would scatter the policy check across each tool function instead of keeping it as one shared, auditable evaluation step ahead of any execution, which is what Constitution Principle IV ("AI recommends, policy decides") and FR-011 actually require — a visible decision point, not a convention every tool has to individually remember.

Decision: Model, thinking, and effort — configurable, not hardcoded

  • Decision: AI_SUPPORT_MODEL env var (default claude-opus-5), AI_SUPPORT_EFFORT env var (default medium, one of low/medium/high/xhigh/max). Every call uses thinking: { type: "adaptive" } (Claude Opus 5 runs adaptive thinking by default; this makes it explicit and keeps the code correct if the configured model changes) and output_config: { effort: AI_SUPPORT_EFFORT }.
  • Rationale: Doc 11 §B2 explicitly asks for model choice to be "a configurable policy, not a hardcoded model name" — this satisfies that with the simplest mechanism that fits this feature's scope (one configured model for all calls; see Assumptions in spec.md for why per-call model routing is out of scope). Defaulting to claude-opus-5 follows this codebase's standing default for new Claude integrations; the env var lets it be changed without a code change if the operator wants a different cost/quality tradeoff.
  • Alternatives considered: Hardcoding the model string inline at each call site — rejected, directly contradicts doc 11 §B2 and Constitution Principle II (config over hardcoding).

Decision: Confidence-band policy — new AIConfidencePolicy config table, most-specific-match fallback

  • Decision: A new model, AIConfidencePolicy (productId String?, categoryId String?, highThreshold Float, lowThreshold Float, maxClarifyingQuestions Int), unique on (productId, categoryId). Lookup order: exact (productId, categoryId) match → (productId, null) match → hardcoded system defaults from env (AI_SUPPORT_DEFAULT_HIGH_CONFIDENCE/AI_SUPPORT_DEFAULT_LOW_CONFIDENCE/ AI_SUPPORT_DEFAULT_MAX_CLARIFYING_QUESTIONS). A diagnosis with confidence >= highThreshold → proceed; confidence < lowThreshold → escalate; otherwise → ask.
  • Rationale: FR-005 requires per-product (optionally per-category) configuration "without requiring a code deploy" — a DB-backed config row an admin can write via an endpoint satisfies that directly, and the most-specific-match-with-fallback pattern means an operator never has to pre-seed every product before the feature works; the env-var defaults exist for exactly the "no configuration exists yet" case (FR-005's explicit fallback requirement).
  • Alternatives considered: A single global config row with no per-product override — rejected, doesn't satisfy FR-005's "per product (and optionally per category)" requirement, and doc 03 §4 explicitly frames thresholds as "tunable per product/category." Storing thresholds as JSON on Product itself — rejected; a dedicated table keeps the same admin-CRUD-with-versioning-free-writes shape as every other config surface in this codebase and supports the category-scoped override doc 03 asks for without denormalizing Product.

Decision: Tool system — code-defined registry, DB-backed per-product enablement not needed for this phase's tool set

  • Decision: Tools are defined in code (name, description, Zod input schema, permission string, risk level, auditRequired) as a small, fixed registry — src/modules/ai-support/tools/constants/tool-registry.ts. Four tools ship in this feature:
    • getTicketSnapshot (low risk) — reads the ticket, its problem, and recent messages; real, read-only, always available.
    • searchProductKnowledge (low risk) — re-runs knowledgeService.retrieve(...) (same in-process call as above) with a model-supplied feature/category refinement mid-conversation; real, read-only.
    • verifyProductResolution (low risk) — a documented, fail-closed placeholder, exactly the same pattern already established by ticketing/attachments's UnimplementedPlaceholderScanner: it always returns { confirmed: false, status: "unknown" }, never a fabricated success, because there is no real product-side signal to check yet (doc 11 §A2 — outbound webhooks from an integrated product don't exist in this codebase). This is what makes FR-018 ("never mark AI-resolved from customer confirmation alone") hold mechanically today: until a real per-product status check replaces this placeholder, verification can never auto-pass, so genuinely automatic AI-resolution won't happen in practice — which is the correct, honest behavior for a system with no real evidence source yet, not a bug to work around.
    • escalateToHuman (low risk, always policy-approved) — a real, functional action: ends the session the same way FR-020/FR-022 describe, lets the AI act on an explicit customer request for a human as a first-class, audited proposal rather than free-text the app has to sniff for. A fifth, high-risk tool, overrideTicketPriority (mutates Ticket.priority/severity based on the AI's assessment), exists in the registry specifically so SC-003 ("100% of high-risk proposals blocked from automatic execution... across every risk level") is actually exercised by a real tool, not vacuously true. Per FR-012, its risk tier requires the "stronger control path" — this feature implements that path as: the proposal and policy evaluation are recorded as pending_approval and the action never executes automatically. A human-approval UI is Phase 10 (Agent/Admin UI) work, out of scope here — this is the same class of explicitly documented known limitation as fastify.authenticate's auth stub, not a silent gap. Per-product enablement uses each tool's static supportedProducts field ("*" for all-products tools, or an explicit external-product-id allowlist) — no separate DB table for tool enablement in this phase, since FR-010/FR-011 only require the scope to be declared and checked, not admin-editable without a deploy (unlike confidence thresholds, which FR-005 explicitly requires to be).
  • Rationale: Matches FR-010 (permission/risk/product-scope declared per tool) and FR-013 (every proposal + evaluation + result durably recorded) exactly, while keeping the registry a plain, typed, code-reviewed artifact — appropriate for a small, fixed tool set, consistent with how Zod schemas generally aren't meant to be data-driven in this codebase. Reusing the fail-closed-placeholder pattern for verifyProductResolution is a direct, deliberate echo of an already-accepted precedent in this same codebase, not a new pattern being introduced.
  • Alternatives considered: A DB-backed dynamic tool registry (schemas as JSON, admin-editable) — rejected as significant unrequired complexity; nothing in spec.md's FRs asks for tools themselves to be admin-configurable, only for thresholds (FR-005) to be. Inventing a working fake product API to make verifyProductResolution "really" verify something — rejected; would misrepresent evidence this system doesn't actually have, which is precisely what FR-018 exists to prevent.

Decision: Deterministic policy gate — one shared evaluation function, ahead of any execution

  • Decision: evaluateToolProposal(tool, sessionContext) is a single function every proposed tool call passes through before anything executes: checks (a) the tool exists in the registry, (b) supportedProducts includes the session's product, (c) risk level — low → approved automatically; medium/high → recorded pending_approval, not executed. The AI's own message text is never consulted by this function — only the tool name and the session's actual product/ permission context (FR-011, FR-024).
  • Rationale: This is Constitution Principle IV made concrete for this feature, and directly answers doc 11 §A4's prompt-injection concern: since the gate never reads free-text content, content injected into a customer message or attachment has no path to influence what executes, no matter how it's phrased.
  • Alternatives considered: Per-tool ad hoc checks inside each tool's handler — rejected, same reasoning as the tool-runner rejection above: a shared gate is auditable and impossible to accidentally skip for a new tool; scattered checks are not.

Decision: Session triggering — enqueued on ticket creation, advanced over HTTP per customer turn

  • Decision: TicketsService.createFromInboundRequest (003-ticketing, src/modules/ticketing/tickets/service/tickets.service.ts) enqueues a QueueName.AI_SESSION job ({ ticketId }) right after a new ticket is created (not on an idempotent replay). A new worker (src/jobs/ai-session/index.ts, registered in src/bootstrap/queue.bootstrap.ts alongside the existing attachment worker) picks it up and runs the first diagnosis turn asynchronously. Every subsequent turn (a customer's reply, a runbook step's outcome) is driven by an explicit HTTP call — POST /tickets/:ticketId/ai-session/messages — which runs synchronously and returns the AI's next turn in the same response.
  • Rationale: The first turn happens off the hot path of POST /v1/support/requests (002's inbound integration boundary) — that caller shouldn't wait on an LLM round-trip just to get a ticket-created acknowledgment, and this codebase already has a real, working queue+worker pattern for exactly this shape of "do this after the request returns" work (the attachment malware-scan worker). Subsequent turns are naturally request/response — a customer reply is already an HTTP call into this system (matching 003's existing POST .../messages shape), and there's no reason to make the AI's response to it async when the caller is already waiting for an HTTP response.
  • Alternatives considered: Running the first diagnosis synchronously inside createFromInboundRequest — rejected; would add LLM latency (and a new failure mode: an LLM timeout) to every inbound SaaS integration request, which is the trust-boundary endpoint 002 already established as latency-sensitive. Polling instead of a queued worker — rejected, this codebase already has BullMQ wired up for exactly this "background work after a DB write" purpose.

Decision: AISupportSession.status drives Ticket.status through the existing state machine

  • Decision: src/modules/ticketing/tickets/mapper/ticket-state-machine.ts (built in 003-ticketing, before this feature existed) already defines NEW → AI_ANALYZING → AI_TROUBLESHOOTING → AI_VERIFYING → AI_RESOLVED, with HUMAN_ESCALATION reachable from every AI_* state and HUMAN_ESCALATION → IN_PROGRESS as the human hand-off edge — a near-exact match for doc 06's AISupportSession.status enum (analyzing | troubleshooting | verifying | resolved | escalated). This feature does not invent a parallel status concept: every time a session's own status changes, it calls the existing ticketsService.updateStatus(ticketId, newTicketStatus, expectedVersion, 'ai') (003) to drive the ticket through the matching status (analyzing→AI_ANALYZING, troubleshooting→AI_TROUBLESHOOTING, verifying→AI_VERIFYING, resolved→AI_RESOLVED, escalated→HUMAN_ESCALATION), reusing 003's own optimistic-concurrency handling (expectedVersion/409) rather than adding a second one. The session's own status field still exists separately (data-model.md) because it carries session-scoped values the ticket state machine doesn't need to know about (ended_by_agent — see FR-023 below — never appears on Ticket), but for every value the two share, the ticket is the caller-visible source of truth and the session record is the AI-internal detail behind it.
  • Rationale: Ticket.status is what every other part of this codebase (agents, SLA, orchestration once built, the ticket-status contract in 003) already reads to know where a ticket stands — a session-only status field that never touched Ticket.status would make the ticket lie about its own state while an AI session was quietly doing something else internally. Reusing 003's state machine and its updateStatus method also means this feature inherits 003's already-tested transition validation and concurrency guarantee for free, rather than re-deriving both.
  • Alternatives considered: A session-only status with no Ticket.status linkage — rejected per the rationale above. Building a second, AI-specific transition table — rejected; 003's table already defines exactly these states and edges, and doc 06's AISupportSession.status values were clearly authored to match it in the first place.
  • FR-023 implementation note: ticketsService.updateStatus gains one additional check — when it's called with an actor other than 'ai' (a human agent action) while an AISupportSession for that ticket is still active, the session is ended (status: ended_by_agent) as part of the same call, before the ticket's own status update commits. This is the concrete mechanism behind "a human agent takes ownership ends the AI session the same way an escalation does" (FR-023) — there's no separate "agent claims ticket" endpoint yet (orchestration/assignment is a later phase), so any human-actor status transition is the signal this feature has available today.

Decision: Runbook engine — the app selects the step, the model only phrases and interprets it

  • Decision: AISupportSession gains activeRunbookKey String? and currentStepIndex Int? (refining doc 06's conceptual AISupportSession, same "refine during Phase 1 modeling" convention 004 already established for KnowledgeEntry/Runbook). When a diagnosis's problemType matches an active Runbook.key for the ticket's product (004's RunbooksRepository.findCurrentByKey), the session sets activeRunbookKey and currentStepIndex = 0. The application reads steps[currentStepIndex] from the runbook's JSON and passes only that one step's content into the reasoning call's prompt as an instruction ("present this step, then interpret the customer's response against it") — the model is never given the full step list or asked to choose a step. On an outcome that means "try the next step," the app increments currentStepIndex deterministically; the model cannot set or advance the index itself (it has no tool for that).
  • Rationale: This is the literal requirement in FR-015/doc 03 §6 ("the AI cannot skip, reorder, or invent a step the runbook doesn't define") — the only way to guarantee that mechanically is for the index to be state the application owns and advances, with the model never seeing (and therefore never able to act on) any step but the current one.
  • Alternatives considered: Giving the model the full runbook and trusting a system-prompt instruction ("only present steps in order") — rejected; doc 03 §6 explicitly says not to trust the LLM to "improvise" here, and a prompt instruction is not a guarantee, it's a request the model could depart from under distribution shift or adversarial input (doc 11 §A4).

Decision: Verification and resolution — a session can only close on tool evidence

  • Decision: AISupportSession.status (and, via the decision above, Ticket.status) transitions verifying → resolved (AI_VERIFYING → AI_RESOLVED) only when a verifyProductResolution tool result exists on the session with confirmed: true — which, given that tool's current fail-closed placeholder implementation (above), means resolution through this exact tool never actually auto-fires yet in this deployment. Customer confirmation (a message recorded during the session) is stored and surfaced in the escalation/resolution summary, but the status transition's guard checks tool evidence only, never message content. Once Ticket.status reaches AI_RESOLVED, this feature's own responsibility ends — whether the ticket then moves to RESOLUTION_PENDING_CUSTOMER or straight to RESOLVED is 003-ticketing's existing generic status-update surface, not something this feature further automates.
  • Rationale: FR-018/FR-019 verbatim — resolution requires verification evidence, customer confirmation is secondary-only. Guarding the state transition on a structured tool-result field (not on parsing what the AI "said" about the outcome) keeps this enforceable in code, not just in prompt instructions.
  • Alternatives considered: Letting the reasoning call's own structured output include a resolved: boolean field the app trusts — rejected; this delegates a MUST-level policy decision (FR-018) to the model's own judgment, exactly the inversion Constitution Principle IV forbids.

Decision: Module placement — sessions, tools, troubleshooting, escalation under ai-support

  • Decision: Four new submodules under the existing src/modules/ai-support/ group (which 004 created with only knowledge populated): sessions/ (AISupportSession, AIDiagnosis, AIInteraction, AIConfidencePolicy, the two-call reasoning orchestration), tools/ (the registry, the policy gate, AIAction/AIActionResult), troubleshooting/ (the runbook-step engine, consuming 004's RunbooksRepository through knowledge's public index.ts), escalation/ (hand-off summary construction, ending a session, ticket-ownership hand-off). Doc 07's fuller list (agents/diagnosis/tool-execution/verification) is intentionally not built as separate submodules — diagnosis and verification are concerns inside sessions and tools respectively (an AIDiagnosis is produced by a session turn, not by an independent subsystem; verification is one tool's evaluated result, not a separate engine), and tool-execution is the same concern as tools split for no reason this feature's requirements give. agents (a registry of distinct AI "personas"/agent configs) has no requirement in spec.md at all — this feature has exactly one reasoning agent, not a multi-agent registry.
  • Rationale: Same reasoning 004 already used for not pre-building every doc 07 submodule speculatively — build what the current feature's FRs actually require, not the full documented taxonomy ahead of need.
  • Alternatives considered: One flat ai-support/agent/ module holding everything — rejected; four genuinely distinct responsibilities (session/diagnosis orchestration, tool policy/ execution, runbook stepping, escalation hand-off) benefit from the same controller/service/repository separation every other module in this codebase already uses, and cross-module imports must go through index.ts either way (Principle III) — collapsing them into one module wouldn't reduce real coupling, just hide the boundaries.

Decision: Admin endpoint authentication and message visibility — reuse existing conventions

  • Decision: The confidence-policy admin CRUD endpoint (PUT /admin/products/:externalProductId/ai-policy, optionally .../categories/:categoryId/ai-policy) is gated by fastify.authenticate, same known-limitation stub as every prior admin surface. AI-authored messages are written through 003-ticketing's existing messagesService.post(...) with type: 'AI_MESSAGE' — that type and its customer-visible mapping already exist in the message-type→visibility map (specs/003-ticketing/research.md); this feature adds no new message-visibility rule.
  • Rationale: Consistency with established precedent; introducing a different auth mechanism or a parallel message-writing path for this feature alone would be unjustified inconsistency.
  • Alternatives considered: None — direct reuse of existing, already-accepted conventions.