003-ticketing's ticket-state-machine.ts already defines AI_ANALYZING/ AI_TROUBLESHOOTING/AI_VERIFYING/AI_RESOLVED/HUMAN_ESCALATION ticket statuses, clearly authored anticipating this feature. Corrects the plan before implementation: AISupportSession.status now drives Ticket.status through the existing ticketsService.updateStatus (reusing its optimistic concurrency), instead of an isolated status field the rest of the system never sees. Also clarifies that knowledge retrieval is an in-process service call through knowledge's index.ts, not an HTTP loopback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
23 KiB
23 KiB
Phase 0 Research: AI Support Agent
Decision: LLM provider integration — Anthropic SDK, manual loop, two calls per reasoning turn
- Decision: Use
@anthropic-ai/sdkdirectly (new Anthropic(), credential fromANTHROPIC_API_KEY). Each session turn is two model calls, not one, plus a deterministic step between them:- Classify+diagnose —
client.messages.parse()withoutput_config.format(Zod schema viazodOutputFormat) against the conversation so far. No tools. Output: the structuredAIDiagnosisshape (product/feature/problemType/severity/confidence/possibleCauses). Structured output and tool use are not combined in the same call — keeping diagnosis as a pure structured-output call means it can never emit a stray tool proposal, and keeps the confidence score honest (it's the model's stated belief about the classification, not entangled with whatever it also did with tools that turn). - Deterministically (no model call): call
knowledgeService.retrieve(...)(004, imported throughai-support/knowledge's publicindex.ts— Principle III; an in-process call, not an HTTP loopback to this same service's ownGET /knowledge/retrieveroute) scoped to the diagnosis's product/feature and the ticket's category, and apply the confidence-band policy (FR-004) to the diagnosis. This decides the branch: ask / proceed / escalate — this decision is application code, never delegated to the model. - Reasoning/response — only on the "proceed" branch (or to word a clarifying question on
the "ask" branch): a regular
client.messages.create()call, given the diagnosis, the retrieved knowledge (as context, explicitly framed as data), the conversation so far, and — only on "proceed" — the tool registry's currently-enabled tools for this product. The model may returntool_useblocks here; a manual loop (not the SDK's beta tool runner) executes them, because policy evaluation has to happen before execution and has to be a first- class, auditable step of its own — see the tool-gating decision below — which the tool runner'srun()-function-level gating pattern would bury inside each tool rather than express as a shared, visible gate.
- Classify+diagnose —
- Rationale: Directly implements doc 03 §1's four-step flow (Classification → Knowledge Retrieval → Reasoning → Next Action) as written, rather than collapsing it into one prompt that both classifies and acts — which would make confidence-band policy (a MUST per FR-004) something the model influences by how it phrases one big answer, instead of something applied deterministically to a discrete classification output.
- Alternatives considered: One combined call producing diagnosis + response + tool calls
together — rejected; makes it impossible to apply the confidence-band policy before the model
has already committed to a response/tool calls, which is backwards from FR-004's "confidence
decides what happens next." The SDK's beta tool runner — rejected for the reasoning call
specifically (not for tool definition, which still uses the same
Anthropic.Toolshape); the runner's per-toolrun()gating pattern would scatter the policy check across each tool function instead of keeping it as one shared, auditable evaluation step ahead of any execution, which is what Constitution Principle IV ("AI recommends, policy decides") and FR-011 actually require — a visible decision point, not a convention every tool has to individually remember.
Decision: Model, thinking, and effort — configurable, not hardcoded
- Decision:
AI_SUPPORT_MODELenv var (defaultclaude-opus-5),AI_SUPPORT_EFFORTenv var (defaultmedium, one oflow/medium/high/xhigh/max). Every call usesthinking: { type: "adaptive" }(Claude Opus 5 runs adaptive thinking by default; this makes it explicit and keeps the code correct if the configured model changes) andoutput_config: { effort: AI_SUPPORT_EFFORT }. - Rationale: Doc 11 §B2 explicitly asks for model choice to be "a configurable policy, not a
hardcoded model name" — this satisfies that with the simplest mechanism that fits this feature's
scope (one configured model for all calls; see Assumptions in spec.md for why per-call model
routing is out of scope). Defaulting to
claude-opus-5follows this codebase's standing default for new Claude integrations; the env var lets it be changed without a code change if the operator wants a different cost/quality tradeoff. - Alternatives considered: Hardcoding the model string inline at each call site — rejected, directly contradicts doc 11 §B2 and Constitution Principle II (config over hardcoding).
Decision: Confidence-band policy — new AIConfidencePolicy config table, most-specific-match fallback
- Decision: A new model,
AIConfidencePolicy(productId String?,categoryId String?,highThreshold Float,lowThreshold Float,maxClarifyingQuestions Int), unique on(productId, categoryId). Lookup order: exact(productId, categoryId)match →(productId, null)match → hardcoded system defaults from env (AI_SUPPORT_DEFAULT_HIGH_CONFIDENCE/AI_SUPPORT_DEFAULT_LOW_CONFIDENCE/AI_SUPPORT_DEFAULT_MAX_CLARIFYING_QUESTIONS). A diagnosis withconfidence >= highThreshold→ proceed;confidence < lowThreshold→ escalate; otherwise → ask. - Rationale: FR-005 requires per-product (optionally per-category) configuration "without requiring a code deploy" — a DB-backed config row an admin can write via an endpoint satisfies that directly, and the most-specific-match-with-fallback pattern means an operator never has to pre-seed every product before the feature works; the env-var defaults exist for exactly the "no configuration exists yet" case (FR-005's explicit fallback requirement).
- Alternatives considered: A single global config row with no per-product override — rejected,
doesn't satisfy FR-005's "per product (and optionally per category)" requirement, and doc 03 §4
explicitly frames thresholds as "tunable per product/category." Storing thresholds as JSON on
Productitself — rejected; a dedicated table keeps the same admin-CRUD-with-versioning-free-writes shape as every other config surface in this codebase and supports the category-scoped override doc 03 asks for without denormalizingProduct.
Decision: Tool system — code-defined registry, DB-backed per-product enablement not needed for this phase's tool set
- Decision: Tools are defined in code (name, description, Zod input schema, permission
string, risk level,
auditRequired) as a small, fixed registry —src/modules/ai-support/tools/constants/tool-registry.ts. Four tools ship in this feature:getTicketSnapshot(low risk) — reads the ticket, its problem, and recent messages; real, read-only, always available.searchProductKnowledge(low risk) — re-runsknowledgeService.retrieve(...)(same in-process call as above) with a model-supplied feature/category refinement mid-conversation; real, read-only.verifyProductResolution(low risk) — a documented, fail-closed placeholder, exactly the same pattern already established byticketing/attachments'sUnimplementedPlaceholderScanner: it always returns{ confirmed: false, status: "unknown" }, never a fabricated success, because there is no real product-side signal to check yet (doc 11 §A2 — outbound webhooks from an integrated product don't exist in this codebase). This is what makes FR-018 ("never mark AI-resolved from customer confirmation alone") hold mechanically today: until a real per-product status check replaces this placeholder, verification can never auto-pass, so genuinely automatic AI-resolution won't happen in practice — which is the correct, honest behavior for a system with no real evidence source yet, not a bug to work around.escalateToHuman(low risk, always policy-approved) — a real, functional action: ends the session the same way FR-020/FR-022 describe, lets the AI act on an explicit customer request for a human as a first-class, audited proposal rather than free-text the app has to sniff for. A fifth, high-risk tool,overrideTicketPriority(mutatesTicket.priority/severitybased on the AI's assessment), exists in the registry specifically so SC-003 ("100% of high-risk proposals blocked from automatic execution... across every risk level") is actually exercised by a real tool, not vacuously true. Per FR-012, its risk tier requires the "stronger control path" — this feature implements that path as: the proposal and policy evaluation are recorded aspending_approvaland the action never executes automatically. A human-approval UI is Phase 10 (Agent/Admin UI) work, out of scope here — this is the same class of explicitly documented known limitation asfastify.authenticate's auth stub, not a silent gap. Per-product enablement uses each tool's staticsupportedProductsfield ("*"for all-products tools, or an explicit external-product-id allowlist) — no separate DB table for tool enablement in this phase, since FR-010/FR-011 only require the scope to be declared and checked, not admin-editable without a deploy (unlike confidence thresholds, which FR-005 explicitly requires to be).
- Rationale: Matches FR-010 (permission/risk/product-scope declared per tool) and FR-013
(every proposal + evaluation + result durably recorded) exactly, while keeping the registry a
plain, typed, code-reviewed artifact — appropriate for a small, fixed tool set, consistent with
how Zod schemas generally aren't meant to be data-driven in this codebase. Reusing the
fail-closed-placeholder pattern for
verifyProductResolutionis a direct, deliberate echo of an already-accepted precedent in this same codebase, not a new pattern being introduced. - Alternatives considered: A DB-backed dynamic tool registry (schemas as JSON, admin-editable)
— rejected as significant unrequired complexity; nothing in spec.md's FRs asks for tools
themselves to be admin-configurable, only for thresholds (FR-005) to be. Inventing a working
fake product API to make
verifyProductResolution"really" verify something — rejected; would misrepresent evidence this system doesn't actually have, which is precisely what FR-018 exists to prevent.
Decision: Deterministic policy gate — one shared evaluation function, ahead of any execution
- Decision:
evaluateToolProposal(tool, sessionContext)is a single function every proposed tool call passes through before anything executes: checks (a) the tool exists in the registry, (b)supportedProductsincludes the session's product, (c) risk level —low→ approved automatically;medium/high→ recordedpending_approval, not executed. The AI's own message text is never consulted by this function — only the tool name and the session's actual product/ permission context (FR-011, FR-024). - Rationale: This is Constitution Principle IV made concrete for this feature, and directly answers doc 11 §A4's prompt-injection concern: since the gate never reads free-text content, content injected into a customer message or attachment has no path to influence what executes, no matter how it's phrased.
- Alternatives considered: Per-tool ad hoc checks inside each tool's handler — rejected, same reasoning as the tool-runner rejection above: a shared gate is auditable and impossible to accidentally skip for a new tool; scattered checks are not.
Decision: Session triggering — enqueued on ticket creation, advanced over HTTP per customer turn
- Decision:
TicketsService.createFromInboundRequest(003-ticketing,src/modules/ticketing/tickets/service/tickets.service.ts) enqueues aQueueName.AI_SESSIONjob ({ ticketId }) right after a new ticket is created (not on an idempotent replay). A new worker (src/jobs/ai-session/index.ts, registered insrc/bootstrap/queue.bootstrap.tsalongside the existing attachment worker) picks it up and runs the first diagnosis turn asynchronously. Every subsequent turn (a customer's reply, a runbook step's outcome) is driven by an explicit HTTP call —POST /tickets/:ticketId/ai-session/messages— which runs synchronously and returns the AI's next turn in the same response. - Rationale: The first turn happens off the hot path of
POST /v1/support/requests(002's inbound integration boundary) — that caller shouldn't wait on an LLM round-trip just to get a ticket-created acknowledgment, and this codebase already has a real, working queue+worker pattern for exactly this shape of "do this after the request returns" work (the attachment malware-scan worker). Subsequent turns are naturally request/response — a customer reply is already an HTTP call into this system (matching 003's existingPOST .../messagesshape), and there's no reason to make the AI's response to it async when the caller is already waiting for an HTTP response. - Alternatives considered: Running the first diagnosis synchronously inside
createFromInboundRequest— rejected; would add LLM latency (and a new failure mode: an LLM timeout) to every inbound SaaS integration request, which is the trust-boundary endpoint 002 already established as latency-sensitive. Polling instead of a queued worker — rejected, this codebase already has BullMQ wired up for exactly this "background work after a DB write" purpose.
Decision: AISupportSession.status drives Ticket.status through the existing state machine
- Decision:
src/modules/ticketing/tickets/mapper/ticket-state-machine.ts(built in 003-ticketing, before this feature existed) already definesNEW → AI_ANALYZING → AI_TROUBLESHOOTING → AI_VERIFYING → AI_RESOLVED, withHUMAN_ESCALATIONreachable from every AI_* state andHUMAN_ESCALATION → IN_PROGRESSas the human hand-off edge — a near-exact match for doc 06'sAISupportSession.statusenum (analyzing | troubleshooting | verifying | resolved | escalated). This feature does not invent a parallel status concept: every time a session's own status changes, it calls the existingticketsService.updateStatus(ticketId, newTicketStatus, expectedVersion, 'ai')(003) to drive the ticket through the matching status (analyzing→AI_ANALYZING,troubleshooting→AI_TROUBLESHOOTING,verifying→AI_VERIFYING,resolved→AI_RESOLVED,escalated→HUMAN_ESCALATION), reusing 003's own optimistic-concurrency handling (expectedVersion/409) rather than adding a second one. The session's ownstatusfield still exists separately (data-model.md) because it carries session-scoped values the ticket state machine doesn't need to know about (ended_by_agent— see FR-023 below — never appears onTicket), but for every value the two share, the ticket is the caller-visible source of truth and the session record is the AI-internal detail behind it. - Rationale:
Ticket.statusis what every other part of this codebase (agents, SLA, orchestration once built, the ticket-status contract in 003) already reads to know where a ticket stands — a session-only status field that never touchedTicket.statuswould make the ticket lie about its own state while an AI session was quietly doing something else internally. Reusing 003's state machine and itsupdateStatusmethod also means this feature inherits 003's already-tested transition validation and concurrency guarantee for free, rather than re-deriving both. - Alternatives considered: A session-only status with no
Ticket.statuslinkage — rejected per the rationale above. Building a second, AI-specific transition table — rejected; 003's table already defines exactly these states and edges, and doc 06'sAISupportSession.statusvalues were clearly authored to match it in the first place. - FR-023 implementation note:
ticketsService.updateStatusgains one additional check — when it's called with an actor other than'ai'(a human agent action) while anAISupportSessionfor that ticket is still active, the session is ended (status: ended_by_agent) as part of the same call, before the ticket's own status update commits. This is the concrete mechanism behind "a human agent takes ownership ends the AI session the same way an escalation does" (FR-023) — there's no separate "agent claims ticket" endpoint yet (orchestration/assignment is a later phase), so any human-actor status transition is the signal this feature has available today.
Decision: Runbook engine — the app selects the step, the model only phrases and interprets it
- Decision:
AISupportSessiongainsactiveRunbookKey String?andcurrentStepIndex Int?(refining doc 06's conceptualAISupportSession, same "refine during Phase 1 modeling" convention 004 already established forKnowledgeEntry/Runbook). When a diagnosis'sproblemTypematches an activeRunbook.keyfor the ticket's product (004'sRunbooksRepository.findCurrentByKey), the session setsactiveRunbookKeyandcurrentStepIndex = 0. The application readssteps[currentStepIndex]from the runbook's JSON and passes only that one step's content into the reasoning call's prompt as an instruction ("present this step, then interpret the customer's response against it") — the model is never given the full step list or asked to choose a step. On an outcome that means "try the next step," the app incrementscurrentStepIndexdeterministically; the model cannot set or advance the index itself (it has no tool for that). - Rationale: This is the literal requirement in FR-015/doc 03 §6 ("the AI cannot skip, reorder, or invent a step the runbook doesn't define") — the only way to guarantee that mechanically is for the index to be state the application owns and advances, with the model never seeing (and therefore never able to act on) any step but the current one.
- Alternatives considered: Giving the model the full runbook and trusting a system-prompt instruction ("only present steps in order") — rejected; doc 03 §6 explicitly says not to trust the LLM to "improvise" here, and a prompt instruction is not a guarantee, it's a request the model could depart from under distribution shift or adversarial input (doc 11 §A4).
Decision: Verification and resolution — a session can only close on tool evidence
- Decision:
AISupportSession.status(and, via the decision above,Ticket.status) transitionsverifying → resolved(AI_VERIFYING → AI_RESOLVED) only when averifyProductResolutiontool result exists on the session withconfirmed: true— which, given that tool's current fail-closed placeholder implementation (above), means resolution through this exact tool never actually auto-fires yet in this deployment. Customer confirmation (a message recorded during the session) is stored and surfaced in the escalation/resolution summary, but the status transition's guard checks tool evidence only, never message content. OnceTicket.statusreachesAI_RESOLVED, this feature's own responsibility ends — whether the ticket then moves toRESOLUTION_PENDING_CUSTOMERor straight toRESOLVEDis 003-ticketing's existing generic status-update surface, not something this feature further automates. - Rationale: FR-018/FR-019 verbatim — resolution requires verification evidence, customer confirmation is secondary-only. Guarding the state transition on a structured tool-result field (not on parsing what the AI "said" about the outcome) keeps this enforceable in code, not just in prompt instructions.
- Alternatives considered: Letting the reasoning call's own structured output include a
resolved: booleanfield the app trusts — rejected; this delegates a MUST-level policy decision (FR-018) to the model's own judgment, exactly the inversion Constitution Principle IV forbids.
Decision: Module placement — sessions, tools, troubleshooting, escalation under ai-support
- Decision: Four new submodules under the existing
src/modules/ai-support/group (which 004 created with onlyknowledgepopulated):sessions/(AISupportSession,AIDiagnosis,AIInteraction,AIConfidencePolicy, the two-call reasoning orchestration),tools/(the registry, the policy gate,AIAction/AIActionResult),troubleshooting/(the runbook-step engine, consuming 004'sRunbooksRepositorythroughknowledge's publicindex.ts),escalation/(hand-off summary construction, ending a session, ticket-ownership hand-off). Doc 07's fuller list (agents/diagnosis/tool-execution/verification) is intentionally not built as separate submodules —diagnosisandverificationare concerns insidesessionsandtoolsrespectively (anAIDiagnosisis produced by a session turn, not by an independent subsystem; verification is one tool's evaluated result, not a separate engine), andtool-executionis the same concern astoolssplit for no reason this feature's requirements give.agents(a registry of distinct AI "personas"/agent configs) has no requirement in spec.md at all — this feature has exactly one reasoning agent, not a multi-agent registry. - Rationale: Same reasoning 004 already used for not pre-building every doc 07 submodule speculatively — build what the current feature's FRs actually require, not the full documented taxonomy ahead of need.
- Alternatives considered: One flat
ai-support/agent/module holding everything — rejected; four genuinely distinct responsibilities (session/diagnosis orchestration, tool policy/ execution, runbook stepping, escalation hand-off) benefit from the same controller/service/repository separation every other module in this codebase already uses, and cross-module imports must go throughindex.tseither way (Principle III) — collapsing them into one module wouldn't reduce real coupling, just hide the boundaries.
Decision: Admin endpoint authentication and message visibility — reuse existing conventions
- Decision: The confidence-policy admin CRUD endpoint (
PUT /admin/products/:externalProductId/ai-policy, optionally.../categories/:categoryId/ai-policy) is gated byfastify.authenticate, same known-limitation stub as every prior admin surface. AI-authored messages are written through 003-ticketing's existingmessagesService.post(...)withtype: 'AI_MESSAGE'— that type and its customer-visible mapping already exist in the message-type→visibility map (specs/003-ticketing/research.md); this feature adds no new message-visibility rule. - Rationale: Consistency with established precedent; introducing a different auth mechanism or a parallel message-writing path for this feature alone would be unjustified inconsistency.
- Alternatives considered: None — direct reuse of existing, already-accepted conventions.