Files
support_backend/specs/015-reporting-dashboards/spec.md
T
saqib mirandClaude Sonnet 5 4a159725c2 docs(015-reporting-dashboards): plan, research, data model, contract, quickstart
Documents the exact Prisma query per dashboard figure, the one new
durable table this feature needs (ErrorCodeLookup — 014's own equivalent
metric is process-lifetime, unusable for a historical report), the
"no data -> null, never NaN" convention, and why the AI dashboard's
confidence distribution deliberately uses the system-default threshold
rather than resolving a per-diagnosis policy (AIDiagnosis has no
reliable FK back to which policy applied).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 11:14:29 +05:30

16 KiB

Feature Specification: Reporting and Analytics Dashboards

Feature Branch: 015-reporting-dashboards

Created: 2026-09-09

Status: Draft

Input: User description: "Reporting and analytics dashboards: real, read-only aggregation endpoints backing the four dashboards named in docs/09-testing-observability-cicd.md (Management, Product, Support, AI) — wiring up the pre-scaffolded but never-implemented platform/reports module into actual database-backed aggregation queries, admin-gated, with a date-range filter."

User Scenarios & Testing (mandatory)

User Story 1 - Management sees organization-wide support health (Priority: P1)

An admin or team lead opens a single view showing how support is doing overall for a chosen period: how many cases came in, how many were resolved (by AI vs. by a human), how many are still open, whether SLA commitments are being met, and how escalation is trending.

Why this priority: This is the one dashboard covering the whole roadmap's own top-level success criteria (docs/10-implementation-roadmap.md's checklist) in one place — the first thing anyone asks about a support operation is "how are we doing," and today there is no way to answer that except querying the database by hand.

Independent Test: Can be fully tested by creating a known set of tickets in various terminal states (AI-resolved, human-resolved, still open) plus a mix of met/breached SLA runs and escalations within a chosen date range, then requesting the Management dashboard for that range and confirming every figure matches what was actually created.

Acceptance Scenarios:

  1. Given a mix of tickets created within a chosen date range — some AI-resolved, some human-resolved, some still open — When the Management dashboard is requested for that range, Then total cases, AI-resolved count, human-escalated count, resolved count, and open count all match the actual data exactly.
  2. Given SLA runs that completed on time and others that breached within the range, When the dashboard is requested, Then SLA compliance (a rate) and SLA breach count both reflect the real outcomes.
  3. Given some tickets have a recorded first agent response and a resolution timestamp, When the dashboard is requested, Then average response time and average resolution time are computed only from tickets that actually reached those milestones within the range (a still-open ticket contributes to "open count" but never a fabricated resolution time).
  4. Given a date range with zero activity, When the dashboard is requested, Then every count is zero and every average is reported as "no data" rather than a computed zero or a division-by-zero error.

User Story 2 - See support broken down by product (Priority: P1)

An admin viewing support data for a specific product (or comparing products) sees volume, problem-type breakdown, which problems recur most, how well AI is resolving that product's issues versus escalating them, and which error codes come up most often.

Why this priority: SupportHub serves multiple SaaS products (Constitution Principle I); a number that isn't broken out by product hides which integration actually needs attention — this is as fundamental as the Management view, just sliced differently.

Independent Test: Can be fully tested by creating tickets/problems/error-code lookups across two distinct products within a date range, requesting the Product dashboard for each product, and confirming each one's figures include only its own product's data.

Acceptance Scenarios:

  1. Given tickets exist for two different products in the same date range, When the Product dashboard is requested scoped to one product, Then support volume and every other figure reflect only that product's tickets, never the other product's.
  2. Given problems in several categories for one product, When the dashboard is requested, Then the problem-type breakdown and "recurring problems" ranking both reflect the real category distribution, most-frequent first.
  3. Given a mix of AI-resolved and human-escalated tickets for one product, When the dashboard is requested, Then AI resolution rate and human escalation rate are both computed as a percentage of that product's own total, not the platform-wide total.
  4. Given several known-error-code lookups for one product, some codes looked up more than others, When the dashboard is requested, Then "top errors" lists those codes ranked by lookup frequency.

User Story 3 - Support sees team workload and performance (Priority: P2)

An admin or team lead sees how much work is currently assigned across agents, which tickets are at SLA risk, how much escalation is happening, and how quickly the team is responding to and resolving tickets.

Why this priority: This view is about ongoing operational load, not historical trend — useful for day-to-day team management, but the organization can already see whether it's healthy overall from User Story 1 without this one; P2 reflects that it adds an operational lens rather than a new class of information.

Independent Test: Can be fully tested by assigning several tickets to known agents (some close to SLA breach, some not), then requesting the Support dashboard and confirming workload per agent and the SLA-risk count both match reality.

Acceptance Scenarios:

  1. Given several tickets are currently assigned across two agents, When the Support dashboard is requested, Then each agent's current open-assignment count matches what was actually assigned to them (not a stale count from a previous, now-unassigned period).
  2. Given a ticket's SLA run is running and past a configurable risk threshold of its resolution due date (but not yet breached), When the dashboard is requested, Then it is counted as "at risk," distinct from both "on track" and "breached."
  3. Given response and resolution durations for several resolved tickets in the period, When the dashboard is requested, Then response-performance and resolution-performance figures are computed only from tickets that actually reached those milestones.

User Story 4 - See how well the AI is performing (Priority: P2)

An admin sees, for a chosen period, how often the AI resolves issues on its own versus escalating them, how often its attempted troubleshooting fails outright, how often it finds relevant knowledge, how confident its diagnoses tend to be, how reliably its tools succeed, and how often it ultimately hands off to a human.

Why this priority: This is the dashboard that validates the AI-first design's core premise (Constitution Principle IV) is actually working in practice — valuable, but a narrower audience than the org-wide and per-product views above, hence P2.

Independent Test: Can be fully tested by running several AI sessions to different terminal outcomes (resolved, escalated, escalated-after-failed-troubleshooting) with a mix of tool successes/failures and confidence levels recorded, then requesting the AI dashboard and confirming every figure matches the real session data.

Acceptance Scenarios:

  1. Given a mix of AI sessions ending resolved vs. escalated in the period, When the AI dashboard is requested, Then AI resolution rate and human-handoff rate both reflect the real outcome mix as percentages of total sessions.
  2. Given some AI tool invocations succeeded and others failed in the period, When the dashboard is requested, Then tool success/failure figures reflect the real invocation outcomes.
  3. Given diagnoses were recorded with a range of confidence values, When the dashboard is requested, Then the confidence distribution groups them into the same proceed/ask/escalate bands the AI support module's own confidence-policy service already classifies each diagnosis into (005-ai-support), not a newly-invented scheme.
  4. Given some AI sessions' knowledge-retrieval step found matching entries and others found none, When the dashboard is requested, Then knowledge-match rate reflects the real match/no-match mix.

Edge Cases

  • What happens when no from/to date range is given? Defaults to a reasonable trailing window (see Assumptions) rather than scanning the entire history unbounded on every request.
  • What happens when from is after to? Rejected as a validation error, not silently swapped or silently returning empty data.
  • What happens when a requested productId (Product dashboard) doesn't exist? Rejected with a clear not-found error, not an empty-but-200 response that looks like "this product has zero activity."
  • What happens when an average would divide by zero (no tickets reached that milestone in the range)? Reported as an explicit "no data" value, never NaN, null silently coerced to 0, or a thrown error.
  • What happens when a ticket's SLA run was paused for part of the period? SLA-risk/compliance figures use the run's own already-durable due dates (008-sla-escalation's pause/resume already accounts for paused time) rather than this feature re-deriving elapsed time itself.
  • Who can see these dashboards? Same admin-only gate as every other admin configuration/reporting surface introduced since 010-identity-auth — no new role is introduced.

Requirements (mandatory)

Functional Requirements

  • FR-001: System MUST provide a Management dashboard summarizing, for a given date range: total cases created, cases resolved by AI, cases escalated to a human, total resolved (either path), total still open, SLA compliance rate, SLA breach count, escalation count, average first-response time, and average resolution time.
  • FR-002: System MUST provide a Product dashboard summarizing, for a given date range and a specific product: support volume, a breakdown by problem category, a ranked list of the most recurring problem categories, AI resolution rate, human escalation rate, and a ranked list of the most frequently looked-up error codes.
  • FR-003: System MUST provide a Support dashboard summarizing, for a given date range: current per-agent open-assignment workload, count of tickets at SLA risk (past a configurable risk threshold of their resolution due date but not yet breached), count of tickets already breached, escalation count, average response performance, and average resolution performance.
  • FR-004: System MUST provide an AI dashboard summarizing, for a given date range: AI resolution rate, rate of sessions that escalated after at least one failed troubleshooting attempt, knowledge-match rate, a distribution of diagnosis confidence across the existing proceed/ask/escalate bands, tool invocation success/failure counts, and human-handoff rate.
  • FR-005: Every dashboard endpoint MUST accept an optional from/to date range; when omitted, it MUST default to a documented trailing window rather than scanning unbounded history.
  • FR-006: The Product dashboard MUST require a valid productId and MUST reject an unknown one with a clear not-found error rather than returning an empty-but-successful response.
  • FR-007: Every rate/average figure MUST be computed only from tickets/sessions/runs that actually reached the relevant milestone within the range; a metric with no qualifying data MUST be reported as an explicit "no data" value, never a computed 0, null, or NaN.
  • FR-008: All four dashboard endpoints MUST be admin-gated, consistent with every other admin-only reporting/configuration surface in this codebase.
  • FR-009: This feature MUST NOT alter the meaning or shape of any existing endpoint, event, or table — nearly every figure is derived read-only from data already durably recorded by the modules that own it (003 ticketing, 005 AI support, 007 orchestration, 008 SLA/escalation, 009 problem resolution). The one exception is FR-011: a small new durable record needed only because no existing table can answer "which error codes are looked up most" historically.
  • FR-011: System MUST durably record each known-error-code lookup (product, error code, timestamp) at the point it already happens (the existing error-code lookup call site) so the Product dashboard's "top errors" ranking (FR-002) can be computed historically — the equivalent live, in-process counter this project already exposes on /metrics (014-full- observability) is process-lifetime and reset on every restart, unusable for a historical dashboard.
  • FR-010: This feature is backend-only; presenting these figures in a UI is a separate, explicitly out-of-scope follow-on (see Assumptions).

Key Entities

  • Dashboard response: A read-only, computed JSON summary for one of the four dashboards over a requested date range (and, for the Product dashboard, one product) — never itself persisted; recomputed fresh on every request from existing durable records.
  • Date range: An inclusive from/to pair (calendar dates or timestamps) scoping every aggregation query; not a stored entity, a request parameter.
  • Error code lookup record (new, FR-011): a durable, append-only audit row — which product, which error code, when — written at the existing lookup call site; exists solely so "top errors" can be computed over a historical range, never read or written anywhere else.
  • Confidence band: The existing proceed/ask/escalate classification 005-ai-support already applies to a diagnosis's confidence score — reused here for the AI dashboard's distribution, not redefined.

Success Criteria (mandatory)

Measurable Outcomes

  • SC-001: For any chosen date range, every figure on all four dashboards can be independently verified against the underlying ticket/session/SLA-run/escalation-event records and matches exactly — no discrepancy between what a dashboard reports and what actually happened.
  • SC-002: An admin can answer "how is support doing right now" (Management), "how is this specific product doing" (Product), "who's overloaded and what's at risk" (Support), and "is the AI actually helping" (AI) each from a single request, with no manual database query needed.
  • SC-003: A dashboard request for a period with no matching activity returns clean, explicit "no data" results in well under a second — never an error, a stall, or a misleading zero.

Assumptions

  • Presentation is out of scope for this feature. The user's own explicit direction was to build the backend aggregation capability first (the established pattern this project has followed for every prior feature that touched both repos — identity/auth, the agent ticket queue, and full observability were each built backend-first). A supporthub-web dashboard UI consuming these endpoints is a natural, separate follow-on, not bundled into this spec.
  • The default trailing window when no date range is given is the last 30 days, matching common reporting-dashboard convention; CONFIGURABLE via the same admin-config env-driven pattern this project already uses for every other business-policy value (Constitution Principle II), not hardcoded as a magic number in application logic.
  • "SLA risk" needs a threshold (how close to the due date counts as "at risk") that the business has not specified — CONFIGURABLE, not invented as a hardcoded percentage, consistent with docs/10-implementation-roadmap.md's own "never hardcode a placeholder value and ship it as final" instruction.
  • These endpoints compute their figures synchronously, on request, directly from the existing tables — no new pre-aggregation table, no scheduled batch job, and no use of the pre-scaffolded ANALYTICS queue (src/jobs/analytics), which remains an inert stub outside this feature's scope. Live query performance at current data volumes is assumed adequate; a future feature can introduce pre-aggregation if and when it's actually needed (load/concurrency testing, a separate not-yet-started Phase 11 sub-area, is where that question would be validated).
  • "Top errors"/"recurring problems" rankings return a bounded top-N list (CONFIGURABLE limit, defaulting to 10) rather than the full distribution, matching how a dashboard is actually consumed.
  • Dashboard responses are computed fresh per request (no caching layer) — acceptable given the assumed data volumes and consistent with not prematurely optimizing ahead of the load-testing phase.