docs(015-reporting-dashboards): feature spec and quality checklist

Phase 11's third sub-area (reporting/analytics dashboards), per explicit
user direction. Backend-first scope (four read-only aggregation
endpoints wiring up the pre-scaffolded platform/reports module),
following the same backend-before-frontend pattern already established
for 010/011/014 this session — a supporthub-web dashboard UI is a
separate, not-yet-started follow-on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
saqib mir
2026-09-09 11:10:00 +05:30
co-authored by Claude Sonnet 5
parent ea50e3596a
commit c4a2faa6e3
2 changed files with 295 additions and 0 deletions
@@ -0,0 +1,49 @@
# Specification Quality Checklist: Reporting and Analytics Dashboards
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-09-09
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- This is `docs/10-implementation-roadmap.md`'s own Phase 11, third sub-area, per explicit user
direction (013 was the security pass, 014 was full observability). Backend-first scope
(Assumptions) follows the same pattern already established three times this session
(010-identity-auth, 011-agent-ticket-queue, and 014-full-observability's own frontend-free
scope) — a `supporthub-web` dashboard UI is a natural, separate follow-on, not re-litigated
here via a fresh question.
- The pre-scaffolded-but-inert `platform/reports` module (`ReportsService.generateSummaryReport`
currently returns `{}`) and the `ANALYTICS` queue stub (`src/jobs/analytics`, logs only) were
both confirmed via direct code inspection before writing this spec — the same
"provisioned before this session's rebuild but never wired up" pattern found repeatedly this
session. This feature wires up the former; the Assumptions section explicitly keeps the latter
out of scope (synchronous queries, no pre-aggregation job, for this first cut).
- All items pass; no revision iterations were needed. No [NEEDS CLARIFICATION] markers were
required — every open question (default date window, SLA-risk threshold, top-N limit) had a
reasonable, documented, CONFIGURABLE default (see Assumptions), matching the roadmap's own
"never hardcode a placeholder value and ship it as final" instruction.
+246
View File
@@ -0,0 +1,246 @@
# Feature Specification: Reporting and Analytics Dashboards
**Feature Branch**: `015-reporting-dashboards`
**Created**: 2026-09-09
**Status**: Draft
**Input**: User description: "Reporting and analytics dashboards: real, read-only aggregation endpoints backing the four dashboards named in docs/09-testing-observability-cicd.md (Management, Product, Support, AI) — wiring up the pre-scaffolded but never-implemented platform/reports module into actual database-backed aggregation queries, admin-gated, with a date-range filter."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Management sees organization-wide support health (Priority: P1)
An admin or team lead opens a single view showing how support is doing overall for a chosen
period: how many cases came in, how many were resolved (by AI vs. by a human), how many are
still open, whether SLA commitments are being met, and how escalation is trending.
**Why this priority**: This is the one dashboard covering the whole roadmap's own top-level
success criteria (`docs/10-implementation-roadmap.md`'s checklist) in one place — the first
thing anyone asks about a support operation is "how are we doing," and today there is no way to
answer that except querying the database by hand.
**Independent Test**: Can be fully tested by creating a known set of tickets in various terminal
states (AI-resolved, human-resolved, still open) plus a mix of met/breached SLA runs and
escalations within a chosen date range, then requesting the Management dashboard for that range
and confirming every figure matches what was actually created.
**Acceptance Scenarios**:
1. **Given** a mix of tickets created within a chosen date range — some AI-resolved, some
human-resolved, some still open — **When** the Management dashboard is requested for that
range, **Then** total cases, AI-resolved count, human-escalated count, resolved count, and
open count all match the actual data exactly.
2. **Given** SLA runs that completed on time and others that breached within the range,
**When** the dashboard is requested, **Then** SLA compliance (a rate) and SLA breach count
both reflect the real outcomes.
3. **Given** some tickets have a recorded first agent response and a resolution timestamp,
**When** the dashboard is requested, **Then** average response time and average resolution
time are computed only from tickets that actually reached those milestones within the range
(a still-open ticket contributes to "open count" but never a fabricated resolution time).
4. **Given** a date range with zero activity, **When** the dashboard is requested, **Then** every
count is zero and every average is reported as "no data" rather than a computed zero or a
division-by-zero error.
---
### User Story 2 - See support broken down by product (Priority: P1)
An admin viewing support data for a specific product (or comparing products) sees volume,
problem-type breakdown, which problems recur most, how well AI is resolving that product's
issues versus escalating them, and which error codes come up most often.
**Why this priority**: SupportHub serves multiple SaaS products (Constitution Principle I); a
number that isn't broken out by product hides which integration actually needs attention — this
is as fundamental as the Management view, just sliced differently.
**Independent Test**: Can be fully tested by creating tickets/problems/error-code lookups across
two distinct products within a date range, requesting the Product dashboard for each product,
and confirming each one's figures include only its own product's data.
**Acceptance Scenarios**:
1. **Given** tickets exist for two different products in the same date range, **When** the
Product dashboard is requested scoped to one product, **Then** support volume and every other
figure reflect only that product's tickets, never the other product's.
2. **Given** problems in several categories for one product, **When** the dashboard is
requested, **Then** the problem-type breakdown and "recurring problems" ranking both reflect
the real category distribution, most-frequent first.
3. **Given** a mix of AI-resolved and human-escalated tickets for one product, **When** the
dashboard is requested, **Then** AI resolution rate and human escalation rate are both
computed as a percentage of that product's own total, not the platform-wide total.
4. **Given** several known-error-code lookups for one product, some codes looked up more than
others, **When** the dashboard is requested, **Then** "top errors" lists those codes ranked by
lookup frequency.
---
### User Story 3 - Support sees team workload and performance (Priority: P2)
An admin or team lead sees how much work is currently assigned across agents, which tickets are
at SLA risk, how much escalation is happening, and how quickly the team is responding to and
resolving tickets.
**Why this priority**: This view is about ongoing operational load, not historical trend — useful
for day-to-day team management, but the organization can already see whether it's healthy
overall from User Story 1 without this one; P2 reflects that it adds an operational lens rather
than a new class of information.
**Independent Test**: Can be fully tested by assigning several tickets to known agents (some
close to SLA breach, some not), then requesting the Support dashboard and confirming workload
per agent and the SLA-risk count both match reality.
**Acceptance Scenarios**:
1. **Given** several tickets are currently assigned across two agents, **When** the Support
dashboard is requested, **Then** each agent's current open-assignment count matches what was
actually assigned to them (not a stale count from a previous, now-unassigned period).
2. **Given** a ticket's SLA run is running and past a configurable risk threshold of its
resolution due date (but not yet breached), **When** the dashboard is requested, **Then** it
is counted as "at risk," distinct from both "on track" and "breached."
3. **Given** response and resolution durations for several resolved tickets in the period,
**When** the dashboard is requested, **Then** response-performance and resolution-performance
figures are computed only from tickets that actually reached those milestones.
---
### User Story 4 - See how well the AI is performing (Priority: P2)
An admin sees, for a chosen period, how often the AI resolves issues on its own versus escalating
them, how often its attempted troubleshooting fails outright, how often it finds relevant
knowledge, how confident its diagnoses tend to be, how reliably its tools succeed, and how often
it ultimately hands off to a human.
**Why this priority**: This is the dashboard that validates the AI-first design's core premise
(Constitution Principle IV) is actually working in practice — valuable, but a narrower audience
than the org-wide and per-product views above, hence P2.
**Independent Test**: Can be fully tested by running several AI sessions to different terminal
outcomes (resolved, escalated, escalated-after-failed-troubleshooting) with a mix of tool
successes/failures and confidence levels recorded, then requesting the AI dashboard and
confirming every figure matches the real session data.
**Acceptance Scenarios**:
1. **Given** a mix of AI sessions ending resolved vs. escalated in the period, **When** the AI
dashboard is requested, **Then** AI resolution rate and human-handoff rate both reflect the
real outcome mix as percentages of total sessions.
2. **Given** some AI tool invocations succeeded and others failed in the period, **When** the
dashboard is requested, **Then** tool success/failure figures reflect the real invocation
outcomes.
3. **Given** diagnoses were recorded with a range of confidence values, **When** the dashboard is
requested, **Then** the confidence distribution groups them into the same high/medium/low
bands the AI support module itself already uses (005-ai-support), not a newly-invented scheme.
4. **Given** some AI sessions' knowledge-retrieval step found matching entries and others found
none, **When** the dashboard is requested, **Then** knowledge-match rate reflects the real
match/no-match mix.
---
### Edge Cases
- What happens when no `from`/`to` date range is given? Defaults to a reasonable trailing window
(see Assumptions) rather than scanning the entire history unbounded on every request.
- What happens when `from` is after `to`? Rejected as a validation error, not silently swapped or
silently returning empty data.
- What happens when a requested `productId` (Product dashboard) doesn't exist? Rejected with a
clear not-found error, not an empty-but-200 response that looks like "this product has zero
activity."
- What happens when an average would divide by zero (no tickets reached that milestone in the
range)? Reported as an explicit "no data" value, never `NaN`, `null` silently coerced to `0`,
or a thrown error.
- What happens when a ticket's SLA run was paused for part of the period? SLA-risk/compliance
figures use the run's own already-durable due dates (008-sla-escalation's pause/resume
already accounts for paused time) rather than this feature re-deriving elapsed time itself.
- Who can see these dashboards? Same admin-only gate as every other admin configuration/reporting
surface introduced since 010-identity-auth — no new role is introduced.
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST provide a Management dashboard summarizing, for a given date range:
total cases created, cases resolved by AI, cases escalated to a human, total resolved
(either path), total still open, SLA compliance rate, SLA breach count, escalation count,
average first-response time, and average resolution time.
- **FR-002**: System MUST provide a Product dashboard summarizing, for a given date range and a
specific product: support volume, a breakdown by problem category, a ranked list of the most
recurring problem categories, AI resolution rate, human escalation rate, and a ranked list of
the most frequently looked-up error codes.
- **FR-003**: System MUST provide a Support dashboard summarizing, for a given date range:
current per-agent open-assignment workload, count of tickets at SLA risk (past a configurable
risk threshold of their resolution due date but not yet breached), count of tickets already
breached, escalation count, average response performance, and average resolution performance.
- **FR-004**: System MUST provide an AI dashboard summarizing, for a given date range: AI
resolution rate, rate of sessions that escalated after at least one failed troubleshooting
attempt, knowledge-match rate, a distribution of diagnosis confidence across the existing
high/medium/low bands, tool invocation success/failure counts, and human-handoff rate.
- **FR-005**: Every dashboard endpoint MUST accept an optional `from`/`to` date range; when
omitted, it MUST default to a documented trailing window rather than scanning unbounded
history.
- **FR-006**: The Product dashboard MUST require a valid `productId` and MUST reject an unknown
one with a clear not-found error rather than returning an empty-but-successful response.
- **FR-007**: Every rate/average figure MUST be computed only from tickets/sessions/runs that
actually reached the relevant milestone within the range; a metric with no qualifying data MUST
be reported as an explicit "no data" value, never a computed `0`, `null`, or `NaN`.
- **FR-008**: All four dashboard endpoints MUST be admin-gated, consistent with every other
admin-only reporting/configuration surface in this codebase.
- **FR-009**: This feature MUST NOT alter what any existing endpoint, event, or table stores —
every figure is derived read-only from data already durably recorded by the modules that own
it (003 ticketing, 005 AI support, 007 orchestration, 008 SLA/escalation, 009 problem
resolution).
- **FR-010**: This feature is backend-only; presenting these figures in a UI is a separate,
explicitly out-of-scope follow-on (see Assumptions).
### Key Entities
- **Dashboard response**: A read-only, computed JSON summary for one of the four dashboards over
a requested date range (and, for the Product dashboard, one product) — never itself persisted;
recomputed fresh on every request from existing durable records.
- **Date range**: An inclusive `from`/`to` pair (calendar dates or timestamps) scoping every
aggregation query; not a stored entity, a request parameter.
- **Confidence band**: The existing high/medium/low classification 005-ai-support already applies
to a diagnosis's confidence score — reused here for the AI dashboard's distribution, not
redefined.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: For any chosen date range, every figure on all four dashboards can be independently
verified against the underlying ticket/session/SLA-run/escalation-event records and matches
exactly — no discrepancy between what a dashboard reports and what actually happened.
- **SC-002**: An admin can answer "how is support doing right now" (Management), "how is this
specific product doing" (Product), "who's overloaded and what's at risk" (Support), and "is the
AI actually helping" (AI) each from a single request, with no manual database query needed.
- **SC-003**: A dashboard request for a period with no matching activity returns clean, explicit
"no data" results in well under a second — never an error, a stall, or a misleading zero.
## Assumptions
- **Presentation is out of scope for this feature.** The user's own explicit direction was to
build the backend aggregation capability first (the established pattern this project has
followed for every prior feature that touched both repos — identity/auth, the agent ticket
queue, and full observability were each built backend-first). A `supporthub-web` dashboard UI
consuming these endpoints is a natural, separate follow-on, not bundled into this spec.
- The default trailing window when no date range is given is the last 30 days, matching common
reporting-dashboard convention; CONFIGURABLE via the same admin-config env-driven pattern this
project already uses for every other business-policy value (Constitution Principle II), not
hardcoded as a magic number in application logic.
- "SLA risk" needs a threshold (how close to the due date counts as "at risk") that the business
has not specified — CONFIGURABLE, not invented as a hardcoded percentage, consistent with
`docs/10-implementation-roadmap.md`'s own "never hardcode a placeholder value and ship it as
final" instruction.
- These endpoints compute their figures synchronously, on request, directly from the existing
tables — no new pre-aggregation table, no scheduled batch job, and no use of the pre-scaffolded
`ANALYTICS` queue (`src/jobs/analytics`), which remains an inert stub outside this feature's
scope. Live query performance at current data volumes is assumed adequate; a future feature can
introduce pre-aggregation if and when it's actually needed (load/concurrency testing, a
separate not-yet-started Phase 11 sub-area, is where that question would be validated).
- "Top errors"/"recurring problems" rankings return a bounded top-N list (CONFIGURABLE limit,
defaulting to 10) rather than the full distribution, matching how a dashboard is actually
consumed.
- Dashboard responses are computed fresh per request (no caching layer) — acceptable given the
assumed data volumes and consistent with not prematurely optimizing ahead of the load-testing
phase.