docs: spec for ticket creation, messages & attachments feature

/speckit-specify output for 003-ticketing (roadmap Phase 5): 3 user
stories (immediate ticket/problem creation with idempotency, typed
messages with enforced internal-note privacy, secure attachment
pipeline) and 15 functional requirements. Explicitly scoped to Phase
5 only -- investigation/root-cause/solution/resolution (Phase 9) and
AI diagnosis (Phase 4) are out of scope. Quality checklist passes
with no NEEDS CLARIFICATION markers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
saqib mir
2026-09-02 14:46:59 +05:30
co-authored by Claude Sonnet 5
parent 55253287b3
commit be2eb9907a
2 changed files with 278 additions and 0 deletions
@@ -0,0 +1,41 @@
# Specification Quality Checklist: Ticket Creation, Messages & Attachments
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-09-02
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- Scope is deliberately Phase 5 only (per docs/10-implementation-roadmap.md): ticket/problem
creation, messages, attachments, lifecycle state machine. Investigation/root cause/solution/
resolution (Phase 9) and AI diagnosis (Phase 4) are explicitly out of scope — see Assumptions.
- Recurring-problem matching is intentionally left to an explicit caller-supplied reference for
this feature; real fuzzy/semantic matching is deferred to the future AI-support feature.
- Malware-scanner choice and RLS adoption are left to `/speckit-plan` / business confirmation
respectively, not decided here.
- All items pass; no revision iterations were needed.
+237
View File
@@ -0,0 +1,237 @@
# Feature Specification: Ticket Creation, Messages & Attachments
**Feature Branch**: `003-ticketing`
**Created**: 2026-09-02
**Status**: Draft
**Input**: User description: "Phase 5 of docs/10-implementation-roadmap.md: Ticket + Problem
models (kept separate), message types, attachment pipeline (object storage, scanning, expiring
URLs), ticket lifecycle state machine. Per docs/04-ticketing-and-problem-management.md and
docs/06-database-schema.md."
## User Scenarios & Testing *(mandatory)*
### User Story 1 - A trusted request creates a durable ticket immediately (Priority: P1)
The moment a validated inbound request (from the SaaS integration trust boundary) describes a
customer's problem, SupportHub creates a durable ticket right away — before any diagnosis,
before any human is involved. The ticket starts in a `NEW` status. If the same underlying problem
recurs for the same product/tenant, the new ticket is linked to the existing `Problem` record
rather than creating a duplicate; if it's a genuinely new problem, a new `Problem` record is
created alongside the ticket. A retried request (same idempotency key) returns the
already-created ticket instead of creating a second one.
**Why this priority**: This is the foundational principle of the whole product ("ticket created
at the start of the journey, not after AI gives up") and everything else in this feature — and
every future feature (AI, orchestration, resolution) — depends on the ticket/problem records
existing first.
**Independent Test**: Send a valid inbound request through the trust boundary and confirm a
`Ticket` (status `NEW`) and a `Problem` exist immediately, correctly linked to the validated
product/tenant/user context; send the same request again with the same idempotency key and
confirm no second ticket is created.
**Acceptance Scenarios**:
1. **Given** a validated inbound request describing a problem, **When** it's processed, **Then**
a `Ticket` is created in `NEW` status and a `Problem` is created (or an existing one reused —
see Scenario 3), both linked to the validated product/tenant/user context, before any further
processing occurs.
2. **Given** a ticket was just created, **When** its record is inspected, **Then** it has a
human-referenceable code (e.g. `DQB-2026-00567`-style), the originating product, and the
trusted customer reference — never a raw, unvalidated value from the request.
3. **Given** a customer reports what is recognizably the same underlying problem again (same
product, same recognizable symptoms/context) as an existing open `Problem`, **When** a new
ticket is created for it, **Then** the new ticket links to the *existing* `Problem` record
rather than creating a duplicate one.
4. **Given** an inbound request carries an idempotency key that was already used for a
successfully created ticket, **When** the request is retried, **Then** the existing ticket is
returned and no second ticket or problem is created.
---
### User Story 2 - Ticket messages are typed, and internal notes are never visible to customers (Priority: P2)
A ticket accumulates a timeline of messages — from the customer, from AI (in future), from
agents, from the system, and internal-only notes (investigation/solution notes, general internal
notes). Every message has a type, and the API enforces — not just the UI — that
customer-invisible message types can never reach a customer-scoped read.
**Why this priority**: Without a message timeline there's no record of the interaction to show
anyone; without enforced internal-note privacy, an agent's private note becomes a customer-facing
leak the moment someone builds a UI that forgets to filter client-side.
**Independent Test**: Post one of each message type on a ticket, then read the ticket's messages
as a customer-scoped caller and confirm only customer-visible types appear; read the same
messages as an agent-scoped caller and confirm all types appear.
**Acceptance Scenarios**:
1. **Given** a ticket, **When** a message of any defined type is posted to it, **Then** it's
stored with its type, author reference, body, and a customer-visibility flag derived from its
type (never independently settable per-message in a way that contradicts the type).
2. **Given** a ticket has both customer-visible and internal-only messages, **When** its messages
are read through a customer-scoped endpoint, **Then** only customer-visible messages are
returned — internal notes are absent from the response entirely, not merely hidden by a flag.
3. **Given** the same ticket, **When** its messages are read through an agent-scoped endpoint,
**Then** every message, including internal notes, is returned.
4. **Given** a caller not authorized for a given ticket's tenant, **When** they attempt to read or
post a message on it, **Then** the request is rejected regardless of message type.
---
### User Story 3 - Attachments are safely stored and only ever downloaded through expiring, authorized URLs (Priority: P3)
A customer or agent can attach a file (screenshot, PDF, log, video, document) to a ticket. The
file goes to object storage, never to PostgreSQL — only its metadata and a storage reference are
stored in the database. It is not available for download until it has cleared a malware scan.
Every download happens through a short-lived, authorization-checked URL scoped to that ticket's
tenant/user context — never a permanent or unauthenticated link.
**Why this priority**: Attachments are common (screenshots, logs) but not required for the
minimum ticket flow to work, and getting the security properties right (never in Postgres, never
downloadable pre-scan, never a permanent link) matters more than shipping it first.
**Independent Test**: Upload a file to a ticket and confirm it's rejected for download until scan
status clears; confirm a generated download URL stops working after it expires; confirm a caller
outside the ticket's tenant cannot generate or use a download URL for it.
**Acceptance Scenarios**:
1. **Given** a file upload to a ticket, **When** it's outside the configured type/size limits,
**Then** it's rejected before being sent to object storage.
2. **Given** an accepted upload, **When** it has not yet cleared malware scanning, **Then** it is
not downloadable — its status is visibly `pending`, not silently unavailable.
3. **Given** a file that clears scanning, **When** an authorized caller requests to download it,
**Then** they receive a time-limited URL that stops working after it expires.
4. **Given** a file that fails malware scanning, **When** anyone attempts to download it,
**Then** the download is refused and the failure is visible on the attachment's record.
5. **Given** a caller outside the ticket's tenant/user context, **When** they attempt to generate
or use a download URL for one of its attachments, **Then** the request is rejected.
---
### Edge Cases
- What happens when two requests for the same new problem arrive concurrently (not a literal
retried idempotency key, but a genuine race — e.g. a flaky client double-submits without
reusing the idempotency key)? Out of scope to fully solve here beyond the idempotency-key
mechanism in User Story 1 — true duplicate-problem detection beyond exact idempotency-key reuse
is a knowledge/classification concern for a future AI feature, not this one.
- What happens when a ticket's status is updated by two actors at nearly the same time (e.g. a
customer reopens while an agent is closing)? The update that observes a stale status MUST be
rejected and retried against the current state — not silently overwrite the other actor's
change (Constitution Principle VII).
- What happens when an attachment upload is interrupted mid-transfer? The attachment record MUST
NOT be considered available; a resumed/retried upload is a new attempt, not a partial record
left in a downloadable-looking state.
- What happens when a malware scan itself fails to run (infrastructure error, not "found
malware")? The attachment MUST remain `pending`, never silently promoted to available.
- What happens when a message is posted with a type that doesn't exist in the defined set? The
request MUST be rejected — message type is not free text.
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: The system MUST create a `Ticket` record immediately upon processing a validated
inbound request describing a problem — before any diagnosis, classification, or human
involvement occurs.
- **FR-002**: The system MUST create or reuse a `Problem` record for every ticket: a new `Problem`
when none matches, or the existing `Problem` when the ticket represents a recurrence of one
already open for the same product/tenant context.
- **FR-003**: `Ticket` and `Problem` MUST remain separate, related entities — a ticket references
exactly one problem; a problem may have many tickets (Constitution Principle VIII).
- **FR-004**: The system MUST honor the inbound request's idempotency key: a retried request
carrying a previously-used key returns the already-created ticket rather than creating a new
one.
- **FR-005**: Every ticket MUST have a human-referenceable code, unique, generated by the system
— never supplied by the caller.
- **FR-006**: A ticket's status MUST only ever be one of the defined lifecycle states, and MUST
only transition through valid state changes (an invalid transition is rejected, not silently
coerced).
- **FR-007**: A ticket's status update MUST use optimistic concurrency control: an update based on
a stale prior status is rejected, not applied on top of a change it didn't observe.
- **FR-008**: Every message posted to a ticket MUST have one of the defined message types, and its
customer-visibility MUST be determined by its type, not independently settable in a way that
contradicts the type.
- **FR-009**: The system MUST NOT return customer-invisible message types (internal notes,
investigation notes, solution notes) through any customer-scoped read of a ticket's messages —
enforced at the API/serialization layer, not left to client-side filtering.
- **FR-010**: The system MUST reject reading or posting on a ticket by a caller not authorized for
that ticket's tenant/user context, regardless of message type or attachment involved.
- **FR-011**: Attachment files MUST be stored in object storage, never in PostgreSQL — the
database stores only metadata and a storage reference.
- **FR-012**: The system MUST validate an attachment's type and size against configured limits
before accepting the upload.
- **FR-013**: An uploaded attachment MUST NOT be downloadable until it has cleared malware
scanning; its scan status MUST be visible on its record (`pending` / `clean` / `infected` /
`rejected`).
- **FR-014**: Attachment downloads MUST only be possible through a time-limited,
authorization-checked URL scoped to the ticket's tenant/user context — never a permanent or
unauthenticated link.
- **FR-015**: The system MUST scope every ticket, message, and attachment query by the caller's
validated tenant/user context — a caller-supplied identifier alone is never sufficient
authorization (Constitution Principle I).
### Key Entities
- **Ticket**: The durable, operational record of one support interaction — status, priority,
severity, the product/tenant/customer it belongs to, and links to its problem, messages,
attachments, and (in later features) AI sessions, assignments, and escalation events.
- **Problem**: The underlying issue being solved, which can outlive and span multiple tickets —
statement, symptoms, impact, severity, environment. Deliberately separate from `Ticket`.
- **Ticket Message**: One entry in a ticket's timeline — typed (`CUSTOMER_MESSAGE`, `AI_MESSAGE`,
`AGENT_MESSAGE`, `INTERNAL_NOTE`, `SYSTEM_EVENT`, `INVESTIGATION_NOTE`, `SOLUTION_NOTE`), with
an author reference, body, and a visibility derived from its type.
- **Ticket Attachment**: Metadata for one uploaded file — storage reference (never the file
itself), original filename, MIME type, size, scan status, and who uploaded it.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: 100% of validated inbound requests result in a ticket existing before any further
processing — zero requests that pass the trust boundary without a corresponding ticket.
- **SC-002**: A retried request using the same idempotency key never produces more than one
ticket, regardless of how many times it's retried.
- **SC-003**: Zero internal-only messages ever appear in a customer-scoped read of a ticket's
timeline, verified across every defined message type.
- **SC-004**: Zero attachment files are ever persisted directly in the database — 100% go to
object storage with only a reference stored.
- **SC-005**: An attachment that hasn't cleared malware scanning is rejected for download 100% of
the time it's attempted.
- **SC-006**: A generated attachment download URL becomes unusable after its configured expiry —
verified by attempting to use it past that point.
- **SC-007**: Two concurrent status updates to the same ticket never both apply silently — exactly
one succeeds against the state it observed, and the other is rejected and must retry.
## Assumptions
- This feature covers ticket/problem creation, the message timeline, and the attachment pipeline
only. It does NOT include: AI diagnosis/classification (Phase 4), investigation/root
cause/solution/resolution workflows (Phase 9), orchestration/assignment/SLA (Phases 6-8), or the
customer-confirmation/auto-close/reopen *workflow* automation (Phase 9) — though the `REOPENED`
status itself is part of the lifecycle state machine this feature defines, since doc 04 lists it
as a core ticket status.
- "Recognizably the same underlying problem" (FR-002/User Story 1 Scenario 3) is intentionally
left without a precise matching algorithm here — real recurring-problem detection is a
knowledge/classification capability that belongs to the AI-support feature (Phase 4). For this
feature, an explicit, caller-supplied reference (e.g. a prior ticket/problem id in the inbound
request's `referenceIds`, per docs/02 §3) is sufficient grounds to link to an existing `Problem`
— this feature does not attempt fuzzy/semantic matching on its own.
- AI-driven status transitions (`AI_ANALYZING`, `AI_TROUBLESHOOTING`, `AI_VERIFYING`,
`AI_RESOLVED`) are part of the lifecycle state machine's defined states (FR-006), but nothing in
this feature *automatically drives* a ticket into them — that requires the AI-support feature
(Phase 4), which doesn't exist yet. This feature only guarantees the state machine itself is
correct and that transitions can be triggered (e.g. by an authorized caller or a future feature)
without corrupting ticket state under concurrency.
- Row-level security (Postgres RLS) as a defense-in-depth layer under the application-level tenant
scoping in FR-015 (per `docs/11-architect-additions-gaps-and-recommendations.md` §A3) is a
valuable hardening step but is deliberately deferred — `REQUIRES BUSINESS/PLATFORM
CONFIRMATION` on whether/when to adopt it, not invented here. Application-level scoping (FR-015)
is the enforced control for this feature.
- Malware scanning integration specifics (which scanner/service) are a technical decision left to
planning — this spec only requires that the scan gate and its visible states exist.