* fix(claude): restore namespaced /gsd:<command> references * test(claude): align slash-command expectations to /gsd: form * test(claude): align generated command references to /gsd: * test(claude): finish /gsd: namespace expectation updates
6.7 KiB
AI-SPEC — Phase {N}: {phase_name}
AI design contract generated by
/gsd:ai-integration-phase. Consumed bygsd-plannerandgsd-eval-auditor. Locks framework selection, implementation guidance, and evaluation strategy before planning begins.
1. System Classification
System Type:
Description:
Critical Failure Modes:
1b. Domain Context
Researched by
gsd-domain-researcher. Grounds the evaluation strategy in domain expert knowledge.
Industry Vertical:
User Population:
Stakes Level:
Output Consequence:
What Domain Experts Evaluate Against
Known Failure Modes in This Domain
Regulatory / Compliance Context
Domain Expert Roles for Evaluation
| Role | Responsibility |
|---|---|
2. Framework Decision
Selected Framework:
Version:
Rationale:
Alternatives Considered:
| Framework | Ruled Out Because |
|---|---|
Vendor Lock-In Accepted:
3. Framework Quick Reference
Fetched from official docs by
gsd-ai-researcher. Distilled for this specific use case.
Installation
# Install command(s)
Core Imports
# Key imports for this use case
Entry Point Pattern
# Minimal working example for this system type
Key Abstractions
| Concept | What It Is | When You Use It |
|---|---|---|
Common Pitfalls
Recommended Project Structure
project/
├── # Framework-specific folder layout
4. Implementation Guidance
Model Configuration:
Core Pattern:
Tool Use:
State Management:
Context Window Strategy:
4b. AI Systems Best Practices
Written by
gsd-ai-researcher. Cross-cutting patterns every developer building AI systems needs — independent of framework choice.
Structured Outputs with Pydantic
# Pydantic output model for this system type
Async-First Design
Prompt Engineering Discipline
Context Window Management
Cost and Latency Budget
5. Evaluation Strategy
Dimensions
| Dimension | Rubric (Pass/Fail or 1-5) | Measurement Approach | Priority |
|---|---|---|---|
| Code / LLM Judge / Human | Critical / High / Medium |
Eval Tooling
Primary Tool:
Setup:
# Install and configure
CI/CD Integration:
# Command to run evals in CI/CD pipeline
Reference Dataset
Size:
Composition:
Labeling:
6. Guardrails
Online (Real-Time)
| Guardrail | Trigger | Intervention |
|---|---|---|
| Block / Escalate / Flag |
Offline (Flywheel)
| Metric | Sampling Strategy | Action on Degradation |
|---|---|---|
7. Production Monitoring
Tracing Tool:
Key Metrics to Track:
Alert Thresholds:
Smart Sampling Strategy:
Checklist
- System type classified
- Critical failure modes identified (≥ 3)
- Domain context researched (Section 1b: vertical, stakes, expert criteria, failure modes)
- Regulatory/compliance context identified or explicitly noted as none
- Domain expert roles defined for evaluation involvement
- Framework selected with rationale documented
- Alternatives considered and ruled out
- Framework quick reference written (install, imports, pattern, pitfalls)
- AI systems best practices written (Section 4b: Pydantic, async, prompt discipline, context)
- Evaluation dimensions grounded in domain rubric ingredients
- Each eval dimension has a concrete rubric (Good/Bad in domain language)
- Eval tooling selected — Arize Phoenix default confirmed or override noted
- Reference dataset spec written (size ≥ 10, composition + labeling defined)
- CI/CD eval integration specified
- Online guardrails defined
- Production monitoring configured (tracing tool + sampling strategy)