# AI-SPEC — Phase {N}: {phase_name} > AI design contract generated by `/gsd-ai-integration-phase`. Consumed by `gsd-planner` and `gsd-eval-auditor`. > Locks framework selection, implementation guidance, and evaluation strategy before planning begins. --- ## 1. System Classification **System Type:** **Description:** **Critical Failure Modes:** 1. 2. 3. --- ## 1b. Domain Context > Researched by `gsd-domain-researcher`. Grounds the evaluation strategy in domain expert knowledge. **Industry Vertical:** **User Population:** **Stakes Level:** **Output Consequence:** ### What Domain Experts Evaluate Against ### Known Failure Modes in This Domain ### Regulatory / Compliance Context ### Domain Expert Roles for Evaluation | Role | Responsibility | |------|---------------| | | | --- ## 2. Framework Decision **Selected Framework:** **Version:** **Rationale:** **Alternatives Considered:** | Framework | Ruled Out Because | |-----------|------------------| | | | **Vendor Lock-In Accepted:** --- ## 3. Framework Quick Reference > Fetched from official docs by `gsd-ai-researcher`. Distilled for this specific use case. ### Installation ```bash # Install command(s) ``` ### Core Imports ```python # Key imports for this use case ``` ### Entry Point Pattern ```python # Minimal working example for this system type ``` ### Key Abstractions | Concept | What It Is | When You Use It | |---------|-----------|-----------------| | | | | ### Common Pitfalls 1. 2. 3. ### Recommended Project Structure ``` project/ ├── # Framework-specific folder layout ``` --- ## 4. Implementation Guidance **Model Configuration:** **Core Pattern:** **Tool Use:** **State Management:** **Context Window Strategy:** --- ## 4b. AI Systems Best Practices > Written by `gsd-ai-researcher`. Cross-cutting patterns every developer building AI systems needs — independent of framework choice. ### Structured Outputs with Pydantic ```python # Pydantic output model for this system type ``` ### Async-First Design ### Prompt Engineering Discipline ### Context Window Management ### Cost and Latency Budget --- ## 5. Evaluation Strategy ### Dimensions | Dimension | Rubric (Pass/Fail or 1-5) | Measurement Approach | Priority | |-----------|--------------------------|---------------------|----------| | | | Code / LLM Judge / Human | Critical / High / Medium | ### Eval Tooling **Primary Tool:** **Setup:** ```bash # Install and configure ``` **CI/CD Integration:** ```bash # Command to run evals in CI/CD pipeline ``` ### Reference Dataset **Size:** **Composition:** **Labeling:** --- ## 6. Guardrails ### Online (Real-Time) | Guardrail | Trigger | Intervention | |-----------|---------|--------------| | | | Block / Escalate / Flag | ### Offline (Flywheel) | Metric | Sampling Strategy | Action on Degradation | |--------|------------------|----------------------| | | | | --- ## 7. Production Monitoring **Tracing Tool:** **Key Metrics to Track:** **Alert Thresholds:** **Smart Sampling Strategy:** --- ## Checklist - [ ] System type classified - [ ] Critical failure modes identified (≥ 3) - [ ] Domain context researched (Section 1b: vertical, stakes, expert criteria, failure modes) - [ ] Regulatory/compliance context identified or explicitly noted as none - [ ] Domain expert roles defined for evaluation involvement - [ ] Framework selected with rationale documented - [ ] Alternatives considered and ruled out - [ ] Framework quick reference written (install, imports, pattern, pitfalls) - [ ] AI systems best practices written (Section 4b: Pydantic, async, prompt discipline, context) - [ ] Evaluation dimensions grounded in domain rubric ingredients - [ ] Each eval dimension has a concrete rubric (Good/Bad in domain language) - [ ] Eval tooling selected — Arize Phoenix default confirmed or override noted - [ ] Reference dataset spec written (size ≥ 10, composition + labeling defined) - [ ] CI/CD eval integration specified - [ ] Online guardrails defined - [ ] Production monitoring configured (tracing tool + sampling strategy)