Files
msd-core/get-shit-done/templates/AI-SPEC.md
Tom Boucher a60e05c714 fix(claude): restore namespaced /gsd:<command> references (#3452)
* fix(claude): restore namespaced /gsd:<command> references

* test(claude): align slash-command expectations to /gsd: form

* test(claude): align generated command references to /gsd:

* test(claude): finish /gsd: namespace expectation updates
2026-05-12 21:53:24 -04:00

6.7 KiB

AI-SPEC — Phase {N}: {phase_name}

AI design contract generated by /gsd:ai-integration-phase. Consumed by gsd-planner and gsd-eval-auditor. Locks framework selection, implementation guidance, and evaluation strategy before planning begins.


1. System Classification

System Type:

Description:

Critical Failure Modes:


1b. Domain Context

Researched by gsd-domain-researcher. Grounds the evaluation strategy in domain expert knowledge.

Industry Vertical:

User Population:

Stakes Level:

Output Consequence:

What Domain Experts Evaluate Against

Known Failure Modes in This Domain

Regulatory / Compliance Context

Domain Expert Roles for Evaluation

Role Responsibility

2. Framework Decision

Selected Framework:

Version:

Rationale:

Alternatives Considered:

Framework Ruled Out Because

Vendor Lock-In Accepted:


3. Framework Quick Reference

Fetched from official docs by gsd-ai-researcher. Distilled for this specific use case.

Installation

# Install command(s)

Core Imports

# Key imports for this use case

Entry Point Pattern

# Minimal working example for this system type

Key Abstractions

Concept What It Is When You Use It

Common Pitfalls

project/
├── # Framework-specific folder layout

4. Implementation Guidance

Model Configuration:

Core Pattern:

Tool Use:

State Management:

Context Window Strategy:


4b. AI Systems Best Practices

Written by gsd-ai-researcher. Cross-cutting patterns every developer building AI systems needs — independent of framework choice.

Structured Outputs with Pydantic

# Pydantic output model for this system type

Async-First Design

Prompt Engineering Discipline

Context Window Management

Cost and Latency Budget


5. Evaluation Strategy

Dimensions

Dimension Rubric (Pass/Fail or 1-5) Measurement Approach Priority
Code / LLM Judge / Human Critical / High / Medium

Eval Tooling

Primary Tool:

Setup:

# Install and configure

CI/CD Integration:

# Command to run evals in CI/CD pipeline

Reference Dataset

Size:

Composition:

Labeling:


6. Guardrails

Online (Real-Time)

Guardrail Trigger Intervention
Block / Escalate / Flag

Offline (Flywheel)

Metric Sampling Strategy Action on Degradation

7. Production Monitoring

Tracing Tool:

Key Metrics to Track:

Alert Thresholds:

Smart Sampling Strategy:


Checklist

  • System type classified
  • Critical failure modes identified (≥ 3)
  • Domain context researched (Section 1b: vertical, stakes, expert criteria, failure modes)
  • Regulatory/compliance context identified or explicitly noted as none
  • Domain expert roles defined for evaluation involvement
  • Framework selected with rationale documented
  • Alternatives considered and ruled out
  • Framework quick reference written (install, imports, pattern, pitfalls)
  • AI systems best practices written (Section 4b: Pydantic, async, prompt discipline, context)
  • Evaluation dimensions grounded in domain rubric ingredients
  • Each eval dimension has a concrete rubric (Good/Bad in domain language)
  • Eval tooling selected — Arize Phoenix default confirmed or override noted
  • Reference dataset spec written (size ≥ 10, composition + labeling defined)
  • CI/CD eval integration specified
  • Online guardrails defined
  • Production monitoring configured (tracing tool + sampling strategy)