Files
msd-core/gsd-core/templates/AI-SPEC.md
Tom Boucher 463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00

247 lines
6.7 KiB
Markdown

# AI-SPEC — Phase {N}: {phase_name}
> AI design contract generated by `/gsd:ai-integration-phase`. Consumed by `gsd-planner` and `gsd-eval-auditor`.
> Locks framework selection, implementation guidance, and evaluation strategy before planning begins.
---
## 1. System Classification
**System Type:** <!-- RAG | Multi-Agent | Conversational | Extraction | Autonomous Agent | Content Generation | Code Automation | Hybrid -->
**Description:**
<!-- One-paragraph description of what this AI system does, who uses it, and what "good" looks like -->
**Critical Failure Modes:**
<!-- The 3-5 behaviors that absolutely cannot go wrong in this system -->
1.
2.
3.
---
## 1b. Domain Context
> Researched by `gsd-domain-researcher`. Grounds the evaluation strategy in domain expert knowledge.
**Industry Vertical:** <!-- healthcare | legal | finance | customer service | education | developer tooling | e-commerce | etc. -->
**User Population:** <!-- who uses this system and in what context -->
**Stakes Level:** <!-- Low | Medium | High | Critical -->
**Output Consequence:** <!-- what happens downstream when the AI output is acted on -->
### What Domain Experts Evaluate Against
<!-- Domain-specific rubric ingredients — in practitioner language, not AI jargon -->
<!-- Format: Dimension / Good (expert accepts) / Bad (expert flags) / Stakes / Source -->
### Known Failure Modes in This Domain
<!-- Domain-specific failure modes from research — not generic hallucination, but how it manifests here -->
### Regulatory / Compliance Context
<!-- Relevant regulations or constraints — or "None identified" if genuinely none apply -->
### Domain Expert Roles for Evaluation
| Role | Responsibility |
|------|---------------|
| <!-- e.g., Senior practitioner --> | <!-- Dataset labeling / rubric calibration / production sampling --> |
---
## 2. Framework Decision
**Selected Framework:** <!-- e.g., LlamaIndex v0.10.x -->
**Version:** <!-- Pin the version -->
**Rationale:**
<!-- Why this framework fits this system type, team context, and production requirements -->
**Alternatives Considered:**
| Framework | Ruled Out Because |
|-----------|------------------|
| | |
**Vendor Lock-In Accepted:** <!-- Yes / No / Partial — document the trade-off consciously -->
---
## 3. Framework Quick Reference
> Fetched from official docs by `gsd-ai-researcher`. Distilled for this specific use case.
### Installation
```bash
# Install command(s)
```
### Core Imports
```python
# Key imports for this use case
```
### Entry Point Pattern
```python
# Minimal working example for this system type
```
### Key Abstractions
<!-- Framework-specific concepts the developer must understand before coding -->
| Concept | What It Is | When You Use It |
|---------|-----------|-----------------|
| | | |
### Common Pitfalls
<!-- Gotchas specific to this framework and system type — from docs, issues, and community reports -->
1.
2.
3.
### Recommended Project Structure
```
project/
├── # Framework-specific folder layout
```
---
## 4. Implementation Guidance
**Model Configuration:**
<!-- Which model(s), temperature, max tokens, and other key parameters -->
**Core Pattern:**
<!-- The primary implementation pattern for this system type in this framework -->
**Tool Use:**
<!-- Tools/integrations needed and how to configure them -->
**State Management:**
<!-- How state is persisted, retrieved, and updated -->
**Context Window Strategy:**
<!-- How to manage context limits for this system type -->
---
## 4b. AI Systems Best Practices
> Written by `gsd-ai-researcher`. Cross-cutting patterns every developer building AI systems needs — independent of framework choice.
### Structured Outputs with Pydantic
<!-- Framework-specific Pydantic integration pattern for this use case -->
<!-- Include: output model definition, how the framework uses it, retry logic on validation failure -->
```python
# Pydantic output model for this system type
```
### Async-First Design
<!-- How async is handled in this framework, the one common mistake, and when to stream vs. await -->
### Prompt Engineering Discipline
<!-- System vs. user prompt separation, few-shot guidance, token budget strategy -->
### Context Window Management
<!-- Strategy specific to this system type: RAG chunking / conversation summarisation / agent compaction -->
### Cost and Latency Budget
<!-- Per-call cost estimate, caching strategy, sub-task model routing -->
---
## 5. Evaluation Strategy
### Dimensions
| Dimension | Rubric (Pass/Fail or 1-5) | Measurement Approach | Priority |
|-----------|--------------------------|---------------------|----------|
| | | Code / LLM Judge / Human | Critical / High / Medium |
### Eval Tooling
**Primary Tool:** <!-- e.g., RAGAS + Langfuse -->
**Setup:**
```bash
# Install and configure
```
**CI/CD Integration:**
```bash
# Command to run evals in CI/CD pipeline
```
### Reference Dataset
**Size:** <!-- e.g., 20 examples to start -->
**Composition:**
<!-- What scenario types the dataset covers: critical paths, edge cases, failure modes -->
**Labeling:**
<!-- Who labels examples and how (domain expert, LLM judge with calibration, etc.) -->
---
## 6. Guardrails
### Online (Real-Time)
| Guardrail | Trigger | Intervention |
|-----------|---------|--------------|
| | | Block / Escalate / Flag |
### Offline (Flywheel)
| Metric | Sampling Strategy | Action on Degradation |
|--------|------------------|----------------------|
| | | |
---
## 7. Production Monitoring
**Tracing Tool:** <!-- e.g., Langfuse self-hosted -->
**Key Metrics to Track:**
<!-- 3-5 metrics that will be monitored in production -->
**Alert Thresholds:**
<!-- When to page/alert -->
**Smart Sampling Strategy:**
<!-- How to select interactions for human review — signal-based filters -->
---
## Checklist
- [ ] System type classified
- [ ] Critical failure modes identified (≥ 3)
- [ ] Domain context researched (Section 1b: vertical, stakes, expert criteria, failure modes)
- [ ] Regulatory/compliance context identified or explicitly noted as none
- [ ] Domain expert roles defined for evaluation involvement
- [ ] Framework selected with rationale documented
- [ ] Alternatives considered and ruled out
- [ ] Framework quick reference written (install, imports, pattern, pitfalls)
- [ ] AI systems best practices written (Section 4b: Pydantic, async, prompt discipline, context)
- [ ] Evaluation dimensions grounded in domain rubric ingredients
- [ ] Each eval dimension has a concrete rubric (Good/Bad in domain language)
- [ ] Eval tooling selected — Arize Phoenix default confirmed or override noted
- [ ] Reference dataset spec written (size ≥ 10, composition + labeling defined)
- [ ] CI/CD eval integration specified
- [ ] Online guardrails defined
- [ ] Production monitoring configured (tracing tool + sampling strategy)