* feat(templates): add optional Business Context section to PROJECT.md template
Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.
Refs #72
* chore(changeset): set pr number for #72 fragment
* test(#72): add source-text-is-the-product exemption marker
Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#991): inject configured agent_skills into code-review family subagents
code-review.md, code-review-fix.md, and eval-review.md spawned their
subagents (gsd-code-reviewer / gsd-code-fixer / gsd-eval-auditor) without
querying or injecting the project-configured agent_skills, while ~20 sibling
workflows do. Subagents don't inherit the orchestrator's auto-loaded context,
so this injection is the only channel — reviewers/fixers/auditors silently ran
without the configured rule/skill context.
Mirror the established sibling idiom: add
`VAR=$(gsd_run query agent-skills <agent-type>)` in each workflow's initialize
step and interpolate `${VAR}` into every Agent() spawn of that type. This
covers all spawn sites, including code-review-fix.md's --auto loop which
re-spawns gsd-code-reviewer in addition to the two gsd-code-fixer spawns.
Regression test reads the workflow text (source-text-is-the-product) and
asserts each file queries agent-skills for every agent type it spawns and
interpolates the result at least once per spawn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#991): add changeset for code-review agent_skills injection fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works
The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.
Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.
Closes#978
* chore(#978): backfill changeset pr number (982)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)
Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).
Closes#985
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)
The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.
Closes#981
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.
- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
reproduces the removed case EXACTLY (query +--budget, status, diff, build,
hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
config:{graphify.enabled default false}, commands:[{family:graphify, module,
router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.
Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.
Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.
Closes#972
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1)
Build the capability command mechanism per ADR-959: the `commands` declaration
field on the feature role, the registry commandFamilies index, and a real
dispatchCapabilityCommand consulted in runCommand's default case (replacing the
dead _dispatchNonFamily shim's role).
The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a
default-case command, dispatch enforces a bare-.cjs-basename, resolves the module
under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and
own-property-guards the router export before calling it. A require.main===module
guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged.
Additive: commandFamilies is empty today, so the default case is behavior-
preserving for every command; the 10 dead _dispatchNonFamily sites are untouched
(future migration markers). The graphify cutover is the separate next step.
Closes#961
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#961): surface capability router failures structurally + enforce sync contract
Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an
unexpected (non-ExitError) throw is converted to a structured, attributed
error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw
stack trace; an intentional ExitError propagates unchanged. An async router
(returns a thenable) is rejected loudly with a structured error (the contract is
synchronous, like the 12 host routers). Not shipped as "consistent with existing
behavior": the host's pre-existing version of this gap is filed as #965.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#950): emit status: complete in quick-task SUMMARY frontmatter
Add `status: complete` to all four SUMMARY templates (summary.md,
summary-minimal.md, summary-standard.md, summary-complex.md), to the
executor agent's documented frontmatter field list, and to the quick.md
executor constraints block. The audit-open milestone-close scanner
(scanQuickTasks) reads this field to decide whether a quick task is done;
without it the scanner falls back to `[unknown]` and false-flags finished
tasks as open. Writer-side fix; the scanner is correct and unchanged.
Blast-radius: no other scanner reads `status:` from phase-plan SUMMARY
files. Phase disk_status is derived from file-count heuristics only.
Adding the field to the shared template is therefore safe and the value
`complete` is semantically accurate for a finished plan.
Regression test: tests/bug-950-quick-summary-status-complete.test.cjs
- RED: 4 template-contract tests fail before fix, behavioral tests pass
- GREEN: all 8 tests pass after fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add changeset for fix/950-quick-summary-status-complete (#951)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#950): assert writer-path contract + scope template checks to YAML frontmatter (adversarial review)
- Add `// allow-test-rule: source-text-is-the-product` at file top (before block comment)
- Add `extractFrontmatter()` helper that handles both leading-frontmatter files
(summary-minimal/standard/complex.md) and fenced-frontmatter files (summary.md,
whose frontmatter is embedded inside a ```markdown fence) — assertions now
target the actual YAML block, not the whole file
- Scope all four [TEMPLATE CONTRACT] tests through extractFrontmatter() so a stray
`status: complete` in prose/examples cannot produce a false green; error messages
now print the extracted block to aid diagnosis
- Add [WRITER-PATH] quick.md test: asserts the <constraints> block instructs the
executor to write `status: complete` in SUMMARY frontmatter
- Add [WRITER-PATH] gsd-executor.md test: asserts the Frontmatter spec documents
`status: complete` as a required field
- Sanity-checked: guards fail when `status: complete` is removed from a template
or from quick.md, and pass once restored
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Make install + surface read the registry's derived profileMembership/
capabilityClusters so a capability's tier drives what installs + surfaces.
resolveProfile (when given the registry) unions capability skills for the
profiles its tier implies before the requires: closure; resolveSurface merges
capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the
capability-state resolver all thread the registry.
Shipped as a proven no-op: the UI capability is reconciled to tier:full (its
skills were full-only in the hand-authored profiles), so it contributes only to
the full profile (already the '*' sentinel) and core/standard are unchanged.
Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/
capability-state are identical with vs without the registry; the core-alias
staging path is verified equivalent (empty manifest → raw PROFILES.core).
Closes#949
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.
Additive: install/surface/workflows untouched; consumed by nothing.
Closes#945
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Establish `tier` as the source of install-profile + cluster membership
(ADR-894 §4). The registry now derives two views: capabilityClusters
(capId → its skills) and profileMembership (capId → {tier, profiles}, where
profiles is the PROFILE_RANK suffix from the capability's tier). A consistency
gate cross-checks them against the hand-authored PROFILES/CLUSTERS: HARD (throws)
on a capId-matching-a-cluster-name with a different skill set; SOFT
(pending-reconciliation stderr warning, not serialized) when a capability skill
isn't yet in the closure-resolved hand-authored profile.
The SOFT gate loads the real skills manifest and resolves each profile's closure
once so transitively-included skills don't false-warn; warnings are de-duped to
one per (capability, skill); both derived views are scoped to skill-owning
capabilities; serialized with a global capId sort for determinism; reserved-name
guards at every write site; lazy requires of the built constants.
Behavior-preserving: install/surface untouched; derived views consumed by
nothing. Full generation rides along with the phase-6 migration.
Closes#942
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#934): reapply verifier handles missing pristine baseline post-rename
Gap 1 (verify-reapply-patches.cjs): when backup-meta.json records a
pristine_hash for a file but gsd-pristine/ has no corresponding snapshot
on disk, the verifier fell to over-broad mode and produced false
FAIL_USER_LINES_MISSING. Fix: return advisory OK_NO_BASELINE (non-blocking,
exit 0) so the verifier does not block on files it cannot reason about.
Gap 2 (new migration 004): migration 003 removed legacy get-shit-done/
runtime files but left gsd-pristine/get-shit-done/ orphan snapshots in
place. Those stale snapshots referenced get-shit-done/... key paths that
no longer match the active gsd-core/... layout. Fix: add migration
004-prune-stale-pristine-get-shit-done (NOT editing 003, preserving its
checksum — ref #670 guard) to remove all files under
gsd-pristine/get-shit-done/ as GSD-managed pristine snapshots.
Includes tests: bug-934 OK_NO_BASELINE assertions in the verifier test,
new installer-migration-prune-stale-pristine.test.cjs, updated
installer-migrations baseline-lock checksum for 004.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#934): rename migration to satisfy legacy-name guard + mark intentional path refs
Rename src/installer-migrations/004-prune-stale-pristine-get-shit-done.cts
→ 004-prune-stale-pristine-snapshots.cts so the filename no longer contains the
forbidden token. Update .gitignore and eslint.config.mjs to track the new built
path. Add gsd-allow-legacy-name markers to the remaining intentional uses of the
legacy path string in the migration body (lines 3 and 100) and in tests
(installer-migration-prune-stale-pristine.test.cjs lines 202 and 226; and the
baseline-lock key in installer-migrations.test.cjs:1469). Update the baseline
checksum for migration 2026-06-09-prune-stale-pristine-get-shit-done to reflect
the two new marker comments added to its body.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Both sites in plan-review-convergence.md that wrapped gsd-plan-phase in
Agent() (initial planning + replan loop) are now bare Skill() calls at depth 0.
On Claude Code, a depth-1 Agent has no Agent tool so wrapped plan-phase could
never spawn gsd-planner/gsd-plan-checker — the replan loop silently produced no
revised plan when HIGHs were found. Running plan-phase inline from the depth-0
orchestrator (which retains the Agent tool) restores the full sub-agent chain.
A full audit of all workflow files confirmed these two sites were the only
instances of the anti-pattern (no other workflow wraps a spawner orchestrator
in Agent() without a RUNTIME carve-out).
Added structural guard test bug-936-no-nested-spawner-wrap.test.cjs that
dynamically derives the spawner set (workflows containing subagent_type=) and
asserts no workflow wraps a spawner inside Agent() without a RUNTIME != claude
carve-out — prevents silent regression. Test passes on fixed code, would fail
on pre-fix code at the two de-wrapped sites.
Also applied two low-severity prose nits flagged in review:
- commands/gsd/plan-review-convergence.md: orchestrator role updated to
describe inline plan-phase + Agent for review (was generic "spawn Agents")
- gsd-core/workflows/plan-review-convergence.md success_criteria: narrowed
"Each Agent fully completes" to the review Agent (plan-phase is inline now)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- bin/install.js now copies scripts/changeset/ and scripts/lib/ into
<configDir>/scripts/ so $GSD_DIR/scripts/changeset/cli.cjs resolves
at runtime; aborts install with an explicit failure if the source
directory is missing from the package.
- gsd-core/workflows/update.md: corrected path from
gsd-core/scripts/changeset/cli.cjs to scripts/changeset/cli.cjs;
added an explicit [ ! -f ] guard so a missing CLI surfaces a clear
message rather than silently swallowing the error; stderr captured
via 2>&1 sentinel so node errors are visible in the preview output.
- release.yml's changeset-CLI invocations (node scripts/changeset/cli.cjs)
remain at the repo-root path and are unaffected by this change.
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`context: fork` strips the `Agent` tool from a subagent's environment.
Spawning orchestrators (`/gsd-autonomous`, `/gsd-execute-phase`,
`/gsd-plan-phase`) depend on `Agent` to dispatch sub-agents; running
them forked silently disables the core capability they exist to provide
(#921). Remove `context: fork` from all three command frontmatter files.
`effort: xhigh` (introduced by #769) is preserved.
The `<runtime_compatibility>` Agent-availability guard added by #913 was
checking whether `Agent` was present *before* attempting the call. On
runtimes where the tool list is dynamically resolved this produced
false-negative aborts in sessions that have the tool (#922). Replace
the introspection-based pattern with an attempt-based gate: always
attempt the `Agent()` call; stop only if a real tool-unavailable error
is returned. This preserves #853's backgrounded-session close-off and
#913's intent of preventing inline role-collapse, while eliminating
false negatives.
Tests updated: enh-769-context-fork-effort.install.test.cjs asserts the
three orchestrators lack `context: fork` and that the converter still
passes the field through for non-orchestrator commands; plan-phase-drift-
guard.test.cjs adds four assertions for the attempt-based gate language;
workflow-size-budget unchanged (budgets not exceeded).
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add the loop.render-hooks resolver: the first registry-consuming query.
`gsd-tools loop render-hooks <point>` validates the point against the
authoritative canonical 12, reads the registry's materialized byLoopPoint
hooks, filters them by activation, and emits a JSON envelope {point,
activeHooks, rendered} with ordered markdown.
Activation resolves each hook's `when` key by precedence: loadConfig value
(post-cutover federated) -> raw config.json workstream/root single-key lookup
(pre-cutover central override) -> registry configSchema default (so a
default:true capability hook is active out-of-the-box) -> inactive. Guarded
single-value reads only (no merged object built from untrusted keys).
Registry-only: no workflow calls the resolver yet (wiring is the phase-6
cutover). Completes the phase-3 trio.
Closes#918
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.
Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.
Closes#910
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Three-part fix for the top-level inline collapse bug:
1. plan-phase.md: add <runtime_compatibility> block after
</available_agent_types> that makes the Agent-availability
requirement explicit; workflow fails-closed (stops with a clear
log) in genuinely Agent-less contexts.
2. plan-phase.md: rename 7 "ORCHESTRATOR RULE — CODEX RUNTIME"
labels to "ALL RUNTIMES" so the spawn guard applies universally
(not just when Codex is detected).
3. execute-phase.md: scope the existing "Other runtimes" inline-
fallback prose to non-Claude contexts, preserving the #853
backgrounded-agent behaviour for Claude Code background agents.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- Updated `gsd-core/workflows/_runtime-launcher.snippet.sh` with 15 new
`elif` arms covering Hermes, Cursor, Codex, Gemini, Copilot, Windsurf,
Augment, Trae, Qwen, CodeBuddy, Cline, Grok, Antigravity, OpenCode, and
Kilo (respecting each runtime's env-var override with a `$HOME`-relative
default).
- Re-ran `scripts/sync-runtime-launcher.cjs` to propagate the expanded
snippet into all `gsd-core/workflows/*.md` files (~70 files).
- Manually applied the same snippet update to `commands/gsd/import.md`
(1 occurrence) and `commands/gsd/graphify.md` (5 occurrences) — these
are not covered by the sync script.
- Updated `tests/workflow-size-budget.test.cjs` budgets (XL/LARGE/DEFAULT
+ discuss-phase target) to account for the ~3 KB snippet expansion.
- Added regression test `tests/bug-891-non-claude-runtime-home-fallback.test.cjs`
(6 tests: structural probe presence, ordering, behavioral HERMES_HOME
env-var + default-path stubs, resolution order, and workflow propagation).
- Added `.changeset/891-launcher-non-claude-runtime-homes.md` (Fixed).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry
generator with a generated-from-workflows contract (ADR-894 §3). The contract
is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five
step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs,
and required by gen-capability-registry.cjs — one source of truth, no drift.
Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step
must declare exactly its canonical loop points), multiple-block + duplicate-key
hard errors, and a word-boundary agent-role cross-check. Contract content is
byte-identical to the former constant; registry-only, nothing wired into the
live loop.
Closes#903
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl)
First phase-3 code: the Capability Registry generation pipeline, built against
the ADR-894 contract and NOT wired into the live loop (registry-only, per the
staged-cutover design).
- capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2
skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation.
- scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled
schema validation (envelope + role-typed feature/runtime bodies + typed
steps/contributions/gates + when + gate-check variants); cross-capability
invariants (single ownership; requires exist+acyclic+tier-monotone; config-key
ownership exclusive, collision-vs-central as a pending-migration warning);
hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its
source to the generated-from-workflows contract); GLOBAL point-ordered
consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes
topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned
indexes + requiresClosure). Prototype-pollution guards (Object.create(null) +
inline literal key checks) + fragment.path traversal guard.
- gsd-core/bin/lib/capability-registry.cjs — committed generated artifact
(mirrors package-identity.cjs: script-generated, tracked, linted, regenerated
on build, drift-tested), wired via the new `gen:capability-registry` build step.
- tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook +
ordering + adversarial (path-traversal, proto-pollution, runtime body,
self-consume, cycles, collisions) + committed-file staleness guard.
New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md
"Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop.
Gates: lint, code-review (4 bugs fixed), security-review (path-traversal +
prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed,
confirmed sound), clean-build docker 13190 pass / 0 fail.
Closes#896
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows)
The committed capability-registry.cjs staleness guard failed on Windows CI only:
git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the
generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize
line endings on both sides of the --check comparison (no .gitattributes change,
no change to the LF the generator writes). Adds a regression test simulating the
Windows CRLF checkout.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#853): gate manager/autonomous bg dispatch by runtime
/gsd-manager and /gsd-autonomous --interactive dispatched Plan/Execute
via Agent(run_in_background=true). On Claude Code a backgrounded agent
has no Agent/Task tool, so it cannot spawn the nested subagents those
pipelines need — per-plan worktree-isolated executors, the plan-checker,
and the verifier. The phases reported complete but isolation and
independent verification silently never ran, even with use_worktrees /
plan_check / verifier enabled.
Both workflows now resolve the runtime (config-get runtime, default
claude) before dispatching: run plan/execute INLINE on Claude Code so
the nested pipeline runs, and background-dispatch only on runtimes where
a backgrounded agent can still nest. Mirrors execute-phase.md's existing
Codex fail-closed precedent. Reconciles the stale unconditional
background/overlap/lean-context claims elsewhere in both workflows and
in the docs. Adds a content regression test pinning the gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#853): add changeset for runtime-gated bg dispatch
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior.
Closes#815
* feat(#773): add --ephemeral and --dangerously-bypass-hook-trust to automated codex exec invocations
Automated codex exec calls in the review workflow now carry --ephemeral
(no session-state accumulation across CI runs) and
--dangerously-bypass-hook-trust (skip hook-trust prompts for hooks
whose provenance gsd-core already controls). Both flags were verified
present in the installed codex CLI (codex exec --help).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#773): correct changeset pr: reference to #824
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#651): consolidate verification-status routing into one queryable seam
The passed/gaps_found/human_needed verification status was re-encoded as
bare strings across three prose surfaces (gsd-verifier emits, execute-phase
routes, ship gates), each independently deciding the per-status next action
with no parity coupling — the DEFECT.GENERATIVE-FIX class.
Give the enum one home: src/verification.cts (-> bin/lib/verification.cjs)
exposing `gsd_run query verification.status <phaseDir>` returning a typed
{status, next_action, next_command}. ship.md and execute-phase.md now consume
the query instead of re-deriving the routing in prose; gsd-verifier.md points
at the shared vocabulary as the single emitter (values unchanged).
Also fixes the latent broad-grep status misread (DEFECT.FRONTMATTER-SCALAR-
BROAD-GREP): execute-phase.md read `grep "^status:"` over the whole report, so
a body `status:` line could misroute a valid phase. Extraction is now
frontmatter-scoped in one place. A parity test fails if a verifier status
gains no route. Lands the two CONTEXT.md DEFECT entries captured on the issue.
Closes#651
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#651): set changeset pr to 755
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#52): add agent_skills_security.trusted_global_roots allowlist
Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.
- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
project-relative and dangerously broad roots (filesystem/UNC root, homedir),
realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
(hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
feature, negative, broad-root hardening, stderr NOTE).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#52): add changeset fragment for trusted_global_roots (#754)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#703): add --granularity override flag to /gsd:plan-phase
Add a `--granularity <coarse|standard|fine>` flag to /gsd:plan-phase that
overrides the configured planning granularity for a single invocation.
The override is a new highest-priority tier above the existing precedence
chain (granularities[phaseType] -> granularity -> planning.granularity ->
'standard') in resolveGranularityInternal; when the flag is absent, resolution
is byte-for-byte unchanged. cmdInitPlanPhase now resolves with phaseType
'planning' so granularities.planning participates, and emits the resolved
value in the init JSON, which the plan-phase workflow forwards to the planner
prompt. Invalid values are rejected at the CLI boundary via a shared
assertValidGranularityOverride helper.
Closes#703
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#703): set changeset pr to 750
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch
Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.
- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
(origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
auto-degrades the run to sequential on the main tree when a base mismatch
is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
.claude/settings.local.json (no-clobber, respecting an explicit shared
settings.json value); upgrades print an opt-in notice pointing at
`gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees
Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure
The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.
The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)
tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.
All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#720): lazy-load MVP-only reference bodies (eager @-import → gated Read)
Convert eager @-imports of MVP-only reference bodies into lazy "Read" instructions
gated on MVP_MODE / WALKING_SKELETON / MVP+TDD, so non-MVP planning/execution runs
no longer pull MVP guidance into context. Covers both the workflow files and the
planner/executor agent definitions (the dominant context-cost path):
- workflows/plan-phase.md: planner-mvp-mode.md + skeleton-template.md (L146/936/937/941)
- workflows/execute-phase.md: execute-mvp-tdd.md halt-report ref, now gated on gate-trip (L191)
- agents/gsd-planner.md: planner-mvp-mode.md, user-story-template.md, skeleton-template.md
- agents/gsd-executor.md: execute-mvp-tdd.md
The dedicated always-MVP mvp-phase workflow keeps its eager imports (intentional).
Behaviour is unchanged; non-MVP runs simply carry less loaded context. Adds a
regression guard mirroring the discuss-phase lazy-load test, and documents the
conformance in docs/ARCHITECTURE.md.
Refs #720
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#720): add changeset fragment (pr #746)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Part 2 of 2 of the n/no-process-exit cleanup (completes umbrella #738; part 1
was #739/scripts). Converts the 20 flagged process.exit() calls in the three
hand-written gsd-core/bin CLI entrypoints and flips n/no-process-exit to error.
- New src/cli-exit.cts -> gsd-core/bin/lib/cli-exit.cjs (ExitError + runMain),
the gsd-core-side equivalent of scripts/lib/cli-exit.cjs; registered in
.gitignore, eslint ignores, and the inventory manifest like its siblings.
- gsd-tools.cjs: 13 apply-prompt-budget exits -> throw ExitError; main()->runMain.
- verify-reapply-patches.cjs: 6 exits -> throw ExitError / return verdict; runMain.
- check-latest-version.cjs: 1 exit -> return verdict; runMain.
- eslint.config.mjs: n/no-process-exit warn -> error.
Scope note: the gsd-core/bin/lib/*.cjs modules (core, state, profile-pipeline,
roadmap-command-router, adr-parser, ui-safety-gate) are tsc-generated and
eslint-ignored (ADR-457), so their process.exit calls were never flagged and are
intentionally left untouched. Only the linted hand-written entrypoints are in scope.
Exit codes verified unchanged for all three entrypoints.
Closes#738
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).
No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.
Closes#732
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase
When RESEARCH.md already exists in research-only mode and neither --research
nor --view is passed, emit a one-line notice and exit cleanly instead of
prompting update/view/skip. This matches the promptless auto-use of standard
/gsd:plan-phase <N> (§5.1) and removes the §5.0/§5.1 inconsistency, making
AI-agent and CLI invocations non-interactive in the common case. The two
explicit-flag escape hatches (--research to refresh, --view to print) cover
any deviation.
Closes#159
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#159): point changeset fragment at PR #718
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#159): tighten research-phase reference register (Diataxis)
Make the 'no modifier' research-phase entries descriptive rather than
imperative and drop the trailing 'pass --research/--view' clauses, which
duplicated the adjacent --research/--view documentation. Reference docs
describe; the recovery flags are documented in their own entries. The
emitted runtime notice in the workflow keeps naming the flags (in-band
recovery), unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#656): add Research Store module (content-addressed cache, TTL staleness)
Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1.
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#656): add Research Provider module (waterfall + confidence + plan)
Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests.
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional)
Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1).
Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter
config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged.
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy)
Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335).
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset)
Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping).
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#656): sync inventory for research modules
Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT).
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457)
research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage.
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#656): backfill changeset pr number to #664
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#656): satisfy eslint lint-tests gate
Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green.
Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): harden package legitimacy per review (W1/W2/I3/I4)
W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green.
Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4)
I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green.
Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3)
Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests.
Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache)
HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green.
Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): close code-review correctness findings
(1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green.
Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(#657): extract researcher documentation_lookup to shared @-reference
6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse.
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(#657): extract researcher philosophy + verification-protocol to shared @-references
philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A.
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1)
The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation).
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1)
project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer.
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2)
Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green.
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3)
scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change).
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles
Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first).
Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#656): make classifyConfidence verification-evidence-driven (W3)
Confidence conflated provider authority with claim verification — context7/ref
stamped HIGH purely by provider identity, and the only verification lever was a
self-set --verified flag. Split into two axes: provider authority (static) +
verification evidence (code-computed). HIGH now requires ground-truth
corroboration (legitimacyVerdict OK), independent of provider; authority alone
caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a
MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a
correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI;
updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent).
Addresses davesienkowski's W3 review on #664.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#656): bind classify-confidence verdict to code, closing CLI self-grading
Adversarial review found the new --legitimacy-verdict flag was caller-supplied,
so an agent could self-assert OK->HIGH without any real legitimacy check —
reintroducing the exact self-grading hole W3 closes. Remove the free flag; the
CLI now computes the verdict via checkPackages only when --package/--ecosystem
is given (code-computed, not agent-asserted). Update the stale CLI test
(context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#669): /gsd-review --cursor actually invokes cursor-agent
The Cursor reviewer branch in review.md never ran the agent:
- detection probed `cursor` (the IDE launcher) instead of the headless
`cursor-agent` binary
- the invocation used the two-token `cursor agent` (the IDE treats `agent`
as a file-path argument, so the agent never starts)
- the prompt was piped via stdin, but `cursor-agent -p` reads the prompt
from a command-line argument, and `2>/dev/null` hid the empty result
Probe `cursor-agent`; invoke `cursor-agent -p --mode ask --trust
--output-format text` with the prompt passed as a file-path-reference
argument (avoids the OS arg-length limit on large prompts); capture stderr
so failures are diagnosable. Invert tests/cursor-reviewer.test.cjs to assert
the corrected contract, with negative guards against the two-token form and
the stdin pipe. The sibling `agy` reviewer already used the argument form.
Closes#669
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#669): set changeset pr number to 686
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`/gsd-review --agy` hung indefinitely on large prompts. agy's print mode runs the
full tool-enabled agent, and on a big, file-path-rich prompt its agentic Cascade
loops on the code_search/grep tool and never converges; the transcript fallback
only runs after agy exits, so it can't recover a run that never exits.
The agy CLI exposes no per-tool deny (that lives in the Antigravity SDK), but it
does expose --print-timeout — agy's native print-mode cap. Pass it explicitly so a
stalled run self-terminates through the tool's own mechanism; a non-zero exit
discards any partial output so the existing transcript fallback / "review failed"
stub take over. Adds a regression test.
Closes#687
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run
The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" <cmd>` form
(fixed for workflows in #621/#637) survived in agent/command surfaces and
misresolves on global/shim-only installs. Route every agent-executed
invocation through the resolved `gsd_run` launcher in gsd-phase-researcher,
gsd-planner (load_graph_context extracted to a shared reference to stay under
the planner size budget), import, and graphify. Add a regression guard over
agents/ + commands/ + gsd-core/references/ bash blocks. User-facing display
messages and docs are intentionally left untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#705): use repo changeset fragment format (type: Fixed, pr: 707)
The hand-written fragment used the standard changesets package format
(package: bump) which lacks the type:/pr: frontmatter the repo's
docs-required lint consumes (fail_malformed_fragment / missing_type).
Regenerated via scripts/changeset/new.cjs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#637): route 3 more workflows through gsd_run launcher (hardcoded $HOME sweep)
The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs"` invocation form
fixed in plan-phase.md (#621) survived in three more workflows. Same bug class:
on a global/shim-only install with no project-local runtime, the hardcoded path
can miss a working install, so the step reports the tool "not found" instead of
resolving it via the launcher. #3668 introduced gsd_run resolution; these sites
were missed.
- plan-review-convergence.md: convert the 3 hardcoded invocations (init,
roadmap get-phase, state planned-phase) to gsd_run. File already carried the
canonical preamble (first gsd_run is the earlier convergence-enabled check).
- ingest-docs.md, spec-phase.md: convert their hardcoded invocations to gsd_run
and inject the canonical launcher preamble via
`node scripts/sync-runtime-launcher.cjs` (these files previously had no
gsd_run and no preamble). The injected preamble is byte-equal to
_runtime-launcher.snippet.sh and precedes the first gsd_run call, per
runtime-launcher-parity invariant (B).
- Add tests/bug-637-workflow-no-hardcoded-home-tool.test.cjs: repo-wide
regression guard asserting NO workflow .md invokes gsd-tools via a hardcoded
$HOME path. Generalizes the plan-phase-only guard from #621 — the parity test
guards retired $GSD_SDK / bare /gsd-tools tokens but not this form, which is
how it survived across four files. Fails on the pre-fix files, passes after.
runtime-launcher-parity 7/7; full unit suite green (3477 pass / 0 fail).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#637): add changeset fragment for PR #642
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#637): update stale bug-2801 assertion to expect gsd_run
bug-2801 pinned ingest-docs.md to the hardcoded node "$HOME/.../gsd-tools.cjs" init form, which #637 replaces with the gsd_run launcher. Flip the assertion to expect gsd_run init ingest-docs; the bare-gsd-tools rejection and CLI-handler tests are unchanged, and bug-637's repo-wide guard now owns the no-hardcoded-$HOME invariant.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#690): promote CHANGELOG 1.3.x + gate release-notes promotion
/gsd:update showed an empty "What's New" preview after updating to 1.3.1
because CHANGELOG.md's 1.3.x content was never promoted out of [Unreleased]
into dated sections, so `scripts/changeset/cli.cjs extract` returned exit 2
("no releases in range").
- CHANGELOG.md: split [Unreleased] into dated [1.3.0] and [1.3.1] sections
(1.3.1 = hono advisory bump + installer-migration checksum self-heal, #670;
1.3.0 = the feature release), restoring an empty [Unreleased].
- scripts/changeset/cli.cjs: new `verify` subcommand that exits non-zero when
CHANGELOG has no dated `## [x.y.z]` heading for a version; hoist shared
stripV/resolveChangelogPath helpers used by extract + verify.
- .github/workflows/release.yml: gate the finalize job on `verify` (after the
build, before tag/publish) so an unpromoted CHANGELOG can never ship again.
- gsd-core/workflows/update.md: move `rm -f $CHANGELOG_TMP` after the
human-readable extract re-run so the preview no longer degrades to
"(changelog unavailable)".
- tests: regression guard for the 1.3.x headings + extract range + verify
command coverage (present/absent/undated/v-prefixed/--json/prerelease).
Closes#690
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#690): add changeset fragment for #694
Fixed-type fragment for the user-facing /gsd:update preview fix and the
release-notes promotion gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#586): make ship PHASE_VERIFICATION_INCOMPLETE actionable, drop dead `pass` arm
The ship preflight gate blocked with PHASE_VERIFICATION_INCOMPLETE but named no
next step, and accepted a `pass` status the verifier never emits. Capture the
verification status and route per value (gaps_found / human_needed / missing),
mirroring execute-phase's status table; accept only `passed`.
Closes#586
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#586): backfill changeset PR number 650
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#586): scope ship verification status to frontmatter only
Codex adversarial review of PR #650 flagged that the status gate grepped
`^status:` over the entire VERIFICATION.md, so a `status:` line in the report
body (a code block / copied artifact) concatenates into a non-matching value and
blocks a genuinely-passed phase with the wrong next action. Restrict extraction
to the leading YAML frontmatter block, first match only. Adds a behavioral
regression test that runs the gate's own bash pipeline against a passing report
whose body contains decoy `status:` lines.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#586): drop manual PR ref from changeset body
The changelog renderer auto-appends `(#<pr>)` from the fragment's pr: field
(scripts/changeset/serialize.cjs, github-release-notes.cjs). The manual trailing
`(#586)` produced a double, mismatched ref (issue #586 + auto PR #650); remove it
to match the sibling-fragment convention.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#586): make ship-586 bash-fence regex Windows-safe (CRLF)
The behavioral test extracted the gate's bash block with /```bash\n.../ — a
literal \n that fails to match Windows CRLF checkouts and trips the
windows-test-parity-guard (fenceRegexLiteralNewline). Use ```bash\r?\n and
normalize the captured block to LF before running it. Full unit suite: 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#586): run ship-586 bash-pipeline tests on POSIX only
On Windows CI the behavioral tests failed: git-bash is present (so the old
hasBash guard ran them) but receives a Windows-style tmpdir path it cannot glob,
so extraction returned empty. The extraction logic is platform-independent and
the gate's bash only runs in a POSIX workflow context, so skip the pipeline
execution on win32. POSIX (macOS/Linux) still runs and asserts it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>