verify-phase.md grew 30875 -> 32334 from the test-tier enforcement consumer step
+ descriptor docs. Well under the 40960 DEFAULT hard cap; regenerate the per-file
baseline ratchet (the growth is the load-bearing consumer wiring, not bloat).
The test-tier fail-closed reason string (surfaced to users via the producer) and
the dispositionForProhibition docstring still said the negative-test enforcement
was 'deferred to a follow-up PR'. This PR IS that follow-up, so the claim is now
false. Updated the user-facing reason and the comment to say the producer landed
in #1259. Policy/branching logic is byte-unchanged (only message strings).
Round-2 adversarial review found the lint-rule runner falsely greened an
eslint-IGNORED target: eslint returns a length-1 "File ignored" result (ruleId
null) that passed the >=1-file vacuity guard while nothing was linted — reopening
the vacuous-green class. Fix: buildLintArgs now passes --no-warn-ignored so an
ignored path returns [] -> fails closed (verified + E2E test on an ignored
bin/lib artifact).
Also: wrap runCheck() so even a (test-injected) throwing runner fails closed
(NEW-WR-01, full no-throw contract); document the benign basename-naming
constraint on wired node-test names (NEW-WR-02, fail-closed).
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:
- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
reported test named distinctly from the file — node --test counts an empty file
as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
eslint as --format json and filters by ruleId, so plugin rules (local/*) load
via the flat config — bare --rule cannot load a plugin. local/no-source-grep
now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
producer requires attestation + a genuine non-vacuous pass. Machine-proven
fail-first (needs a violation fixture) is flagged as a tracked follow-up in
ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
- Extend prohibition-probe.verify-tier.test.cjs with the ENFORCEMENT half (#1259, ADR-550 D5d)
- Require the not-yet-built gsd-core/bin/lib/prohibition-enforcement.cjs (RED)
- Cover both wired-check kinds (node-test + no-source-grep lint-rule) and miss/fail hard-gate
- Typed-field assertions only; injected runCheck (no real subprocess)
- Keep the 2 original fail-closed tests verbatim
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents
- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(#1243): document plugin-provided skills in agent_skills
Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.
Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)
- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1243): regenerate agent-size baseline for the Skill-tool grant
The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).
* chore(#1243): add Added changeset fragment
* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase
cmdStatePlannedPhase ran its body-field replacements on the full file content,
so the case-insensitive ^Status: pattern matched the YAML frontmatter `status:`
line before the body `| Status | … |` cell — the cell never advanced to
'Ready to execute' and syncStateFrontmatter re-derived the stale 'planning'
status (the #1230 delta heuristic preserved it). cmdStateBeginPhase had
pipe-table else-branches only for Status/Last activity (#1256), so the Current
Position `| Phase |` / `| Plan |` cells were left stale while a spurious inline
`Phase: N — EXECUTING` line was prepended.
Both handlers now strip frontmatter before body-field replacement and update the
pipe-table cells in place via stateReplaceField, matching inline-format
behaviour. Systemic residual of #1255 / #1256.
Adds 4 regression tests (#1257 block in tests/state.test.cjs) covering both
findings — RED before the fix, GREEN after; full state-area suite (437) and
local unit suite stay green.
Closes#1257
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1257): backfill changeset pr number (#1260)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase
state begin-phase/complete-phase called stateReplaceField on the FULL file content, so its case-insensitive ^Status: plain-pattern matched the YAML frontmatter status: line first (no g flag) and never updated the body pipe-table Status cell; syncStateFrontmatter then re-derived the stale status from the unchanged body, freezing the frontmatter status. Fix: strip frontmatter before the body-field replacements (operate on body only), reassemble with frontmatter preserved, so the body Status cell updates and the frontmatter derives correctly — for inline AND pipe-table body formats. Also corrects the Current Position pipe-table else-branches (Status/Phase/Last-activity) to write bare, consistent cell values. Pipe-table Status is a supported body format (not rewritten to inline).
Closes#1255
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1255): add changeset for pipe-table state status fix (#1256)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1255): fold pipe-table regression into state.test.cjs + Windows-portable frontmatter regex
Per the 2026-06 audit, new tests/bug-NNNN-*.test.cjs files are banned (lint-regression-test-names) — folded the 7 #1255 regressions into tests/state.test.cjs and removed the standalone file + its lint-test-file-count allowlist entry. Also fixed the frontmatter assertions' /^---\n/ anchors to /^---\r?\n/ (windows-test-parity-guard frontmatterAnchorLiteralNewline).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).
Closes#1110
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes#644.
Records the design gate for moving agent conversion off the inline bin/install.js loop onto the descriptor path (ADR-3660): the two-(really three-)path problem, the 10 parity behaviors the descriptor agents path must gain (verified against code — the issue's 7 plus Qwen/Hermes branding, the Codex TOML sidecar, and stale-agent cleanup), an AgentConverterContext contract to fix the leaky (content)=>string converter signature, and an incremental per-runtime cutover gated on byte-for-byte golden parity (full + minimal mode). Status: Proposed. Codex-reviewed for technical accuracy.
Closes#1235
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings
- worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs).
- active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting.
- gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle.
- install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests.
Closes#1191
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1191): add changeset for valid-null settings fix (#1233)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy
The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam
ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset
Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1188): promote no-source-grep to error on tests/**
Apply local/no-source-grep as an error to tests/**/*.test.cjs (was warn on bin/scripts only). Only 2 real violations surfaced (repo-layout.test.cjs structural guard-placement checks on bin/install.js) — marked allow-test-rule with #1188 reason. PART A (branch-coverage floor) follows separately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1188): add c8 branch-coverage floors (60 global, 70 per-file on UNMUTATED modules)
c8 enforced line floors only. Add a --branches 60 global floor (current ~82.5%, margin mirrors the lines-70-vs-91% gap) to test:coverage + test:coverage:unit, plus a chained 'c8 check-coverage --per-file --branches 70' for the high-risk UNMUTATED modules state/phase/verify/init (current min 78.3% on verify). The per-file check reuses the coverage data the suite run just produced (the proven scripts-floor pattern) so it adds no second suite run.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1188): per-module branch check must pin lines/funcs/stmts to 0
Standalone 'c8 check-coverage' defaults unspecified metrics to 90 and enforces them, so --branches 70 alone also failed on lines (verify.cjs 85.29% < default 90). Pin --lines 0 --functions 0 --statements 0 so only the branch floor (70) is enforced on the UNMUTATED modules. Data-reuse confirmed working (it computed verify's real %).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source
Adds a delta heuristic in `readModifyWriteStateMd`: snapshots the body
Status and Stopped At fields before the transform, then after
syncStateFrontmatter runs, restores the existing frontmatter values for
any field whose body source was not changed by this write. Integrates
cleanly with the existing !resync progress-restore block (computes postFm
once, applies both restorations, reconstructs frontmatter once).
Regression tests added to bug-905 test file (same theme, no new file).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1230): add Fixed changeset fragment
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1229): count bullet-only phases + guard against number collision in phase.add
Before this fix, the phase.add number scan only checked ### Phase N: section
headers and on-disk phases/N-* directories. A phase that existed only as a
roadmap bullet (e.g. "- [ ] **Phase 11: ...**") was invisible to both scans,
causing phase.add to silently assign a duplicate number.
Fix: add a bullet-entry regex scan (all checkbox variants: [ ], [x], [~], with
or without ** bold markers) to the set-based phase-number collection in
cmdPhaseAdd. Also added a post-compute collision guard that advances the
candidate past any already-used number.
Regression tests added to tests/phase.test.cjs (bug #1229 describe block):
bullet-only collision, [x]/[~] variants, plain-bullet, and baseline preservation.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1229): add Fixed changeset fragment
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(#1239): ADR for GSD as an embeddable orchestration engine
Records the design to invert GSD from a standalone installer that projects
onto a host into an embeddable orchestration engine a host loads as a plugin,
driven through a negotiated host-integration interface. Unifies ADR-1016
projection (the declarative adapter) with imperative embedding behind one
contract; grounded in a 9-host capability survey.
Closes#1239
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1239): de-slash illustrative command placeholders (docs-parity gate)
Replace illustrative /gsd:x /gsd-x /gsd.x placeholders with namespace-prefix
wording so the docs-parity live-registry check does not parse them as
non-live commands.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ADR-230's branching-model gate (pr-target-validator.yml) decided allowed/blocked PR targets via inline regex in github-script — untestable. Extracted the decision into committed scripts/pr-target-policy.cjs (classifyPrTarget(base,head)->{decision}), and rewired the workflow to checkout the BASE ref (trusted; fork-tamper-safe) + require the module. Behavior-identical (Codex-verified char-by-char regex equivalence + all side-effects preserved). 70 tests incl. an equivalence oracle battery + hyphen-boundary negatives. Added contents:read for the checkout. Re-attribution: no ADR-230 test references exist (issue's '2 misattributed files' claim not borne out).
Closes#1190
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads
Before this fix, bin/install.js copied scripts/changeset/ and scripts/lib/
into the runtime config dir but omitted scripts/fix-slash-commands.cjs.
gsd-core/bin/lib/command-roster.cjs requires this file at module load via
require('../../../scripts/fix-slash-commands.cjs'), so every gsd-tools
command crashed with MODULE_NOT_FOUND on every installed runtime.
Four changes:
- bin/install.js copy step: copy fix-slash-commands.cjs into <configDir>/scripts/
with source-missing hard-fail and verifyFileInstalled smoke check
- bin/install.js writeManifest: track scripts/fix-slash-commands.cjs (not
covered by the changeset/lib subdir loops)
- bin/install.js uninstall: best-effort unlinkSync before scripts/ rmdir
- scripts/fix-slash-commands.cjs readCmdNames(): wrap readdirSync in
try/catch returning [] so skill-based/global installs without a local
commands/gsd/ directory do not throw ENOENT
Tests added to tests/install.test.cjs (6 new tests):
- smoke: install() copies fix-slash-commands.cjs
- e2e: spawned gsd-tools.cjs does not crash with MODULE_NOT_FOUND
- manifest: writeManifest() tracks the file
- uninstall: uninstall() removes the file
- readCmdNames unit: export returns an array
- readCmdNames spawn: absent COMMANDS_DIR returns exit 0 (no throw)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1223): backfill changeset PR number (#1240)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)
ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for progress --converge surface (#1237)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline
CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1217): bound acquireStateLock retries on recoverable errno (no busy-spin)
The recoverable-errno branch (`ACQUIRE_LOCK_RETRY_ERRNOS`) previously called
`continue` directly, skipping both `clock.sleep()` and the 30 000 ms budget
check. A permanently-failing ENOENT (e.g. parent dir removed) would spin at
100% CPU forever with the event loop fully blocked, making the OS-level
`timeout` the only escape.
Fix: extract a `checkBudgetAndSleep(context)` helper and call it from BOTH
the recoverable-errno path and the EEXIST contention path so every retry is
bounded and backed off identically.
Regression tests added to tests/clock-seam.test.cjs:
- persistent ENOENT throws budget-exceeded error (clock must advance via sleep)
- transient ENOENT (2 retries then success) acquires lock normally
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1217): backfill changeset PR number (#1236)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1224): accept --pr 0 placeholder at changeset creation
The required-field guard `!opts.pr` treated the integer 0 as falsy,
rejecting the documented `pr: 0` two-push placeholder with a usage
error (exit 2). Non-numeric `--pr abc` (NaN) was also silently
accepted before (passes `!NaN === true`... actually `!NaN` is true, so
NaN would trigger the guard already). The new explicit checks use
`opts.pr === null` for missing flag and `Number.isNaN` for non-numeric
input, accepting all finite integer values including 0.
The merge-time safety net in parse.cjs (`pr <= 0` → INVALID_PR) is
unchanged — a pr:0 fragment is still rejected at lint/render time.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1224): backfill changeset PR number (#1231)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Removes three unreferenced functions from bin/install.js:
- convertCursorToolName (zero callers across src/, bin/, scripts/, tests/)
- convertWindsurfToolName (zero callers across src/, bin/, scripts/, tests/)
- convertAugmentToolName (zero callers across src/, bin/, scripts/, tests/)
Dead code only — behavior-preserving. All three were superseded when
the conversion logic moved to src/runtime-artifact-conversion.cts.
lint and build both clean.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a black-box regression guard (8 tests in tests/init.test.cjs) asserting the init handlers (execute-phase, plan-phase, phase-op, milestone-op) resolve workstream-scoped planning paths under GSD_WORKSTREAM via planningPaths()/planningDir(), never the flat .planning form.
Each positive case asserts both the scoped value and not-equal-to-flat (genuine guard, not coverage credit); fixtures seeded workstream-scoped; every CLI call pins GSD_WORKSTREAM and GSD_PROJECT for hermeticity. Verified by mutation (flat join fails exactly the 4 positive assertions) and re-verified under a polluted parent env. Test-only; src/ unmodified.
Closes#1189
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route
agents-kind entries through a converter when the descriptor carries a
non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter`
to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9
agent converter names, and adds a fail-first behavioral test suite
(9 tests) proving the new wiring end-to-end.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#1213): Capability State Writer — write-side inverse of the resolver
Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1213): add changeset for Capability State Writer (#1225)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.
Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.
- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
round-robin partition (fileIndex % n === i-1) over the SORTED selected file
list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
pure no-op. The 28K Windows argv chunking is preserved within each shard. A
legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
sharding still hits the discovery hard error. Composes with --suite and is
order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
cross-product (explicit include rows — a base shard dim does not cross-product
with include legs, and a nested matrix.leg.os is unresolvable by the H1
shell-policy linter). Unit suite runs sharded; integration/security run once
per leg (shard 1). The Required tests fan-in is unchanged: it already needs
test-full and checks the matrix-aggregate result, so a failed/cancelled shard
fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
balance, determinism, boundaries, fast-check property) + parseShardArg
validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.
Closes#1212
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes#1165.
Extracts all 9 convertClaudeAgentTo* functions together with their full
dependency closure (claudeToCopilotTools, claudeToGeminiTools,
convertCopilotToolName, convertGeminiToolName) from bin/install.js into
src/runtime-artifact-conversion.cts and adds them to the module's export=
block.
Adds 6 regression tests in tests/copilot-install.test.cjs including a
DEFECT.GENERATIVE-FIX parity guard that asserts claudeToCopilotTools is
identical in module and bin/install.js.
Inline copies in bin/install.js are retained (#1175 will remove them).
This unblocks #1173 (descriptor-driven dispatch).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1203): gate free-form ROADMAP deprecation warning on phase_id_convention
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: backfill changeset PR number for #1218
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1203): use canonical gsd-tools roadmap upgrade command in warning
Match the migration command string to the canonical form used in
verify.cts and roadmap-command-router.cts (gsd-tools roadmap upgrade
--convention milestone-prefixed, dry-run by default) instead of the
non-canonical 'gsd roadmap upgrade --apply'.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1202): make verify key-links wave-aware for planned future files
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: backfill changeset PR number (#1219)
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs
- Add Phase ID Convention section to <phase_identification> block:
documents sequential (default) vs milestone-prefixed forms, and
instructs the agent to read phase_id_convention from config.json
- Update <output_formats> to show both header and checklist forms for
sequential and milestone-prefixed conventions with examples
(e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**)
- Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs
(5 assertions, confirmed fail-first then pass after fix)
- Update tests/agent-size-baseline.json to reflect legitimate growth
- Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: backfill changeset pr: 1215 for fix/1205
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs
lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files;
regression cases must live in the owning module's test file.
Move the 5 phase_id_convention assertions (#1205 regression) from the
removed tests/bug-1205-roadmapper-convention.test.cjs into
tests/roadmapper-granularity.test.cjs as a new describe block, alongside
the existing granularity calibration tests. Also update the allow-test-rule
comment to cover both #163 and #1205 surface contracts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity
- Add issue ref (see #1205) to allow-test-rule comment in
tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs
passes (new exemptions require #NNN per ADR-456)
- Prune stale 'source-text-is-the-product' entry from
scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down;
comment now compliant and no longer needs grandfathering)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1194): correct inverted statusline auto-compact buffer math
The reserved-buffer percentage was computed as (acw/totalCtx)*100 — the
usable fraction — instead of (1 - acw/totalCtx)*100 — the reserved fraction.
When acw == totalCtx this produced buffer=100%, making the usable-range
denominator zero and pinning `used` at a constant 100% regardless of real
remaining context.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: backfill changeset PR number (#1211)
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace `^[0-9]+\.[0-9]+\.[1-9][0-9]*$` with
`^(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.[1-9][0-9]*$` in the hotfix
branch of the validate-version step, consistent with the two sibling
patterns already using the strict (0|[1-9][0-9]*) guard. Adds
regression assertions to tests/adr-218-release-version-validation.test.cjs
that prove `01.2.3` and `1.02.3` are rejected.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook
Claude Code marketplace plugin installs unpack the package into the
version-pinned plugin cache and never run bin/install.js, so
~/.claude/gsd-core/ is never created. Agents, commands, and templates
markdown-@-include the canonical ~/.claude/gsd-core/... path (which
expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to
nothing and agents (e.g. the executor) failed.
Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a
plugin install, symlinks the canonical path's immutable subdirs (bin,
contexts, references, templates, workflows) to the plugin's bundled
gsd-core/ tree. It changes zero @-references, is a no-op in classic
installs, preserves user-generated files (USER-PROFILE.md, STATE.md),
prunes stale links so it self-heals after `claude plugin update`, uses
Windows junctions, and rejects bundled/canonical paths that escape the
resolved plugin root (no traversal, no clobber).
Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json
SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES.
Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#997): backfill changeset PR number to #1207
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The triage-label mapping pointed ready-for-agent at the stale `confirmed`
label. The live verified-bug gate is `confirmed-bug` (RULESET.CONTRIB.CLASSIFY.fix
requires confirmed/confirmed-bug; 200+ issues and bug-remediation workflows use
confirmed-bug). Update the row + note so /triage applies the correct gate, and
mark `confirmed` as legacy/back-compat only.
Closes#1209
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#1187): per-module mutation-score ratchet + graduate core-utils
ADR-456's 80% mutation floor was unenforceable as a single global break=50:
4 of 6 covered modules sit at 63-79% and forcing them to 80 would require
brittle exact-string assertions on equivalent string-literal mutants (a
Goodhart's-Law trap). Instead, each covered module declares a minScore floor
(locked at its measured score, TARGET 80) enforced per CI shard via
stryker --break, ratcheting up over time without brittle tests.
- mutation-matrix.cjs: minScore per module + TARGET_MUTATION_SCORE=80, emitted
in the matrix; require.main guard + exports for testability.
- mutation.yml: per-shard --break <minScore>.
- stryker.config.mjs: global break 50->60 as a local backstop (CI uses minScore).
- Graduated core-utils (measured 77.5%, floor 75).
- context-utilization 79.5->92.3% via behavioral killers (state classification
outputs + error-value contract, not exact-string matches) -> minScore 80 (TARGET).
- ratchet-integrity guard test (28 cases).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1187): pass mutation break via MUTATION_BREAK env (no stryker --break flag)
Adversarial review caught that Stryker 9.x has no --break CLI flag, so the
per-shard 'stryker run --break <minScore>' errored out every mutation shard.
Read the per-module floor from process.env.MUTATION_BREAK in stryker.config.mjs
and set it per shard via env in mutation.yml. Red-green verified: MUTATION_BREAK=99
exits 1, =80 exits 0. Also make the ratchet guard monotonic (RATCHET_BASELINE
floors; lowering a floor now fails the guard unless the baseline is edited).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1187): fail closed on bad MUTATION_BREAK + monotonic ratchet baseline
Code review: Number(env)||60 failed OPEN — an empty/invalid MUTATION_BREAK
(e.g. a future module missing minScore -> matrix expands to '') silently
degraded the shard to break 60, letting a high-floor module regress undetected.
resolveMutationBreak() now returns 60 only when the env is truly unset (local
backstop) and THROWS on present-but-empty/non-numeric/out-of-range (fail closed);
stryker.config.mjs imports it via createRequire. Also make RATCHET_BASELINE an
equality mirror (=== not >=) so any floor change is explicit in review and no
floor can be silently lowered. Tests: 46 (incl resolveMutationBreak cases).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1187): recalibrate config-schema/prompt-budget floors to CI scores
First CI mutation run failed two shards: the floors were set from local Stryker
runs whose TIMEOUTS were counted as kills (env-variable), inflating scores. CI
runs with timeout~0, so the real deterministic scores are lower:
- config-schema: local 69.7% -> CI 54.55% (5 local timeouts vanished) -> floor 52
- prompt-budget: local 99.6% -> CI 68.33% (239 local timeouts vanished) -> floor 66
Calibrate floors from CI (the documented source of truth) and record the lesson
in the comment so future floors aren't set from timeout-inflated local runs.
Baseline updated to match. The other 5 shards passed (deterministic CI scores
above their floors).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1160): resolve capability surface from installed skill layouts
In a global skills-runtime install (e.g. Codex at ~/.codex), gsd-tools.cjs
runs from <configDir>/gsd-core/bin/ and the commands/gsd source tree is
absent — only <configDir>/skills/gsd-<stem>/SKILL.md files exist.
_resolveCommandsGsdDir() returned a path that does not exist there, so
loadSkillsManifest returned an empty Map. resolveSurface then materialised
the '*' (full) profile sentinel by enumerating that empty manifest → empty
surfaced Set → every capability reported surfaced=false/enabled=false
regardless of project config. As a result `loop render-hooks verify:post`
returned activeHooks:[] even with workflow.security_enforcement and
workflow.nyquist_validation enabled, silently disabling the security and
Nyquist gates.
Fix: add _loadInstalledSkillsManifest(configDir) that scans configDir/skills/
for gsd-<stem>/SKILL.md dirs and builds the same Map shape, and
_resolveManifest(commandsGsdDir, configDir) that prefers the source tree when
present (preserving repo-checkout behaviour) and falls back to the installed
skills layout otherwise. Both resolveCapabilityRuntimeState call sites use
_resolveManifest. Both helpers are exported for direct unit-testing.
Tests: capability-state.test.cjs gains a faithful installed-runtime e2e block
that copies gsd-core/bin + scripts + package.json into a temp install root
with no reachable commands/gsd, then runs the real gsd-tools.cjs against an
installed skills/ layout. It asserts capability state reports security &
nyquist enabled and verify:post includes security->secure-phase and
nyquist->validate-phase; a disabled-config negative confirms no
over-activation. This block FAILS before the fix (activeHooks:[]) and PASSES
after. Plus unit coverage for the two new helpers and the empty-surface
pre-fix scenario.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(changeset): backfill PR number
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>