* fix(#1374): surface diagnostic when configured agent skills all fail to resolve
buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.
Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1374): backfill changeset PR number (#1376)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1355): detect-and-warn guard for claude-code agent-teams
GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.
- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
`query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): add changeset for teams-detect guard
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)
The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): register teams-status.cjs in the inventory manifest
The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1356): rewrite bare ~/.claude paths in the Cursor install branch
The Cursor branch of _applyRuntimeRewrites only rewrote the trailing-slash
.claude forms, so bare ~/.claude / $HOME/.claude references survived into
installed Cursor artifacts (skills/gsd-surface, skills/gsd-graphify,
workflows/plan-phase, workflows/autonomous), tripping the post-install
"unreplaced .claude path reference(s)" audit. Same regression class as
#983/#2418/#2545 — every other affected branch was patched; cursor was missed.
- Add the three bare-form rewrites (~/.claude, $HOME/.claude, ./.claude) to the
cursor branch, mirroring cline/trae/augment/codebuddy. They run after the
slash forms (no double-replace) and use (?![\w-]) so .claude-plugin /
.claudeignore are not corrupted.
- SKILL.md content also passes through this stage, so no stage-1 converter
change is needed. Verified: 40 bare refs in the real leaking files → 0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1356): add changeset for Cursor bare-path rewrite fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape
reconcileCodexHooksJsonEvent preserved whatever shape it read, so on an empty,
absent, or legacy top-level hooks.json it wrote top-level event keys
(`{ "SessionStart": [...] }`) that current Codex (deny_unknown_fields) rejects,
instead of the canonical `{ "hooks": { "SessionStart": [...] } }`.
- Lift any top-level event arrays (legacy, empty, or mixed nested+top-level)
into the nested `hooks` table, merging same-named events so no user/legacy
entry is dropped and no stray top-level event key survives. Mirrors
reconcileCursorHooksJson.
- Collapse an empty hook table back to `{}` so removal on an absent file does
not write a spurious `{ "hooks": {} }`.
- Read path still tolerates both shapes; dedup/removal unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1348): add changeset for Codex hooks.json canonicalization
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1342): scope worktree-path-guard to GSD executor runs; fail open for no-repo targets
The PreToolUse worktree-path-guard fired for any Write/Edit in any linked git
worktree, with no check for active GSD work — so Claude Code plan-mode writing
~/.claude/plans/<slug>.md from a manually-created worktree was hard-blocked.
- Gate enforcement on the GSD isolated-executor branch namespace
(^worktree-agent-[A-Za-z0-9._/-]+$, per worktree-branch-check.md #2924); the
guard is a no-op in non-GSD linked worktrees.
- Fail open when a target resolves to no git repository (e.g. ~/.claude/plans/)
instead of blocking — that is not the #260 main-repo vector. A target inside
a .git directory still blocks (git rev-parse --is-inside-git-dir).
- The #260 different-git-root hard block (escape to the main repo) is preserved.
Detached-HEAD executors no-op the gate; this is accepted because they are
fail-closed by worktree-branch-check.md (exit 42) before committing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1342): add changeset for worktree-path-guard scoping fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1342): build dot-dot traversal path portably (Windows drive-letter fix)
The traversal test built its file_path by stripping a leading slash from an
absolute externalDir and path.join-ing it after a `..` chain. On Windows the
drive letter (C:\) is not a leading slash, so it survived and path.resolve
produced an invalid doubled-drive path (C:\C:\Users\...), which resolves to no
git repo — the hook failed open (exit 0) and the test expected a block (exit 2).
Use path.relative(worktreeDir, externalTarget) + string concat so the file_path
carries literal `..` segments that resolve to externalTarget on both posix and
win32 (no drive doubling). Verified with path.win32/path.posix.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1343): parse decision bullets with text before the colon
parseDecisions() silently dropped any `- **D-NN ...:**` decision bullet
whose header had freeform text (a parenthetical, em-dash, or prose) before
the `:**`, so the blocking check.decision-coverage-plan gate computed
coverage over a narrowed set and reported a false pass.
- Broaden bulletRe to tolerate a freeform run before the colon while
preserving the optional [bracket] tag capture (drives `trackable`).
- Add a parse-miss guard: a line that looks like a D-NN bullet but still
fails the regex flushes the current decision and warns instead of
vanishing — the gate-integrity floor.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1343): add changeset for decision-coverage false-pass fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1343): relocate decision-parser regression into owning module test file
CI's lint-regression-test-names bans new bug-NNNN-*.test.cjs files. Move the
9 regression cases from tests/bug-1343-parsedecisions-drop.test.cjs into the
owning parser test file tests/post-planning-gaps-2493.test.cjs (which already
exercises parseDecisions) and delete the banned file.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones
Codex installs wrote an agents/openai.yaml sidecar under every managed gsd-*
skill dir. Recent Codex builds index both SKILL.md and the sidecar, so each
GSD skill appeared twice in autocomplete (canonical gsd-* name + humanized
display_name).
- Replace writeCodexSkillMetadataFiles / generateCodexSkillMetadataYaml with
cleanupCodexSkillMetadataSidecars: Codex-only (if isCodex), removes stale
managed gsd-*/agents/openai.yaml and prunes the now-empty agents/ dir.
- Preserve user-owned dirs (gsd-dev-preferences), non-empty agents/ dirs, and
non-gsd dirs; lstat-guard against symlinked agents/ so a delete can never
escape the skills tree; fail-open per directory.
- Codex relies on SKILL.md alone for /skills discovery.
- Update USER-GUIDE/FEATURES docs and rewrite the #774 emission tests into
cleanup tests.
Scope: the sidecar duplicate only. The separate multi-root (~/.agents/skills
shared-skills) duplicate facet is a distinct concern, not addressed here.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1326): add changeset for Codex sidecar cleanup
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile)
The map-codebase and docs-update workflows collected background sub-agent
results with the deprecated Claude Code `TaskOutput` tool using `block: true`,
which has a confirmed main-session hang after the agent completes
(anthropics/claude-code#20236).
Migrate the collection steps to the upstream-recommended pattern: keep
`run_in_background=true` on the Agent spawn, then `Read` each agent's
`outputFile` (from the `async_launched` result) once it reports completion.
Completion-marker contracts and on-disk verification are unchanged, and the
non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are
preserved byte-for-byte. docs-update's timeout note no longer references the
unwired `workflow.subagent_timeout` key (it kept a literal before).
Regression coverage folded into tests/subagent-timeout.test.cjs (the owning
module for background-subagent collection). Workflow size baseline regenerated
for the justified prose growth.
Refs #1355 (same latent hang surface).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1359): add changeset for TaskOutput migration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.
- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
(blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
fast-check round-trip property extended to the 4th scalar (the contract trek-e
blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
#1346 now tracks only the node-test causation residual.
190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
Addresses the #1314 maintainer review (trek-e):
- Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded
`if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT
point at a missing file; an honest negative test threw ENOENT *inside its
callback* (a failing test named distinctly from the file), which
isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash.
Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric
with the lint-rule path's file-result guard. Regression test pins it (RED without
the guard); a second test pins cwd-relative fixture resolution.
- Major 2 (misleading prose): the #1278 projection carries no violationFixture, so
the deterministic-locate path always hard-gates (fail-closed) until a
check_violation_fixture scalar is threaded through. verify-phase.md and
prohibition-probe.md no longer read as if the projected path produces greens;
the ADR-550 addendum records both items. Tracked as follow-up #1346.
- Documented residual: existence is necessary but not sufficient (a red caused by
the env being set vs the subject's content); recorded as a constraint, in #1346.
- Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'.
64 tests pass; eslint + tsc clean; changeset valid.
ci-prepare-test-scope.cjs's empty-detection FALLBACK hardcoded
tests/core.test.cjs, deleted in #1291. Every scoped lane (scope=targeted|
windows) that hit the fallback wrote the stale path into .ci-selected-tests.txt
and crashed run-tests with "requested test file(s) not found: core.test.cjs".
The full/sharded lanes glob the suite and were immune, so only the scoped
lanes went red (e.g. run 27599149212 on #1308).
Existence-filter the FALLBACK at write time and fall back to the 'unit' suite
sentinel (the #408/#641 path, resolved live by run-tests) when nothing
survives, so a stale reference degrades instead of crashing the lane. Detected
lists still pass through verbatim (they may carry a suite sentinel and are
already filtered by affected-tests-lib). Refactor to an exported, testable
resolveSelection().
Add a generative parity guard (DEFECT.GENERATIVE-FIX) asserting every FALLBACK
entry resolves on disk or is a known suite sentinel — it fails the instant a
refactor deletes a listed file, which #1291 did and CI did not catch — plus
resolveSelection unit tests and an end-to-end subprocess test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Make src/capability-activation.cts the sole owner of the four-level config-key
precedence walk via a new raw-value primitive resolveConfigKey(dotKey,{config,
cwd,registry}); the boolean wrapper _resolveActivationValue and loop-resolver's
resolveConfigValues are both rebuilt on it. loop-resolver deletes its
byte-identical copy and inline re-walk and imports the engine.
resolveCapabilityRuntimeState no longer returns registry/config (leaked internal
detail); the caller loads one fail-closed config snapshot and threads it in via a
new optional configOverride param, so capability `active` and hook when/configValues
resolve against the same object. capability-writer requires the registry module
directly.
Adds a DEFECT.GENERATIVE-FIX parity gate (identity + behavioral matrix +
end-to-end resolveLoopHooks configValues) that fails if the two precedence
surfaces ever diverge. Pure internal refactor, no user-facing change.
Closes#1308
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.
Closes#968
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A violation fixture that crashes the negative test at load emits a file-named
# fail 1 — a crash, not the assertion firing red. Require a failing test named
distinctly from the file (isNonVacuousNodeTestRed), symmetric with the clean-pass
non-vacuity guard and the lint-rule specific-rule-id requirement. Closes the one
soundness asymmetry surfaced by adversarial review (code-review IN-01).
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active
Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1307): add changeset for intel + loop-resolver active gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)
- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.
* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)
- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged
* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)
- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched
* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)
- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
src/prohibition-enforcement.cts: renames the projected flat scalars
check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
re-validate kind/target/rule — an under-specified descriptor reconstructs to
one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
and parseMustHavesBlock are unchanged (additive +36/-0).
* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)
- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof
* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)
- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged
* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)
- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)
* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset
- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)
* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES
- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
to the prohibition-probe reference (flat-scalar keys, projection/read-back,
fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
verification:test vocabulary, no new glossary term introduced
* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix
- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)
* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary
* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review
- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
reconstructs as a string and locates instead of silently un-locating; no
as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.
* chore(1278): set changeset pr to 1301
* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)
RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
- rename tests/_ff_lint_violation.test.cjs -> .cjs so node --test does not execute the lint fixture (ENOENT on intentional lib/foo.cjs); scope local plugin to it so eslint . stays green (violation -> suppressedMessages, prover still proves it)
- migrate prohibition-probe.verify-tier.test.cjs A/B to inject proveFailFirst (FF-08: attestation alone no longer greens post-#1279)
- re-baseline verify-phase.md workflow size (35362 -> 35812) after the descriptor-shape prose grew
- changeset pr: TODO -> 1279 (issue number; update to PR number at open) so changeset + docs-required gates pass
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only
graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1306): add changeset for graphify tri-state gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Add 5 FULL-producer e2e tests to the REAL-runner describe block, exercising
runProhibitionEnforcement with NO injected runCheck/proveFailFirst (the
SHIPPING defaultProveFailFirst + defaultRunCheck spawn real subprocesses).
- lint-rule: greens on the committed no-source-grep violation fixture + clean
target via real eslint; hard-gates on a toothless fixture in both modes.
- node-test: greens via the GSD_PROHIB_SUBJECT convention (red on bad fixture,
clean pass on clean subject); hard-gates without a fixture (FF-05).
- Closes the #1259 BL-01/SF-01 injected-double gap for fail-first (FF-10 capstone).
RED-first fail-closed guards on the prove seam, committed BEFORE the producer
change. All inject (or omit) proveFailFirst, which the current producer ignores
— so each FAILS against the attestation-greens code (actual 'green' vs
notStrictEqual 'green'). RED is INTENTIONAL; Plans 02-03 wire the default prover
and no-throw wrap to turn them green.
New tests:
- FF-05 no-fixture: node-test with no violationFixture, default prover (none
injected) -> fail-closed, never attestation fallback (RED now)
- FF-05 no-throw: a proveFailFirst that THROWS fails closed, never greens (RED)
- FF-04 both-modes: un-provable check hard-gates in interactive AND autonomous,
mode echoed (RED)
No source edited. 5 total #1279 guards now RED, 2 positive controls green.
Pre-existing 5 attestation-green tests untouched and still pass.
RED-first adversarial guards pinning machine-proven fail-first BEFORE any
producer change. They inject a new proveFailFirst option the current producer
does not read, so they FAIL against the attestation-greens code
(src/prohibition-enforcement.cts:411). This RED is INTENTIONAL — Plans 02-03
wire the prover and turn them green.
New tests:
- FF-01: attestation + clean pass but no machine proof -> hard-gate (currently
RED: producer greens on attestation; actual 'green' vs notStrictEqual 'green')
- FF-01 positive: proven-fail-first + clean pass -> green (passes now, pins shape)
- FF-04: passes-on-violation (prover could not prove red) -> hard-gate (RED)
- FF-04: fails-on-clean even when proven -> hard-gate (passes via runCheck:false)
No source edited. 2 new guards RED, 2 positive controls green.
* feat(#1305): add per-capability active tri-state + isCapabilityActive to the Capability State Resolver
CapabilityStateEntry gains active = enabled && configActivation, where
configActivation resolves the capability's optional activationKey via the
shared _resolveActivationValue (absent activationKey -> true). enabled stays
installed && surfaced (unchanged). Each hook's active now also cascades the
capability config gate (active && configured). Adds isCapabilityActive(capId,
cwd) — a thin convenience over resolveCapabilityRuntimeState. cmdCapabilityState
emits active per capability. No consumer cutover yet (graphify/intel: #1306/#1307;
loop-resolver: #1310). Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1305): add changeset for capability active tri-state
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add an optional activationKey to the feature role of capability.json — the
dotted config key that gates the whole capability (e.g. graphify.enabled).
gen-capability-registry validates it (non-empty string, reserved-name guard,
must be declared in the capability's own config slice, feature-only) and emits
it per-capability in the generated registry. Declared on graphify + intel.
No runtime consumption yet (resolver wiring lands in #1305). Part of #1302.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)
When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.
Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.
To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
(the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
menu (which for an isolated run now follows the fail-safe policy). Net effect:
execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
decision (quick.md is not size-capped).
Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.
Closes#1292
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1296): align config docs/prompts/schema with consumers
The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.
- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
said "seconds (default 600)" but the consumer (map-codebase.md) uses
milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
(prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
(test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
build_command in post-merge-gate) and documented, but absent from validKeys so
`config set` rejected them. Registered both in config-schema.manifest.json and
documented them in references/planning-config.md (overview + complete reference).
Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).
Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.
Closes#1296
Refs #1216
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Fixed fragment for #1296 config-surface alignment
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete)
After T0–T6 nothing imports core, so retire the spine and its scaffolding:
- delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact;
remove its .gitignore + eslint-ignore entries)
- delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the
package.json lint:ci chain
- regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface)
- sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired,
callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner,
and false present-tense core.cjs claims in leaf-module docstrings
The ADR-857 decomposition is complete: the former Core god-module is fully
dissolved into its leaf modules; no re-export spine remains. No behaviour change.
Closes#1294
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1294): migrate the computed-path core.cjs importers the literal grep missed
bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed
path, and bin/install.js was never in the convergence lint's scan roots), and
~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG
forms the literal-string migration grep missed. Route install.js's symbols to
their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET->
model-resolver) and repoint/adjust the test references to the leaves. Recovers
the 161 'Cannot find module core.cjs' failures from the spine deletion.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files
still imported core.cjs. Repoint all 33 behaviour importers to the leaf
modules directly (same symbol->leaf map as the src migration; leaves are the
objects core re-exported by reference), delete the now-meaningless
shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs
(forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot
test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs
(the lint is removed in T-final). Dropped the stale core.test.cjs entries from
the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path
pointed at io.cjs.
After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused);
T-final deletes it. No behaviour change.
Closes#1291
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth.
Fixes#1284
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
First leaf-migration tranche of epic #1267 (after T0 #1268). Migrate the
via-core callers of the two leaves T0 created to import from the leaf
modules directly, and stop core re-exporting their symbols:
- checkAgentsInstalled: docs.cts, verify.cts, init.cts -> agent-install-check.cjs
- gitWorktreeInfoInternal: init.cts -> git-base-branch.cjs
- getAgentsDir had no external via-core caller (internal to the leaf)
core no longer re-exports getAgentsDir / checkAgentsInstalled /
gitWorktreeInfoInternal; the now-unused agent-install-check + git-base-branch
requires are dropped from core; the shim-identity assertions for these are
deleted (behaviour tests retained). Convergence lint stays green (0 new).
No behaviour change.
Closes#1277
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier
Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.
The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.
Closes#966
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>