* fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones
Codex installs wrote an agents/openai.yaml sidecar under every managed gsd-*
skill dir. Recent Codex builds index both SKILL.md and the sidecar, so each
GSD skill appeared twice in autocomplete (canonical gsd-* name + humanized
display_name).
- Replace writeCodexSkillMetadataFiles / generateCodexSkillMetadataYaml with
cleanupCodexSkillMetadataSidecars: Codex-only (if isCodex), removes stale
managed gsd-*/agents/openai.yaml and prunes the now-empty agents/ dir.
- Preserve user-owned dirs (gsd-dev-preferences), non-empty agents/ dirs, and
non-gsd dirs; lstat-guard against symlinked agents/ so a delete can never
escape the skills tree; fail-open per directory.
- Codex relies on SKILL.md alone for /skills discovery.
- Update USER-GUIDE/FEATURES docs and rewrite the #774 emission tests into
cleanup tests.
Scope: the sidecar duplicate only. The separate multi-root (~/.agents/skills
shared-skills) duplicate facet is a distinct concern, not addressed here.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1326): add changeset for Codex sidecar cleanup
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.
- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
(blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
fast-check round-trip property extended to the 4th scalar (the contract trek-e
blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
#1346 now tracks only the node-test causation residual.
190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
Addresses the #1314 maintainer review (trek-e):
- Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded
`if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT
point at a missing file; an honest negative test threw ENOENT *inside its
callback* (a failing test named distinctly from the file), which
isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash.
Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric
with the lint-rule path's file-result guard. Regression test pins it (RED without
the guard); a second test pins cwd-relative fixture resolution.
- Major 2 (misleading prose): the #1278 projection carries no violationFixture, so
the deterministic-locate path always hard-gates (fail-closed) until a
check_violation_fixture scalar is threaded through. verify-phase.md and
prohibition-probe.md no longer read as if the projected path produces greens;
the ADR-550 addendum records both items. Tracked as follow-up #1346.
- Documented residual: existence is necessary but not sufficient (a red caused by
the env being set vs the subject's content); recorded as a constraint, in #1346.
- Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'.
64 tests pass; eslint + tsc clean; changeset valid.
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.
Closes#968
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active
Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1307): add changeset for intel + loop-resolver active gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)
- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.
* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)
- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged
* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)
- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched
* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)
- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
src/prohibition-enforcement.cts: renames the projected flat scalars
check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
re-validate kind/target/rule — an under-specified descriptor reconstructs to
one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
and parseMustHavesBlock are unchanged (additive +36/-0).
* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)
- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof
* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)
- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged
* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)
- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)
* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset
- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)
* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES
- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
to the prohibition-probe reference (flat-scalar keys, projection/read-back,
fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
verification:test vocabulary, no new glossary term introduced
* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix
- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)
* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary
* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review
- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
reconstructs as a string and locates instead of silently un-locating; no
as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.
* chore(1278): set changeset pr to 1301
* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)
RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only
graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1306): add changeset for graphify tri-state gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#1305): add per-capability active tri-state + isCapabilityActive to the Capability State Resolver
CapabilityStateEntry gains active = enabled && configActivation, where
configActivation resolves the capability's optional activationKey via the
shared _resolveActivationValue (absent activationKey -> true). enabled stays
installed && surfaced (unchanged). Each hook's active now also cascades the
capability config gate (active && configured). Adds isCapabilityActive(capId,
cwd) — a thin convenience over resolveCapabilityRuntimeState. cmdCapabilityState
emits active per capability. No consumer cutover yet (graphify/intel: #1306/#1307;
loop-resolver: #1310). Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1305): add changeset for capability active tri-state
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)
When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.
Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.
To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
(the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
menu (which for an isolated run now follows the fail-safe policy). Net effect:
execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
decision (quick.md is not size-capped).
Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.
Closes#1292
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1296): align config docs/prompts/schema with consumers
The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.
- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
said "seconds (default 600)" but the consumer (map-codebase.md) uses
milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
(prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
(test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
build_command in post-merge-gate) and documented, but absent from validKeys so
`config set` rejected them. Registered both in config-schema.manifest.json and
documented them in references/planning-config.md (overview + complete reference).
Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).
Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.
Closes#1296
Refs #1216
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Fixed fragment for #1296 config-surface alignment
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete)
After T0–T6 nothing imports core, so retire the spine and its scaffolding:
- delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact;
remove its .gitignore + eslint-ignore entries)
- delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the
package.json lint:ci chain
- regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface)
- sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired,
callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner,
and false present-tense core.cjs claims in leaf-module docstrings
The ADR-857 decomposition is complete: the former Core god-module is fully
dissolved into its leaf modules; no re-export spine remains. No behaviour change.
Closes#1294
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1294): migrate the computed-path core.cjs importers the literal grep missed
bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed
path, and bin/install.js was never in the convergence lint's scan roots), and
~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG
forms the literal-string migration grep missed. Route install.js's symbols to
their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET->
model-resolver) and repoint/adjust the test references to the leaves. Recovers
the 161 'Cannot find module core.cjs' failures from the spine deletion.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier
Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.
The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.
Closes#966
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Maintainer CHANGES_REQUESTED (reviewed e08667e5, pre-portability-fix):
- B1: lint-rule no longer greens an unparseable target (eslintHasFatalError -> fail
closed on any fatal/parse error) or an inline-suppressed violation (eslintJsonHasRule
now scans suppressedMessages too). RED-first + real-runner repros.
- B2: both child spawns get a bounded timeout (30s node / 60s eslint) + 16MiB maxBuffer;
timeout fails closed. Injectable timeoutMs (positive-only — 0/negative can't disable
the bound) enables a fast 1.5s hang test.
- M1: scoped verify-phase.md — the check descriptor is author-supplied for now; filed
#1278 for deterministic auto-locate of the descriptor (the locate half).
- M2: added tests/prohibition-enforcement.property.test.cjs (fast-check fail-closed
invariants; within the <=2-file budget).
- m1: tapTestNames excludes # SKIP/# TODO; parseNodeTestSummary tracks # cancelled;
isNonVacuousNodeTestPass requires cancelled===0.
- m2: scoped the determinism claim to the decision/parse layer (real runner is env-dependent).
- m3: filed #1279 for machine-proven fail-first (violation-fixture probe).
- n1: -- before target in both arg builders (option-injection). n2: dropped dead token.
- B3 (Windows npx) was already fixed in 2af76306 (pushed ~65s after the review).
Verified node 22 + 24; size baseline regenerated for the verify-phase note.
Re-home the 6 implementation functions squatting in the core.cjs re-export
spine (ADR-857) into the modules whose interface they belong to, with core
re-exporting them BY REFERENCE so all 32 callers + the shim-identity tests
keep resolving unchanged:
- worktree-safety: resolveWorktreeRoot, pruneOrphanedWorktrees
- git-base-branch (broadened to the Git Query Module): gitWorktreeInfoInternal
- agent-install-check (new leaf): getAgentsDir, checkAgentsInstalled
- delete the _resetRuntimeWarningCacheForTests wrapper; consumers use a
shared resetRuntimeWarningCaches() helper in tests/helpers.cjs
Add scripts/lint-core-spine-imports.cjs (migration-convergence lint with a
30-importer allowlist, wired into lint:ci) so the staged spine retirement
provably converges: CI fails on any new ./core import. Register the new
generated agent-install-check.cjs in eslint-ignore + .gitignore +
INVENTORY-MANIFEST.json.
No behaviour change. First tranche (T0) of epic #1267.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:
- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
reported test named distinctly from the file — node --test counts an empty file
as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
eslint as --format json and filters by ruleId, so plugin rules (local/*) load
via the flat config — bare --rule cannot load a plugin. local/no-source-grep
now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
producer requires attestation + a genuine non-vacuous pass. Machine-proven
fail-first (needs a violation fixture) is flagged as a tracked follow-up in
ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents
- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(#1243): document plugin-provided skills in agent_skills
Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.
Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)
- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1243): regenerate agent-size baseline for the Skill-tool grant
The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).
* chore(#1243): add Added changeset fragment
* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).
Closes#1110
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes#644.
Records the design gate for moving agent conversion off the inline bin/install.js loop onto the descriptor path (ADR-3660): the two-(really three-)path problem, the 10 parity behaviors the descriptor agents path must gain (verified against code — the issue's 7 plus Qwen/Hermes branding, the Codex TOML sidecar, and stale-agent cleanup), an AgentConverterContext contract to fix the leaky (content)=>string converter signature, and an incremental per-runtime cutover gated on byte-for-byte golden parity (full + minimal mode). Status: Proposed. Codex-reviewed for technical accuracy.
Closes#1235
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam
ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset
Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1239): ADR for GSD as an embeddable orchestration engine
Records the design to invert GSD from a standalone installer that projects
onto a host into an embeddable orchestration engine a host loads as a plugin,
driven through a negotiated host-integration interface. Unifies ADR-1016
projection (the declarative adapter) with imperative embedding behind one
contract; grounded in a 9-host capability survey.
Closes#1239
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1239): de-slash illustrative command placeholders (docs-parity gate)
Replace illustrative /gsd:x /gsd-x /gsd.x placeholders with namespace-prefix
wording so the docs-parity live-registry check does not parse them as
non-live commands.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)
ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for progress --converge surface (#1237)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline
CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1213): Capability State Writer — write-side inverse of the resolver
Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1213): add changeset for Capability State Writer (#1225)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes#1165.
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook
Claude Code marketplace plugin installs unpack the package into the
version-pinned plugin cache and never run bin/install.js, so
~/.claude/gsd-core/ is never created. Agents, commands, and templates
markdown-@-include the canonical ~/.claude/gsd-core/... path (which
expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to
nothing and agents (e.g. the executor) failed.
Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a
plugin install, symlinks the canonical path's immutable subdirs (bin,
contexts, references, templates, workflows) to the plugin's bundled
gsd-core/ tree. It changes zero @-references, is a no-op in classic
installs, preserves user-generated files (USER-PROFILE.md, STATE.md),
prunes stale links so it self-heals after `claude plugin update`, uses
Windows junctions, and rejects bundled/canonical paths that escape the
resolved plugin root (no traversal, no clobber).
Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json
SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES.
Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#997): backfill changeset PR number to #1207
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The triage-label mapping pointed ready-for-agent at the stale `confirmed`
label. The live verified-bug gate is `confirmed-bug` (RULESET.CONTRIB.CLASSIFY.fix
requires confirmed/confirmed-bug; 200+ issues and bug-remediation workflows use
confirmed-bug). Update the row + note so /triage applies the correct gate, and
mark `confirmed` as legacy/back-compat only.
Closes#1209
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1146): single base-branch resolver across forking workflows
Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).
New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.
Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.
Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(changeset): backfill PR number #1198
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1146): drop stray PR-body file from branch
pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1146): add tests for flat base_branch config key and both-branch tie-break
Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
per tryLocalBranch JSDoc) was documented but untested
9/9 tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1146): allowlist workflow-literal guard as runtime-contract exemption
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks
handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.
Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)
Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).
Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).
HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#956): backfill changeset PR number (#1201)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1196): wire discuss loop step for capability hooks
discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and
LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no
`loop render-hooks` dispatch and was absent from the conformance gate's
HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post.
- discuss-phase.md: add minimal discuss:pre (before analyze_phase) and
discuss:post (after write_context) render-hooks dispatch steps that
delegate consumption to a new shared reference (kept under the 32KB
#2551 budget; no inline subagent dispatch token).
- references/loop-hook-dispatch.md: new canonical, point-agnostic contract
for consuming `loop render-hooks --raw` activeHooks (contribution/step/
gate) — single source for hook consumption across host loops.
- gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and
export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host
file) — one source of truth for the host-loop file + wired-point set.
- phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES
and shared scanWiredPoints (was a hand-maintained duplicate omitting
discuss-phase.md + a duplicated regex).
- gen-capability-registry.cjs: add validateHooksWired() gen-time guard that
rejects a capability hook declared at a valid-but-unwired loop point, with
a clear remediation message — failure now surfaces at gen --check/--write
time instead of deep in the full conformance suite.
- tests (capability-registry.test.cjs): regression + anti-pattern parity
guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES;
POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into
the discuss-class gap again.
- docs/INVENTORY*: register the new reference.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1196): backfill changeset PR number (#1199)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: address adversarial-review findings
Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
capability-registry.cjs in place (concurrency hazard) — uses in-memory
checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
/* */ too, matching no-source-grep) so a block comment can't bypass it;
one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address code-review findings (subdir discovery, rule + test gaps)
xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{} equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)
The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: reconcile allow-test-rule allowlist after rebase onto next
Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(#1170): remove hand-maintained INVENTORY count scalars
The `(N shipped)` heading counts in docs/INVENTORY.md were absolute
scalars that collided silently on merge: two branches each bumping the
same integer to N+1 produced a clean git merge whose value the merged
filesystem (N+2) contradicted, hard-failing inventory-counts.test.cjs on
the CI merge commit across all platforms (DEFECT.INVENTORY-MERGE-UNDERCOUNT).
- Strip the six `(N shipped)` heading counts + the two prose footnote
counts; repoint the intro to INVENTORY-MANIFEST.json as the registry.
- Drop the decorative `generated` date from the manifest + its
strip-before-compare branch in gen-inventory-manifest.cjs (it conflicted
on cross-day merges and is read by nothing).
- Delete inventory-counts.test.cjs (scalar-vs-disk gate, the collision
source); its drift protection is subsumed by the merge-safe set-membership
test inventory-manifest-sync.test.cjs, which stays as the sole gate.
- Add inventory-headings-countfree.test.cjs guard (fails if a count is
re-added to a heading).
- Fix already-broken count-bearing cross-doc anchors to stable count-free
slugs in ARCHITECTURE.md + multi-agent-orchestration.md.
- Retire the now-impossible DEFECT.INVENTORY-MERGE-UNDERCOUNT + obsolete
RULESET.DOC-CONSISTENCY in CONTEXT.md; de-count DEFECT.INVENTORY-DRIFT;
correct stale MANIFEST-CANONICAL-KEY (all six families canonical).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1170): backfill changeset PR number (#1179)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink
The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).
Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)
First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:
- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.
Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): migrate profile-pipeline to a command-family Capability
ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).
Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1167): wire execute:wave:post + implement ui.safety-gate check
Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.
Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates
Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.
Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.
Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)
tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.
BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.
Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): migrate schema-gate to a plan:pre contribution Capability
The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)
Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN
Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.
Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract
Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.
Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.
Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)
The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.
Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.
Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)
3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).
Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)
Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.
Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests
The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):
Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
the inline schema_drift_gate step was removed — non-Claude runtimes would
stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
cite the execute:wave:post contract (loop body shrinks below the frozen
pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
baked verbatim into the committed capability-registry.cjs and leaked the
install path on 11 non-Claude runtimes (registry .cjs is copied, not
path-converted). Made the fragment path-free; regenerated the registry. The
phase-6 conformance gate now guards this (no ~/.claude install path in any
capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
step 6 (schema-gate is a plan:pre capability, §5.7 is gone).
Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).
profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.
Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables
Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).
Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.
ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.
federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): remove dead drifted converter dups + address adversarial review
Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.
Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
`node -e` form (node is guaranteed; matches the file's other node-e usages) so
a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
the orphan assertion). Now asserts the removed capability's key is genuinely
not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
`~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
structured-output check.
Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1169): make runtime-homes-descriptor-drive titles environment-independent
The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.
Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)
The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).
Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.
Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Auto-close bug reports opened without a valid GSD Version (Issue Forms enforce required only in the web UI). Bug reports only; version-shaped validation; version-exempt opt-out. Closes#1180
Proposes a runtime-gated, default-off BETA capability (post ADR-857) that
adopts Claude Code's Workflow tool — the engine behind /effort ultracode — as
an optional parallel-execution backend for the GSD loop, and folds the existing
ultraplan plan-offload under the same capability. Restores wave parallelism +
plan-checker + verifier on Claude Code that #853 currently forces inline.
Design-only ADR draft accompanying feature request #1143; implementation is
blocked on #857 being released.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis
The sandboxTier runtime-capability axis was cosmetic: declared and validated
on all 16 descriptors but read by nothing. The codex per-agent sandbox_mode
line was emitted unconditionally from the hardcoded CODEX_AGENT_SANDBOX map, so
the descriptor field drove no behaviour (ADR-857 audit finding F10; the
"rides along in 5e/5g" promise in ADR-1016 §8 never landed).
Make the axis load-bearing:
- resolveInstallPlan projects sandboxTier as a 7th InstallPlan axis and fails
loud (throws) on a missing/invalid value rather than coercing to 'none'.
- installCodexConfig / generateCodexAgentToml gate sandbox_mode emission on
sandboxTier !== 'none'.
- The per-agent CODEX_AGENT_SANDBOX map is kept: it is GSD agent policy, not a
runtime-descriptor property (different layer). Full removal of that map is
tracked under #1138 (phase-6 descriptor-residue removal).
For codex (sandboxTier === 'codex-agent-sandbox') the emitted TOML is
byte-identical to before; for 'none' runtimes sandbox_mode is omitted.
Adds leaf, projection, and installCodexConfig threading-seam regression tests;
updates the enh-1082 InstallPlan golden master with sandboxTier for all 16
runtimes. Confirmed hypothesis: schema-first vocabulary closure outran consumer
wiring, with no conformance gate to catch the orphaned axis.
Closes#1151
Refs #857, #1138
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1151): stamp changeset with PR number 1152
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>