Commit Graph

4768 Commits

Author SHA1 Message Date
Tom Boucher
3eb1cede26 fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent

Failing-first. Encodes the issue's runtime repro: a trailing comma in
.planning/config.json currently yields source:builtin-defaults with
degraded:false - byte-identical to the file not existing - and the user's
entire configuration is silently discarded.

Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather
than diagnostic prose, per the ADR-1411 amendment's test-methodology clause
and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by
monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000
(root bypasses mode bits).

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): distinguish a corrupt config from an absent one

loadConfigResolved wrapped the read, the JSON.parse and the entire config
build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell
through to the same defaults and the branches returned degraded:false -
actively asserting health over discarded configuration. A single trailing
comma in .planning/config.json silently replaced the user's whole config,
reporting source:builtin-defaults degraded:false, byte-identical to having
no config file at all.

ConfigResolution now carries a machine-readable reason. Genuine absence keeps
degraded:false / not_configured; a file that exists but cannot be used sets
degraded:true with config_unparseable or config_unreadable. The same split
applies to the root config and to ~/.gsd/defaults.json.

Control flow is deliberately unchanged. preflight_check reports cyclomatic
141 / cognitive 196 and 93 dependents on this function, with the guidance
that small edits beat one big one, so faults are CAPTURED at the existing
read sites and stamped onto the returns rather than the try/catch being
restructured.

Also carries the ADR-1411 amendment's wiring clause: loadConfig returns
.config alone to ~51 call sites and would never see the new field, so an
unusable file emits a deduplicated stderr diagnostic keyed on resolved path
plus errno. Without it the reason would be an unreachable field and the user
whose config was discarded would still get no signal - the actual defect.

Registers the config-loader seam in lint-resolution-provenance, which until
now guarded only agent-skills.

Caller audit: ConfigResolution.degraded has exactly one consumer outside this
module, cmdAgentSkills (src/init.cts:2259), which destructures
{config, source, degraded} - adding a field does not break it. Its --json IR
now reports degraded:true for a corrupt config, which is the intended fix and
the one observable behavior change.

Closes #1880

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): degrade when any config on the path is unusable, not just the last

Two defects found by isolated adversarial review of the first cut.

BLOCKER: the success-path return did not consult configFault. A corrupt ROOT
config whose workstream override happened to parse returned degraded:false /
reason:resolved - the root's settings silently dropped, which is the exact
failure this issue closes, reappearing for any project using workstreams. The
stderr diagnostic fired, so the out-of-band half worked while the in-band half
reported a clean resolve; a --json consumer saw health.

MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE
- so an empty workstream file inheriting a non-empty root reported resolved
despite carrying no settings. Emptiness is now judged on the file actually
read, snapshotted before normalizeLegacyKeys mutates it.

Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880
degraded contract; adds a fast-check property asserting a PRESENT file is
never reported not_configured whatever its bytes (CONTRIBUTING.md parser
rule); and asserts the literal enum values so the provenance lint's
configured_empty/not_configured markers check real assertions rather than
incidental prose in test titles.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): reject valid JSON that is not a config object at the read seam

The fast-check property added in the previous commit failed on both node
lanes: a config.json containing 0, "str", [], null or true is valid JSON, so
it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer
catch reported not_configured - a PRESENT file reported as absent, which is
precisely the collapse this issue exists to close. The property asserts a
present file is never not_configured, and it caught it.

_readConfigFile now validates shape, not just parseability (ADR-227: check
the semantic shape at a trust boundary, not merely the type). A non-object
JSON document is an unusable config, reported config_unparseable.

Adds named regression cases for each non-object form alongside the property,
so the class is documented and not only randomly sampled.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1880): backfill changeset pr number (pr:0 -> 2688)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2452): record a fetch-time shallow failure instead of crashing

This guard failed CI on ubuntu-24 while passing on ubuntu-22 and
windows-24 for the same commit, and passed on other PRs. Not a flake and not
caused by the change under test - a real fragility in the test.

runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it
assumed the failure mode is always 'fetch succeeds, diff reports no merge
base'. At a shallow boundary that lands short of the merge base, git can
instead fail during the FETCH ('unable to parse commit' - the boundary
commit's parent is not available). Which stage git fails at is version and
transport dependent, so on some runners the error escaped runnerDiff and
crashed the test rather than being recorded as the ok:false the assertions
expect. Both stages mean the same thing for what this guard protects: a
shallow base ref cannot resolve the three-dot diff.

Also drops two assert.match calls against git's stderr prose. 'no merge base'
and 'unable to parse commit' are the same condition reported at different
stages, and CONTRIBUTING prohibits raw text matching on subprocess output.
The typed outcome (ok === false) is the contract; the tests now assert that
plus the presence of a cause.

Found while investigating the red lane on #2688; fixed here per the no-defer
rule rather than filed.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 00:09:13 -04:00
Tom Boucher
c07734216f fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes

The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19
installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped
with no citation and no evidence quote.

Sourcing every axis independently against Kimi Code CLI's own docs (the issue's
explicit requirement — `kimi` and `kimi-code` are distinct products) showed three
values had been inherited from the Python `kimi` descriptor rather than sourced:

- `embeddingMode` imperative -> declarative. Kimi Code plugins are a
  `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic
  API (docs/en/customization/plugins.md) — the same shape as `codex`.
- `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own
  nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md).
  The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not.
- `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth
  bound is published, so the fail-closed sentinel applies over a guessed integer.

`dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact
layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is
registered with the host and `resolveDispatchType` maps every role onto
coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in
docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance
under Documentation gaps instead.

Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/
backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499),
and the install adapter is not selected by `embeddingMode` (install.js:543 always
uses the imperative adapter). The one visible effect is the curated profile pin,
which moves programmatic-cli -> declarative-cli.

Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit
member that has been in the closed vocabulary since kimi-code shipped.

Same defect class and countermeasure as #2598: pin the corrected values and require
the matrix to agree with the descriptor, because a descriptor/matrix disagreement is
how the gap survived.

Closes #2603

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed

Surfaced by the orthogonal review of this change. `negotiateHostCapabilities`
emits a sentinel-specific warning for every dispatch sub-axis carrying the
documented `undocumented` value — namedDispatch, nested, background,
subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell
through to the numeric guard and reported `host dispatch.maxDepth is missing or
not a number — treating as 0`.

That message is indistinguishable from a genuinely malformed descriptor, so a
correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the
sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this
PR's kimi-code correction adds a seventh, which is why it is fixed here rather
than left in place.

The numeric guard keeps firing for genuinely malformed values; both paths still
degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests,
including the boundary case that the sentinel carve-out must not swallow a real
malformed value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2603): backfill changeset PR number (#2687)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:52:57 -04:00
Tom Boucher
a3853472de fix(#2598): declare OpenCode subagent dispatch synchronous, not background (#2682)
* fix(#2598): declare OpenCode subagent dispatch synchronous, not background

capabilities/opencode/capability.json advertised dispatch.background: true and
dispatch.backgroundDispatch: true. negotiateHostCapabilities and every
degradationFor / shouldFlattenDispatch consumer trusts these per-field values, so
declaring a capability the host lacks OVERSTATES it — the opposite of the
fail-closed posture the negotiation is built for.

The issue's own citations needed checking before acting: the host-integration
matrix (ADR-1239's designated deployment source-of-truth) documented `true` with
NEWER evidence than the issue cited, and explicitly marked the issue's
sst/opencode#5887 reference as a stale snapshot superseded by #2087. git log
confirms #2087 deliberately flipped these from false to true, citing OpenCode
v1.15.0/v1.17 as "background subagents enabled by default in all modes". Applying
the issue as filed would, on that evidence, have REGRESSED a deliberate update.

So the claim was verified against current upstream rather than either document.
`packages/opencode/src/effect/runtime-flags.ts` on `dev` today reads:

    experimentalBackgroundSubagents: enabledByExperimental("OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS")

`enabledByExperimental` falls back to the `experimental` flag and `bool()`
defaults to false — the parameter is hidden from the model unless an operator
opts in by env var. Upstream #29638 is still OPEN and confirms the session loop
`tasks.pop()`s one subtask at a time. #2087's "default-on in all modes" reading
does not hold against current dev.

The issue's CONCLUSION is therefore right even though part of its evidence was
superseded: concurrent dispatch cannot be relied on, so both fields are false.

The matrix rows are corrected with the verified citation rather than reverted to
the old #5887 quote, so the record shows why the value is false TODAY rather than
re-asserting evidence that was legitimately superseded. Neighbouring sub-fields
are untouched and pinned by test: namedDispatch, subagentToolkit, and
isolation:'orchestrator-worktree' (which works via `opencode run --dir` at the OS
process level and is unaffected — #2584 does not depend on this value either way).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2598): re-pin the dispatch contract tests to synchronous OpenCode dispatch

gsd-test on the descriptor change came back FAILED (5 unique, both node
versions). The failures were not incidental — they were deliberate contract-pin
tests encoding #2087's decision, one named literally "background UPGRADE":

  tests/host-integration-descriptors.test.cjs
    - EXPECTED_FLATTEN[opencode] === false (background-eligible)
    - the derived background-eligible set pin
  tests/opencode-imperative-reference.test.cjs
    - "descriptor declares background dispatch true/true (v1.15/v1.17 upgrade)"
    - "background UPGRADE changes shouldFlattenDispatch: false now"

So this is a recorded decision being reversed, not drift being corrected, and it
is reversed on evidence: current upstream `dev` gates the capability behind
OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS (default false) and upstream #29638
(OPEN) confirms the session loop still handles one subtask at a time. The issue
is filed by the maintainer and explicitly directs "update golden-parity /
validator fixtures as needed", which sanctions re-pinning.

Behavioral consequence, verified: shouldFlattenDispatch(opencode) now returns
TRUE, so GSD serializes opencode dispatch instead of trusting concurrency it
cannot get. That is the correct fail-closed direction and is safe today — no
shipped GSD flow drives OpenCode background waves (per the issue), and
isolation:'orchestrator-worktree' is unaffected because it works at the OS
process level via `opencode run --dir`, not via the native subagent.

Each re-pinned test now asserts the retracted contract in the opposite
direction — feeding the #2087 axes back in must still yield "would not flatten" —
so a silent re-flip of either field is caught rather than merely un-asserted.

lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2598): backfill changeset pr number (#2682)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:47:54 -04:00
Tom Boucher
a5633bb32f enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures

* feat(#2671): brand raw vs calibrated token types

* refactor(#2671): hoist type-compile into a before() hook

Two review responses:

- The fixture compile ran in the describe() body, so it executed at
  collection time even when the block was filtered out, and a failed
  precondition collapsed eight independent assertions into one opaque
  describe-level failure. A before() hook is this repo's documented
  idiom and preserves per-test granularity.

- parseTokensFlag now records WHY it returns an unbranded number: it
  validates the magnitude of --tokens, but the basis is decided by
  --calibrated, so branding here would be wrong for half its callers.
  The assertion belongs to cmdEstimateCheck, its only caller.

* test(#2671): pin each brand diagnostic to its OFFENDING marker

Adversarial review demonstrated that asserting only exactly-one-diagnostic-
at-code-N is not airtight. Repairing a fixture's brand violation while
injecting an unrelated error of the same code (a string passed as the
budget argument) still yielded exactly one TS2345, so the fixture would
have reported green while no longer testing its regression at all.

Each bad-* fixture now routes its violating value through a const named
OFFENDING, and the test asserts the diagnostic's start offset falls inside
that node — located through the AST, so it survives reformatting and never
pattern-matches source text. Replaying the proof-of-concept against the new
assertion rejects it: the diagnostic lands on the budget literal, not the
marker.

Also corrects a doc comment that claimed the program type-checks all of
src/; it covers phase-estimation.cts and its transitive dependencies.

* chore(#2671): backfill changeset PR number (#2676)
2026-07-26 21:42:50 -04:00
Tom Boucher
0d08c32048 fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable

Every emitted script was rejected. Four invalid constructs, the first fatal on
its own, so the Workflow backend could never dispatch a wave:

  1. no `export const meta = {…}` first statement -> whole script rejected
  2. resumeFromRunId("<id>")  -> "resumeFromRunId is not defined". It is a
     Workflow TOOL INPUT parameter, not a script function. The run id still
     reaches the caller via summary.resumeRunId, to pass as that input.
  3. budget(<n>)              -> "budget is not a function". `budget` is a
     read-only object { total, spent(), remaining() } fed by the caller's token
     directive; a script cannot set it. Recorded as intent in a comment.
  4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions".
     Now parallel([() => agent(…), …]) — passing agent() results directly also
     started every agent eagerly, before parallel() could bound concurrency.

The single-plan stage had its own branch with the same parallel() defect; both
branches are now one array-emitting path. Waves also emit phase() calls whose
titles match meta.phases exactly, so progress groups correctly.

Two secondary defects kept the script from ever being REACHED — which is why
this shipped undetected:

  5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no
     scriptable way" to introspect it and told callers to omit the flag, so
     gate 5 returned agent_sdk_version_unknown on every automated run while
     `capability state` still reported active:true. True for bash, false for
     Node: the router now reads the installed @anthropic-ai/claude-agent-sdk
     version, walking node_modules up the tree and reading package.json
     directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the
     SDK's exports map does not expose ./package.json. Precedence: explicit flag
     > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an
     unresolvable version still declines to inline. A too-old SDK now reports
     the truthful agent_sdk_version_below_floor instead of unknown.
  6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
     from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
     invocation without --runtime reported runtime_not_claude on an ordinary
     Claude project. Now delegates to runtime-slash.resolveRuntime.

The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}`
snippet is removed rather than repaired: it was also shell-dependent — zsh does
not word-split unquoted parameter expansions, so it collapsed to a single argv
element, argValue() never matched, and the run failed into the same
agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto-
resolution removes the need for the construct entirely.

Verified with the issue's own repro: no flags now reaches the version gate; an
SDK above the floor yields backend:"workflow" with a script that parses as a
real ES module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids

Findings from the isolated review, all fixed.

HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was
already RED because of it. The registry embeds the fragment text INLINE, so the
shipped/installed copy still taught the exact broken contract this PR fixes:
the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown"
guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of
checking its exit code, so I recorded a red chain as green — checking $? now.)

HIGH — three existing tests asserted the OLD broken shape and would have failed
CI; none was touched by the first commit:
  tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…")
  tests/claude-orchestration.test.cjs                 — .includes('budget(')
  tests/claude-orchestration-command-router.test.cjs  — .includes('budget(')
Each now asserts the corrected contract: the id/pool reaches the caller via
summary, and neither construct is ever CALLED. Two sibling assertions had also
gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the
new explanatory COMMENT contains that substring, not because anything is wired.
Rewritten to assert the real property.

MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked
within a wave, but nothing checked wave ids across waves. That was harmless
before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a
matching meta.phases entry and the tool matches titles by exact string — two
waves sharing an id would collapse into one progress group and misattribute the
second wave's agents to the first. Rejected at validation, with tests either
side of the boundary.

MEDIUM — the fragment contradicted itself (its "Manifest construction" header
still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never
told the orchestrator to pass summary.resumeRunId as the Workflow tool's
resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to
a tool-invocation input, an implementer following only the fragment would have
silently regressed phase-resume to a no-op. Both fixed.

MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and
docs/explanation/claude-orchestration-capability.md documented
`resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output —
teaching the bug as the feature. Updated to the real contract, including the
required meta block and the thunk-array parallel() form. (The changeset is
`Fixed`, so the docs gate exempts this; it is corrected because it is wrong,
not because a gate demanded it.)

LOW — the router's top-of-file comment still described the divergent
`--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the
fix; and inserting resolveInstalledAgentSdkVersion had orphaned
resolveDetectionArgs' JSDoc above the wrong function. Both repaired.

lint:ci now exits 0 (verified by exit code, not by reading output).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2590): backfill changeset pr number (#2681)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:49:38 -04:00
Tom Boucher
c3958018dd docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern

ADR-1411 reasons only about a resolution miss. It is silent on input that
is present but not usable, which is how five engine read paths (#1879) could
fold an unusable input into the value meaning 'genuinely absent' without
contradicting an Accepted ADR.

Read together, ADR-1411 and ADR-227 converge and do not license throwing as
the cluster's answer: ADR-227 requires malformed input to be coerced rather
than propagated and carves out only genuinely-fatal fields, while ADR-1411
already permits a fallback provided it is 'a visible value, not a silent
substitution'. The defect in these five sites is therefore not that they fall
back but that they fall back invisibly.

Records the pattern that follows: every current return value is preserved, and
the cause is made visible in-band where the result already carries a
provenance envelope, or out-of-band via a deduplicated stderr diagnostic where
it returns a bare value it cannot extend. Throwing stays confined to ADR-227's
genuinely-fatal carve-out, decided per call. Also names the per-applier caller
audit and the lint-resolution-provenance registry gap.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2674): prove the warning-state reset misses the unknown-key dedup set

The two existing cases in this suite only pass because each picks a key
name no other case reuses, so neither can observe whether the reset the
beforeEach calls actually runs.

Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after
_resetRuntimeWarningCacheForTests(). It is not - the helper clears only
_warnedConfigKeys despite documenting itself as resetting per-process
warning state.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2674): reset the unknown-key dedup set with the runtime warning cache

_resetRuntimeWarningCacheForTests documents itself as resetting per-process
warning state but cleared only _warnedConfigKeys, leaving
_warnedUnknownConfigKeys populated across cases. The suite that exists to
test that set - 'loadConfig - unknown-key warning dedup' - calls the helper
in beforeEach expecting exactly this, so the reset was a silent no-op for
it; both cases passed only because each picked a key name the other never
reused. Any later case reusing a key would have had its warning suppressed
by leaked state.

Found while amending ADR-1411, which names this dedup guard as the pattern
five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the
fix would have propagated the footgun to each of them. Folded in here per
CLAUDE.md's no-defer rule rather than filed.

RED verified on 3c4895841 (test only, no fix): linux-node22 reported
'FAIL tests/config-loader.test.cjs - the documented per-process
warning-state reset must clear the unknown-key dedup set too'.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): document src/ in the changeset-lint trigger list

CONTRIBUTING.md presented the Changeset Required trigger list as bin/,
gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which
scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the
TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is
the most-edited user-facing path in the repo and the omission sends any
contributor who touches it into a CI failure the doc says cannot happen.

Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so
running it bare locally reports success without evaluating the branch. This
PR hit exactly that: a local run said ok_fragment_present and CI failed
fail_missing_fragment on the same diff.

Found while opening this PR; folded in per the no-defer rule.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): restore the round-2 review corrections to the amendment

These edits were made in response to the second isolated review pass but
never staged: later commits used targeted `git add <file>` for the test and
the source fix, so the two markdown files stayed dirty and shipped nothing.
The branch carried the round-1 text, including the ADR-227 misquote the
reviewer raised as a blocker.

Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in
was never implemented, so citing it as the precedent was wrong), the dedup
key, #1882 folded into the out-of-band mechanism instead of a fourth
mechanism-less category, the narrowed caller-audit rationale, and the
test-methodology clause.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 20:25:39 -04:00
Tom Boucher
c7c2fe3c2b fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd

gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as
path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI,
hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the
workspace — so the lookup always missed. sessionStart could only ever emit the
"no .planning/ workflow found" nudge and stop's verify-work reminder could never
fire, even with .planning/STATE.md sitting in the workspace. Slash commands were
unaffected, which is why only the hook layer looked blind.

Both hooks already buffered stdin into `raw` and never parsed it; the payload's
workspace_roots carries the real path.

Multi-root was left open in the report ("first root vs any root"). Resolved
forward: prefer the first root that actually carries .planning/STATE.md, so a
workspace whose GSD project is not the first root still resolves — strictly
better than first-root-only and identical to it in the single-root CLI case.
Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE
ever invokes hooks from the workspace.

The resolver is duplicated verbatim across the two scripts rather than shared via
hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be
registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the
installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per
CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion
so the copies cannot drift.

Failing-first, demonstrated by direct invocation with cwd != workspace:
  pre-fix  sessionStart -> "no .planning/ workflow found"   stop -> {}
  post-fix sessionStart -> ".planning/STATE.md is present"  stop -> reminder

tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as
child processes with a cwd lacking .planning/ and workspace_roots pointing at it.
Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON
fail-open, junk-entry filtering, the parity assertion, and a guard that neither
script resolves .planning from cwd again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate

Three findings from the isolated review, all fixed.

1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical
   defect at line 43 — its own header documents workspace_roots in the input
   schema, but it resolved .planning/ from process.cwd() anyway. Under the
   cursor-agent CLI that meant every Cursor subagent (planner, executor,
   verifier) started with "no .planning/ workflow found" and no phase context.
   The report named only sessionStart and stop; the defect class was wider.
   Verified pre-fix vs post-fix by direct invocation with cwd != workspace.

2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and
   fell back to cwd solely when the array was EMPTY. So when roots were supplied
   but none carried .planning/ while cwd did, the hook reported absent — where
   the pre-fix code, which always used cwd, reported present. That contradicted
   the fallback's own stated intent of preserving IDE behavior. cwd is now a
   CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict
   superset of both the old behavior and the CLI fix, never a narrowing.

3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity
   fixtures store a content hash per installed file; these three hooks appear in
   13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the
   diff is exactly the three hook hashes in exactly those 13 runtimes.

Tests extended: subagentStart resolution via workspace_roots; the stop hook's
absent branch (previously only session-start's was covered); an explicit
regression guard that a project at cwd is still found when roots miss; parity now
asserts all THREE copies byte-identical; and the cwd guard sweeps the whole
RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module

The duplicate-plus-parity-test approach was the wrong call. The reported issue
named two hooks; a third (subagentStart) had the identical defect. That is the
signature of a systemic problem, and three copies of a resolver guarded by a
parity assertion is a divergence risk maintained by hand rather than a fix.

hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor
hooks require it; none defines a local copy. Divergence is prevented
structurally instead of by asserting three copies stay byte-identical.

The reason duplication looked necessary was real, and is fixed properly here
rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall
(#2089), so it never reaches the installer's bulk hooks/lib copy — it was the
ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19
golden fixtures: cursor had the hook scripts, no lib). A naive require would
have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging
every session on precisely the runtime this bug is about.

writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib
helpers the staged scripts actually require, discovered by scanning their
require('./lib/…') calls rather than a hardcoded name — so a future helper
cannot be silently omitted. This is narrower than flipping
skipSharedHooksInstall, which would wrongly pull in every shared hook.
cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the
manifest manage it for the runtimes that do receive hooks/lib.

Verified against a REAL install (runMinimalInstall, cursor/global): the helper
is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a
cwd that is not the project.

Also closes the review gap that the stop hook was excluded from the
cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity
test is replaced by a structural guard (every hook requires the shared module,
none redefines it) plus a new install test asserting the helper is staged and
the installed hook actually loads against it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker

Two findings from the installer-focused review.

H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated
the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js
from source, run the cursor install — it exits 0, prints "Done!", and ships the
three hook scripts with an EMPTY hooks/lib/. The installed hook then throws
`Cannot find module './lib/cursor-workspace.js'` at load, before its own
try/catch, wedging every session — and nothing surfaces until a user hits it.
The scan protected against a required-but-UNLISTED helper while leaving
required-but-MISSING wide open (typo, bad rebase, an accidental delete).
It now throws: a missing helper source is a packaging bug and aborts the install.

M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>`
marker that NOTHING substitutes: copyLibDir stamps .sh files only, and
writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified
the literal was reaching disk on both the bulk (--claude) and Cursor
(--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js —
carries no such marker, so this was newly introduced, not inherited. Marker
removed, matching that precedent, with a note on why. (The explanatory comment
deliberately does not spell the token out, or it would reintroduce the literal.)

M2 — the require-scan regex demanded the exact compact form, so
`require( "./lib/x.js" )` would silently fail to stage its helper and compound
H1. Now tolerant of interior whitespace and either quote style.

Regression test added for H1 — the reviewer confirmed the invariant had zero
coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make
writeCursorHooksJson throw rather than produce a broken install.

Re-verified end to end: the missing-source case throws, no unsubstituted literal
ships, and the installed hook still resolves the workspace from a foreign cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2587): backfill changeset pr number (#2680)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 19:57:58 -04:00
Tom Boucher
c87f6f358e enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#1854): offer restore for user-added files backed up on update

Adds a restore-custom-files gsd-tools verb and wires it into update.md as a
restore_custom_files step: plan, compatibility-check against the newly
installed release, then restore only on explicit opt-in. The backup is never
deleted, a shipped path is never overwritten, and a single unwritable entry
does not abort the rest.

Also drops the jq pipe from update-context field extraction (#2589 class,
missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): reject symlinked restore destinations and backup roots

Self-review of the restore path found two write-through holes: copyFileSync
follows a symlinked destination, so a link planted at the restore target wrote
outside the config dir with every ancestor still a real directory; and statSync
on the backup root followed a link, letting the walk read arbitrary files and
present them as the user's own backup. Both now lstat.

Also marks the report's path/detail strings as untrusted data in update.md so
the rendered step cannot carry instructions into the runtime model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): move the update-context jq guard into the #2589 sweep

update.md joins the AUDITED list rather than carrying a duplicate assertion in
the backup-restore suite, and the guard gains a negative-proof companion so
'no jq pipe' cannot pass by the fields simply no longer being read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): validate manifest files map shape before trusting it

Security review flagged that Object.keys on a non-plain-object files field
yields numeric-index keys matching nothing, so the managed-path check dies
silently while manifest_found still reports true. Shape, not just type
(ADR-227): an array or scalar files map is now an unusable manifest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): size the restore prompt by eligible_count

Spec review found the prompt was driven by entries.length, so a backup holding
only blocked entries asked "Restore 1 file(s)?" when accepting would restore
zero. The question now reads eligible_count, and an all-blocked backup reports
its reasons instead of offering a choice that cannot be honored. The decline
path names the resolved backup_dir rather than the bare directory name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): use t.skip on hosts without symlink support

A bare return in a node:test body registers as a PASS, so the four symlink
guards silently reported green on unprivileged Windows instead of skipping.
Adds the dangling-link destination case the security review called out, and
moves outside-dir teardown to t.after so a failing assert cannot leak it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): unfence restore hint, regen goldens, widen install timeout

Three gate failures from the c99d612a5 run, all root-caused:

1. capability-registry (3): update.md's decline message put an instructional
   'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged
   fences as shell blocks. Retagged both display blocks as text and switched
   the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users
   can actually paste.

2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every
   runtime fixture moved. Regenerated; the diff is exactly those two hashes
   per fixture, no other drift.

3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's
   before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22
   lane passed the SAME commit in 12.7s. A full install measures 13-30s idle,
   so 60s was under 2x headroom and shrinks with every file added to the
   payload. Raised to 120s, matching the heavy case already in this file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1854): backfill changeset pr number to 2679

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:50:56 -04:00
Tom Boucher
9ade602219 docs(#2585): single lockfile-driven bootstrap path in CONTRIBUTING.md (#2675)
Getting Started opened with `npm install`, contradicting both the `npm ci` quick
start fifteen lines below it in the same file and docs/contributing/bootstrap.md,
which states `npm ci` is required so installs are reproducible, lockfile-driven,
and fail fast when package-lock.json is out of sync. A contributor taking the
shortest path — the first copyable block — got the wrong bootstrap contract.

Rather than correct `npm install` in place and leave two near-identical blocks,
the duplicate "Bootstrap your environment" quick start is folded into Getting
Started so CONTRIBUTING.md carries exactly ONE fresh-checkout path, in the
canonical order from bootstrap.md (nvm use -> npm run check:env -> npm ci), and
names bootstrap.md as the source of truth for everything else (fnm/asdf/mise, the
environment validator, daily commands, troubleshooting). That satisfies the
issue's "keep bootstrap.md as the source of truth rather than introducing another
variant" — three competing variants become one pointer plus one sequence.

No regression test: the change is prose in a contributor guide, not a runtime
contract, so a readFileSync+includes assertion would be exactly the source-grep
test scripts/lint-no-source-grep.cjs rejects.


Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:38:51 -04:00
Tom Boucher
bd570618d4 feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop

* fix(#2632): calibrate against the raw projection so the loop converges

* test(#2632): add closed-loop convergence guard and codify the feedback-loop rule

* fix(#2632): pair calibration samples per plan; atomic write; amend adr

* chore(#2632): backfill changeset pr to 2672

* fix(#2632): retry renameSync on transient windows errnos and clean up the temp
2026-07-26 16:28:56 -04:00
Tom Boucher
920a5f3f06 fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep (#2673)
* fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep

The reviewer/workflow config lookups resolved scalars and object fields with a
`gsd_run query <cmd> … | jq … 2>/dev/null || <default>` shape. On any machine
without jq (the default on Windows/Git-Bash) the jq stage fails with exit 127,
the failure is swallowed by 2>/dev/null + the trailing || default, and the
variable comes back EMPTY — the configured per-lane model/host/budget is
silently dropped and the lane falls back to CLI defaults with no diagnostic.

gsd-tools ships native flags that do the same job with no external dep:
  config-get <key> --raw          (strips JSON quotes off a scalar)
  resolve-model <id> --pick model (descends an object)
  resolve-execution … --pick <f>  (same)
  verification.status … --pick status

Replaced every jq-piped config/model/verify lookup across review.md (×23),
plan-phase.md, ship.md, debug.md (incl. the redundant boolean coercion — --raw
returns true/false as bare tokens natively), autonomous.md (×2),
ai-integration-phase.md (×4), and eval-review.md. The legitimate structured-JSON
jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile)
are untouched — only the jq-replaceable lookups moved to the native flags.

Adds tests/fix-2589-config-get-no-jq.test.cjs: a source-invariant guard asserting
no audited workflow pipes config-get/resolve-model/resolve-execution/verification.status
to jq (fails-first on the pre-fix text, passes after).

* test(#2589): update autonomous-converge jq assertion to --pick; regen golden fixtures

Two test consequences of the workflow-doc edits in the prior commit:

1. tests/autonomous-converge.test.cjs pinned the OLD jq-dependent shape
   (`verification.status … | jq -r '.status//empty'`) as the canonical routing
   contract. The test's INTENT is correct (route human validation through
   canonical verification.status) but it over-specified the MECHANISM (the jq
   pipe). Updated the assertion to match the new native --pick status shape;
   the contract being guarded (canonical verification.status read before the
   human_needed branch) is unchanged.

2. The golden-install-parity fixtures (19 runtimes) record a content hash of
   every installed workflow .md; the 7 edited workflows changed those hashes.
   Regenerated via `npm run gen:golden` (the test's own failure message
   instructs this). Only the 7 edited workflow hashes changed in each fixture.

* fix(#2589): declare jq a prerequisite for the lanes that still need it; repair test file

Three defects in the first cut of the #2589 fix:

1. tests/autonomous-converge.test.cjs was a JavaScript syntax error. The regex
   literal /...2>\/dev/null .../ left the second slash unescaped, terminating the
   literal early and parsing `null` as regex flags:
     SyntaxError: Invalid regular expression flags
   The whole file failed to load, so every assertion in it — including the #1522
   and #1526 guards — silently stopped running. Replaced with the string-compare
   form already used at line 202 for the sibling shell-snippet assertion.

2. lint:ci failed. tests/fix-2589-config-get-no-jq.test.cjs buckets into the
   capped `config` production module via its `config-get-...` effective prefix,
   making it a novel offender against the 2-file cap. The test is about workflow
   documents, not the config module, so it is renamed to
   fix-2589-workflow-jq-dependency.test.cjs (free prefix) rather than growing the
   allowlist with a module that does not actually need a 5th test file.

3. The fix deleted the repo's only jq-prerequisite declaration. review.md:244
   ("install jq if missing") was the anchor plan-review-convergence.md cites by
   line number, and it went away with the jq pipes — while the ollama, lm_studio,
   llama_cpp, opencode, and agy lanes still hard-require jq to parse HTTP
   /v1/chat/completions responses, opencode's JSONL event stream, and agy's
   conversation cache. On a jq-less host those five lanes swallow exit 127 into
   empty output: the same silent-degradation class #2589 exists to close.

   detect_clis now probes jq alongside the other prerequisites and emits
   jq:available / jq:missing, and the five dependent lanes are treated as
   undetected when it is absent, with an install hint. The six lanes that do not
   need jq stay selectable. plan-review-convergence.md now cites the section by
   name instead of a line number that moves.

Regression guards added to the renamed test file: review.md must keep the jq
probe and must name all five dependent lanes, and no workflow may cite review.md
by line number. Workflow-size baseline and the 19 install-parity goldens
regenerated for the review.md / plan-review-convergence.md edits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2589): decide the ship verification gate on a single verification.status read

Isolated review finding (medium). Pre-fix, ship.md captured verification.status
ONCE into $VERIFICATION and picked status / next_action / next_command off that
cached JSON with three jq calls. --pick takes a single dot-path field, so the
mechanical conversion issued three separate queries up front: three node spawns
that each re-read the phase VERIFICATION.md and re-derive the commit-time vs
mtime staleness comparison, on every ship — including the common passing path
that never uses the two message fields. It also meant the gate's verdict and the
message shown to the user were derived from three reads with no guarantee they
observed the same state.

The gate now reads `status` once and decides. The two message-only fields are
read on the blocking path only, after PHASE_VERIFICATION_INCOMPLETE is already
determined — so the passing path costs one query instead of three, and a
concurrent write between reads can no longer make the gate and its message
disagree, because the block/allow decision no longer depends on them.

Adding a multi-field --pick to gsd-tools would have collapsed this to one query,
but that changes the flag's output contract and belongs in its own change.

Regression guard in tests/fix-2589-workflow-jq-dependency.test.cjs: ship.md must
read verification.status exactly three times total, the block decision must
follow the status read, and next_action / next_command must both appear after
the blocking prose so they cannot drift back onto the passing path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* docs(#2589): document the jq prerequisite for the five reviewer lanes that need it

/gsd-review's ollama, lm_studio, llama_cpp, opencode, and agy lanes parse JSON
GSD does not produce (OpenAI-compatible /v1/chat/completions responses,
OpenCode's JSONL event stream, Antigravity's conversation cache), so they require
jq on PATH. Nothing in docs/ said so. Records which five lanes need it, which six
do not, that reading configured models/hosts/budgets no longer requires jq at
all, and what /gsd-review now does when jq is absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2589): backfill changeset pr number (#2673)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:17:02 -04:00
Tom Boucher
89b673e40d feat(#2631): planner emits estimate and plan-checker surfaces the over-budget flag (#2670)
* test(#2631): failing-first planner estimate emission and over-budget surfacing

* feat(#2631): emit plan estimate and surface the over-budget split recommendation

* fix(#2631): extract sizing prose to references to fit planner and plan-phase caps

* fix(#2631): move estimate check to plan-checker; fix template regex and caps

* fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble

* fix(#2631): invoke estimate-check after the launcher preamble in plan-checker

* fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference

* chore(#2631): backfill changeset pr to 2670

* chore(#2631): backfill changeset pr to 2670
2026-07-26 14:29:02 -04:00
Tom Boucher
115433bba6 fix(#2539): anchor commit phase-token detection; drop silent wrong-branch switch (#2669)
* fix(#2539): anchor commit phase-token extraction to the phases/ segment; drop silent switch-to-existing

cmdCommit auto-detected the commit's phase from --files with an unanchored
`match(/(\d+(?:\.\d+)*)-/)`, which returns the leftmost digit-run-then-hyphen
anywhere in the joined path. A project_code ending in a digit (PROJECT_V2) made
`.planning/phases/PROJECT_V2-07-name/…` match the `2-` inside `V2-` before the
real `07-` token, resolving phase 2. findPhaseInternal also searches archived
milestones, so an existing archived phase 2 produced a real branch name and the
silent `git checkout <existing-branch>` fallback switched the whole working tree
onto the wrong branch in the same call that then committed.

The extraction now anchors to the directory segment immediately under
`.planning/phases/` (or `.planning/milestones/<v>-phases/`) and runs it through
the existing project-code-aware extractPhaseToken helper — the single owner
shared by the other 6 call sites — rather than introducing a fourth independent
copy of phase-token-matching logic. The auto-switch keeps create-if-absent only
(the #1278 intent: ensure the branch exists before the first commit on it); it
no longer force-switches an already-checked-out working branch onto a different
existing branch.

Adds two regression fixtures: a digit-suffixed project_code + an archived phase
whose number collides with the trailing digit (the silent-wrong-branch case),
and a pre-existing phase branch that must not be silently switched onto.

* test(#2539): assert non-silent warning; hoist execFileSync; normalizePhaseName guard

Address orthogonal-review findings on the #2539 fix:

- Spec AC2 ('an auto-checkout mid-commit must never happen silently'): the
  no-switch path now writes a 'Warning: resolved phase branch X already
  exists; committing on Y instead' line to stderr when checkout -b fails
  because the branch already exists. The regression test captures stderr via
  spawnSync and asserts the warning, so neither direction of the branching
  resolution is silent.
- Spec AC3 ('reuse normalizePhaseName/extractPhaseToken/stripProjectCodePrefix'):
  the token-acceptance guard now runs the candidate token through
  normalizePhaseName and accepts it only when it normalizes to a numeric phase
  form, rather than the brittle 'token !== phaseDir && /\d/.test(token)' check
  that leaned on extractPhaseToken's undocumented dirName fallback.
- Standards (Duplicated Code): hoist execFileSync/spawnSync requires to the top
  of the 'commit command' describe block instead of inlining them per test.

* fix(#2539): build phase-token shape from PHASE_NUMBER_TOKEN_SOURCE (#2128 guard)

The acceptance guard regex in detectPhaseNumberFromFiles was a hardcoded
`/^\d+[A-Z]?(?:\.\d+)*$/i` — a literal re-derivation of the canonical
phase-number grammar, which the #2128 phase-id drift guard
(tests/phase-id-drift-guard.test.cjs) rejects unless sanctioned with a
`// phase-id-owner:` marker. Build it from the single-owner
PHASE_NUMBER_TOKEN_SOURCE export instead, so this read-side acceptance check
cannot drift from every other phase-token reader. gsd-test reported this as
2 failures (linux-node22 + linux-node24) on the prior commit.

* docs(#2539): backfill changeset pr: 2669
2026-07-26 14:14:28 -04:00
Tom Boucher
46ba02acde feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs

* fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions

* fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests

* chore(#2630): backfill changeset pr to 2661
2026-07-26 01:42:47 -04:00
Tom Boucher
bf127b7d25 fix(#2565): route generate-claude-profile target through runtime policy (#2659)
* test(#2565): add failing regression for generate-claude-profile runtime target

cmdGenerateClaudeProfile hardcodes .claude/CLAUDE.md (project + global),
ignoring the runtime policy that #3163 wired into the sibling
cmdGenerateClaudeMd handler. These tests pin the parity contract for
project scope (codex -> AGENTS.md, env precedence, --output override,
claude preserved), global scope (codex -> <CODEX_HOME>/AGENTS.md, claude
preserved), and a divergence guard asserting both handlers agree on the
project instruction path for the same runtime.

Failing-first: all seven tests reproduce the bug on unmodified next.

* fix(#2565): route generate-claude-profile target through runtime policy

cmdGenerateClaudeProfile hardcoded .claude/CLAUDE.md for both project and
global scope, ignoring the runtime-aware resolution that #3163 wired into
the sibling cmdGenerateClaudeMd handler. The #3163 fix diverged when it did
not propagate here, so /gsd-profile-user kept writing Claude instruction
files on Codex installs (and other AGENTS-native runtimes: opencode, kilo,
kimi, antigravity, copilot).

Fix mirrors the proven #3163 pattern using existing policy primitives:
- Project scope resolves through getProjectInstructionFile(runtime) and a
  non-claude runtime wins over a stale claude_md_path (AGENTS-native
  projects must never write to CLAUDE.md).
- Global scope derives ~/.<config-home>/<instruction-basename> via
  getGlobalConfigDir + basename(getProjectInstructionFile), so codex lands
  at ~/.codex/AGENTS.md. Claude global is preserved byte-for-byte (no
  env-var drift beyond the prior hardcoded path).
- GSD_RUNTIME env var takes precedence over config.runtime.

A parity test asserts both handlers agree on the project instruction path
for the same runtime, guarding against future re-divergence
(CLAUDE.md 'Generative Fix Divergence' rule).

* docs(#2565): add changeset fragment for generate-claude-profile runtime fix

* docs(#2565): remove parenthetical product description from changeset

The product-name purity guard (#1777) flags 'ProductName (description)'
patterns in changeset fragments because the prose renders verbatim into
CHANGELOG.md at release time. The original fragment had 'Codex (and other
AGENTS-native runtimes)' in the bold header, which matched the banned
pattern. Reworded to drop the parenthetical; also trimmed the per-runtime
mapping (belongs in code comments, not changelog prose).

* docs(#2565): backfill PR number in changeset fragment

* test(#2565): isolate os.homedir() cross-platform in claude global test

The 'global scope: claude runtime writes to ~/.claude/CLAUDE.md' test set
only HOME to redirect os.homedir() at a tmpDir. On Windows, Node's
os.homedir() reads USERPROFILE (not HOME), so the child process still
resolved the real user profile and the path assertion failed (#2659 CI).
Set USERPROFILE alongside HOME so the isolation holds on both POSIX and
Windows. Production code is unchanged — it uses os.homedir() exactly as
the prior hardcoded path did.
2026-07-26 01:06:35 -04:00
Tom Boucher
ec681e3c21 fix(#2567): scope Paused At to ## Session + guard Last Activity date regression (#2660)
* test(#2567): add failing regression for stale state field overwrites

buildStateFrontmatter extracts Last Activity and Paused At from the full
STATE.md body via stateExtractField, which matches the first 'Field:' line
anywhere. Historical archive sections containing stale field-shaped lines
silently overwrite the correct frontmatter value on every sync, and because
the poisoning line stays in the body it regresses on the next write.

Same divergence class as Bug #2444 (which scoped Stopped At to ## Session
but did not propagate). Failing-first: all three tests reproduce the bug
on unmodified next (verified via the dedicated red run on the test-only
commit).

* fix(#2567): scope Paused At to ## Session + guard Last Activity date

Two complementary fixes for the stale-archive-overwrites-frontmatter bug
class, chosen per field semantics:

- Paused At is a session field: scope extraction to ## Session (via the
  existing matchSessionSection helper), exactly mirroring the #2444 fix for
  Stopped At. A stale 'Paused At:' line in an archive section can no longer
  win over the current value. Falls back to full body when no ## Session.

- Last Activity has no single canonical section (it appears in the
  preamble, ## Configuration, and ## Current Position across STATE.md
  layouts), so a section scope cannot reliably exclude archive copies.
  Instead guard the information-losing direction: when the body-derived
  date is OLDER than the existing frontmatter date, keep the existing
  value and description (preferNewerLastActivity). Applied at both the
  write seam (syncStateFrontmatter) and the read seam (cmdStateJson) so
  they agree. Non-date values pass through unchanged.

A first attempt scoped ALL current-state fields to the body preamble, but
that broke STATE.md variants where the fields legitimately live inside
## Configuration / ## Current Position (regressed 4 frontmatter.test.cjs
suites). This minimal fix targets only the two fields the issue names.

* docs(#2567): add changeset fragment for stale state field overwrite fix

* docs(#2567): backfill PR number in changeset fragment
2026-07-26 00:55:00 -04:00
Tom Boucher
4d6e49f4d2 fix(#2654): bump js-yaml past the merge-key DoS advisory (#2655)
* fix(#2654): bump js-yaml past the merge-key DoS advisory

js-yaml was pinned ^4.2.0, inside the vulnerable 4.0.0 - 4.2.0 range of
GHSA-52cp-r559-cp3m (YAML merge-key chains force quadratic CPU, CVSS
7.5). Bump to ^4.2.1; the lockfile resolves 4.3.0.

It is a devDependency with no reachability from shipped runtime code
under gsd-core/bin/ or src/ — the consumers are scripts/workflow-policy.cjs
and five test files. The path worth closing is CI: workflow-policy parses
workflow frontmatter during the Tests workflow, and on a fork PR that
frontmatter is attacker-controlled.

Scoped to js-yaml only. The remaining brace-expansion advisory is not
fixed by this and is deliberately left alone: npm audit fix takes the
high count from 1 to 5, because the three copies nested under eslint
land on 1.1.16, which still compares inside the advisory's <=5.0.7
range. Closing it needs an eslint major or an overrides entry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2654): backfill changeset pr number to 2655

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 23:45:26 -04:00
Tom Boucher
b897070de3 fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard (#2656)
* fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard

The tracked build artifact gsd-core/bin/lib/api-coverage.cjs had drifted
four days behind src/api-coverage.cts: PR 2551 landed the #2366 fix in
the source without regenerating the compiled output, so the module that
actually ships still carried none of it.

Regenerate the artifact, and add scripts/lint-compiled-artifact-sync.cjs
to lint:generated-sync so a tracked compiled artifact can never again
silently diverge from its source. The check derives its file set from
git ls-files rather than a hand-maintained list, and is regime-agnostic:
if these artifacts are later untracked and gitignored per ADR-457, the
tracked set becomes empty and the check passes trivially.

Verified fail-first: the guard exits 1 against the previously-committed
artifact (37731 bytes vs 38634 expected) and 0 after regeneration.

Also fixes a defect this change surfaced in tests/no-phantom-issue-refs:
its PHANTOM list still banned 2551 and 2361, but GitHub numbers issues
and PRs from one shared counter and this repo has since reached 2654, so
both now resolve to merged PRs. The guard was rejecting accurate
citations of them — it failed this very commit for naming PR 2551 as the
drift's provenance. Verified by replaying the guard's scan: 1 offender
under the old list, 0 under the pruned one. Only 3182 is still a 404.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2653): backfill changeset pr number to 2656

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 23:45:11 -04:00
Behruz Nassre Esfahani
a40ee8a5e7 fix(#2479): drop the codex hook-trust bypass flag and its capability probe (#2536)
* fix(#2479): drop the codex hook-trust bypass flag and its capability probe

The /gsd-review codex lane emitted a hook-trust bypass flag via a
capability-probed variable (#1115). Host-harness safety classifiers deny
commands carrying the flag (23/33 sampled invocations), one denial citing
the probe itself as intent, while flagless retries succeeded 32/32 — the
flag only bypasses persisted hook trust, a first-run condition with no
steady-state value. Remove both the flag and the probe per maintainer
direction (no config key). #1115's diagnosability half — stderr to .err,
folded into the lane on empty output — is untouched; its version-gate
becomes vacuous with no flag to gate. The regression test inverts: the
literal flag is now banned file-wide in review.md (covers continuation
lines, carrier variables, and probes), alongside bans on the carrier
variable and any codex help-grep probe shape.

Fixes #2479

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2479): add changeset for PR #2536

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2479): changeset body in house format (bold lead) + post-rebase fixture regen

Review round 2 Minor: wrap the changeset lead clause in the required
**bold** span (reviewer-supplied text, applied verbatim). Rebased onto
current next; golden-install-parity (19 runtimes) + size baseline
regenerated — delta confined to review.md's hash/size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 20:15:49 -04:00
Tom Boucher
6ee4349272 fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored)

* chore(#2537): backfill changeset pr to 2642
2026-07-25 05:30:20 -04:00
Tom Boucher
e4dd0cbdd5 fix(#2523): normalize --files to repo-relative; reject out-of-repo; gate push on git-add exit (#2638)
* test(#2523): absolute + mixed + out-of-repo --files paths

* fix(#2523): normalize --files to repo-relative; reject out-of-repo; gate push on git-add exit

* chore(#2523): backfill changeset pr to 2638
2026-07-25 01:50:36 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
461c744c31 fix(#2522): fold wrapped success-criteria lines into their criterion (#2637)
* test(#2522): wrapped + blank-line success-criteria parse

* fix(#2522): fold wrapped success-criteria lines into their criterion

* chore(#2522): backfill changeset pr to 2637
2026-07-25 00:36:05 -04:00
Tom Boucher
1a9ae7601b docs(#2629): adr — phase effort estimation & calibration design lock (#2636)
* docs(#2629): adr — phase effort estimation & calibration design lock

* docs(#2629): annotate phase-0 status and link adr cross-reference

* docs(#2629): derive estimate confidence from sample count, not self-rating
2026-07-25 00:21:59 -04:00
Tom Boucher
b97b3dad21 fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix) (#2634)
* test(#2517): guard plan/execute-phase omit model= when *_model inherit/empty

* fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix)

* chore(#2517): backfill changeset pr to 2634
2026-07-24 23:44:12 -04:00
Tom Boucher
07270cbdf8 docs(#2614): record general-purpose agent-prompt skills as out-of-scope (#2633)
Routes standalone agent-discipline prompt modules to the capability
ecosystem instead of the core skill layer. Records three grounds: skills/
is a generated 1:1 projection of commands/gsd (gen-plugin-skills.cjs, gated
by lint:generated-sync), ADR-857 D4 contribution hooks exist precisely so
prompt-woven behavior can leave core, and the self-rated-confidence
mechanism these asks center on is measured weak in
references/honest-verifier.md:25-29.

Scopes the denial with an explicit does-NOT-cover section so a future
triage keyword match cannot misapply it to loop-step-attached prompt fixes
or to externally-measured calibration, and makes the revisit condition
measurable against the honest-verifier baseline.

Closes #2614

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 22:51:05 -04:00
Tom Boucher
c2b498e452 fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH (#2628)
* test(#2498): guard branch-creation uses --no-track (all workflows)

* fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH

* chore(#2498): backfill changeset pr to 2628
2026-07-24 22:49:29 -04:00
Tom Boucher
57b2bd8368 fix(#2491): finish todos/done -> todos/completed rename (14 stale refs) + guard (#2626)
* test(#2491): add todos/done rename under-sweep guard

* fix(#2491): finish todos/done -> todos/completed rename (14 stale refs)

* chore(#2491): backfill changeset pr to 2626

* test(#2491): fix lint-legacy-dir-name + allow-test-rule-refs (split legacy token, add issue ref)
2026-07-24 21:36:33 -04:00
Tom Boucher
4a66d62d10 feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver

Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.

worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.

resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.

Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: rebuild tracked state-transition.cjs to match #2400 source

The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit 2bcfaa2e2) added the progress.total_plans frontmatter sync to source but the tracked bin/lib/state-transition.cjs was never rebuilt, so the fix was not shipping to consumers of the compiled artifact. The mandatory build:lib step for Phase 2 surfaced the drift; recompiling makes the already-merged, already-changelogged #2400 fix effective. Artifact-only resync (no source/test change); drift class tracked by #2591.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 21:23:06 -04:00
Tom Boucher
7e1c736a3e fix(#2556): rescue SUMMARY when cat-file reports absent (exit 128, not 1) (#2611)
* test(#2556): correct cat-file stubs to exit 128 + rewrite fail-closed tests to fail-open

* fix(#2556): rescue SUMMARY when cat-file reports absent (exit 128, not 1)

* chore(#2556): backfill changeset pr to 2611
2026-07-24 13:35:02 -04:00
Tom Boucher
0f46fa366f docs(#2606): ratify ADR-612 getMilestoneFromPhaseId bracket return-form (vN.0) (#2607) 2026-07-24 13:00:34 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
0xdhx
bd18a593a9 fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs (#2592)
* fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs

The gemini and claude blocks in gsd-core/workflows/review.md were the only
two of the ten prompt-fed reviewer legs with both `2>/dev/null` and no
empty-output guard. Any failure that wrote no stdout — CLI missing,
unauthenticated, rate-limited, crashed — left a zero-byte review file with
the only diagnostic evidence already discarded.

write_reviews substitutes each file's raw content into a `## <Reviewer>
Review` section with no marker separating "empty/failed" from "ran cleanly,
nothing to report", so a failed lane silently degraded the advertised
N-reviewer consensus to N-1 while present_results reported success. Both
/gsd:review and /gsd:plan-review-convergence share the invoke_reviewers
step, so both were affected.

Both legs now redirect stderr to a `.err` sidecar and write a diagnostic
stub with the captured stderr appended when the review file comes back
empty — the same shape the codex and cursor legs already use.

Scope limit, noted in the block comment: the guard only runs if the block
itself completes. A host Bash-tool timeout that kills the whole block skips
it, the same hard bound the OpenCode block already documents. The existing
timeout guidance (#2194) covers that case and is unchanged.

Regression test extracts the two dispatch blocks verbatim from the workflow
and runs them under bash against a failing CLI stub, asserting the review
file is non-empty and carries a diagnosable message plus the captured
stderr. It fails against pre-fix review.md (4 of 5 cases).

Golden install-parity fixtures and the workflow size baseline are
regenerated: one review.md hash per fixture, one size entry.

Fixes #2494

* fix(#2494): stamp the changeset fragment with the filed PR number

The fragment was authored with the documented `pr: 0` placeholder because
the PR number does not exist until `gh pr create` returns. Now that this
PR is #2592, stamp it — scripts/changeset/parse.cjs requires pr > 0, so
the placeholder would fail the Changeset Required check.
2026-07-24 12:50:34 -04:00
BeeHiggs
bf9fe4630d feat(#2249): bracket phase-id core grammar — parse/render/toDir round-trip pair (epic #612 PR-1) (#2258)
* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards

PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md).
Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single
canonical owner — as a pure, additive extension. The 17 locked exports and
PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical,
so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs)
stays green.

New pure round-trippable model (ADR Decision 4):
- PhaseId { project, milestone, phase, subphase?, plan? }.
- parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token
  `GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens
  (`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new
  parser — normalizePhaseName and every legacy reader keep accepting those
  tokens unchanged (conservative default; no existing path gains a throw).
- renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug`
  with a slug guard that sanitizes path-traversal input.
- getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone
  from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning
  the `vN.0` form (parity with READING-A). The optional parameter keeps the helper
  pure (no config read) and byte-compatible — every existing single-arg caller
  resolves to the unchanged READING-A body (ADR Decision 6).
- extractPhaseToken(dirName, convention?): bracket dir branch GATED on
  convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is
  string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal
  family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no
  string-only discriminator is complete — an ungated auto-detect silently
  reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit
  convention signal keeps every existing convention-less call site byte-identical
  (pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs).
- comparator: no new code — comparePhaseNum already orders the dot-decimal
  `PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a
  PR-2 resolution concern (bracketQualifiedKey), not core grammar.
- SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999}
  non-milestone guard; the bracket-prefix reading is gated the same way (an
  ungated read called `P0.0-foundation` a sentinel), legacy leading-int form
  unchanged.
- BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately
  more permissive than parsePhaseId — a read-tolerance source for PR-2, not the
  emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt
  owner so PR-2 builds every bracket read regex from the canonical source and
  check:phase-id-drift stays green stack-wide.

The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*`
grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the
config permits parses. parsePhaseId has no live callers in PR-1, so this grammar
choice is forward-facing for PR-2 with zero PR-1 behavior impact.

Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips,
full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases),
extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens,
sentinel + slug guards, bare-token rejection, exported-source behavioral
assertions, and two generative fast-check properties: render∘parse identity over
well-formed displays, and the toDir/disk↔display bijection. Plus a #2043
numeric-tail characterization (single- AND multi-digit rows) in
tests/phase-id.test.cjs pinning the convention-less reading byte-identical.

The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish)
and rebuilt by CI, so it is intentionally not committed.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3)

PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4).

B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers,
over-padded numbers, multi-space separators, stray whitespace), so
render(parse(x)) === x did not hold for every well-formed x as ADR-612
Decision 4 requires. Both branches now enforce canonicality by construction:
parse permissively, rebuild the canonical string via the same emit path
(renderPhaseId for display, a hand-rebuilt token for dir/token), and throw
"parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's
entry is removed — the match anchors now reject leading/trailing whitespace
outright, folding into the existing "not a bracket phase id" rejection.

M1 (major): toDir only ever guarded the slug; project/milestone/phase/
subphase were interpolated unsanitized, so a hand-built PhaseId (a
structural, not nominal, type) could smuggle a path-traversal segment onto
disk. Every field is now validated against the exact shape parsePhaseId
itself would produce before use.

M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling
trailing hyphen in the emitted dir name. toDir now throws in that case.

M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable
from the dir-branch's plan tail, so it silently broke the disk<->identity
bijection on read-back. toDir now rejects all-digit slugs.

Nits: toDir now rejects a non-string slug instead of coercing it to the
literal token 'undefined'/'null'; sentinel boundary tests added for
milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999},
not an inclusive range — these were already correct, now locked by test).

Test-first: every new assertion (concrete examples + fast-check mutation
property for B1; concrete cases for M1-M3 and the nits) was written and
confirmed red before the implementation changes, per repo TDD convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): reformat changeset body to house convention (review Mi2)

The fragment added in ab26190a was a plain paragraph — no bold headline,
no trailing issue reference. Reformat to the repo's
`**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g.
.changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md).

Uses (#2249), the issue every commit on this branch references, not the
PR number already carried in frontmatter (`pr: 2258`) — the changelog
serializer appends `(#{pr})` unconditionally, so a body also ending in
`(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered
bullet directly via parseFragment + serializeChangelog: it now reads
`... (#2249) (#2258)`, matching the dominant convention across the other
fragments (frontmatter pr = merged PR, body reference = originating issue).

Also moved the docs-exempt marker back before the paragraph -> after it
(matching the file's original order): the marker sits on its own line and
is stripped before the body is used, but placing it first left a leading
blank line in front of the bold headline once reformatted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
two property-generator coverage gaps the reviewer flagged; no source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged).

Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no
property fed a 3+-digit milestone/phase/subphase/plan through parse/render/
toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen
numArb to 1–999 so the round-trip and disk↔display bijection properties both
span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no
leading zero, so canonicality still holds). Add a concrete regression pinning
the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and
toDirs to 'GSD.100-05-feature' without truncation.

Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered
milestone/phase pad + whitespace mutations but never a subphase pad. Add
unpad-subphase / overpad-subphase to the mutation set and a generated
`includeSub` boolean that decides whether the canonical carries a `.SS`
(forced in for the subphase mutations so there is always a `.SS` to mutate);
non-subphase mutations keep their original no-subphase coverage.

Non-vacuity verified against the compiled lib by temporarily probing each
widened/new property and confirming it fails: round-trip counterexample
["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit
tokens are genuinely generated and reach the body; a no-op unpad-subphase
mutation trips the mutated===canonical guard (counterexample
["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached
with a subphase present. Probes reverted; numRuns unchanged.

Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail;
`npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major)

BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that
re-derived the grammar as an unbounded `\d+` literal instead of consuming
PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket
path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug
whose first word is a year) over-collected the token as `01-14-2026` instead of
`01-14`.

Interpolating the cap verbatim at every position was rejected on evidence: the
bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent.
The exactly-2 cap at the others would under-collect ids toDir itself emits —
`PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit
sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and
`[GSD.100] 05` is a pinned regression. Those positions are delimiter-
disambiguated (a required field separator; a dot a slug can never contain),
not heuristically recognized, so they have no year collision to defend against.
Upstream draws the same line for the same reason: core-utils/phase cap the
paired PLAN component while the leading phase component stays unbounded.

So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and
each position takes the width its delimiter affords: leading unbounded, dash-1
and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner
seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of
the token grammar.

Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE
instead of re-spelling it as a literal, so the emit-side gate and the read-side
token source are one rule — the same single-owner discipline this fix is about.
Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant).

Refs #2249

* test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major)

The comment block alone cannot hold the divergence: src/phase-id.cts is exempt
from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not
catch the bracket token source drifting from the seam. Per the Generative Fix
Divergence rule, the divergence is pinned behaviorally instead.

Surface 6 joins the existing #2232 parity gate rather than starting a rival one:
the review named the bracket token source "a sixth continuation-recognition
site", and continuation-grammar-parity.test.cjs is already the invariant-named
home where the five #2043 sites agree with the owner on a shared width corpus.
Surface 6 asserts the same contract at the bracket run's slug-adjacent position
(`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so
the bracket path now fails the same gate the other five do.

A second block pins the DELIBERATE half — the wider canonical width at the
delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out
of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would
look like a cleanup rather than a regression.

The generative property ties the READ side to the EMIT side metamorphically: for
every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that
id's numeric run — no more, no less. It needed a new arbitrary: the existing
slugArb generates one [a-z0-9] word and so can never produce the number-leading
slug the collision requires.

Probe-falsified, both directions (probes reverted):
- reverting the source to the old unbounded `\d+` fails 8: the parity gate
  reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the
  review's scenario verbatim — and the property shrinks to
  ["A",1,1,undefined,"100-a"].
- interpolating the seam at EVERY position (the rejected verbatim option) leaves
  the repro and parity green but fails the divergence gate `'02' !== '02-105'`
  and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the
  evidence that a verbatim cap under-collects ids toDir emits.
Width 2 stays green under both probes — the corpus agrees with the owner exactly
where the old and new rules coincide, so the gate discriminates rather than
merely mirroring the regex.

Refs #2249

* docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major)

* test(#2249): pin deterministic grammar boundary cases (re-review m1)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
the m1 proof gap — the grammar's bounds were exercised only incidentally
through the fast-check domain (1-999, [a-z0-9] slugs). No source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged).

Adds a deterministic boundary block (7 describe groups, +22 tests) pinning
the compiled lib's CURRENT behavior — a proof gap, not a behavior gap:

- m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse
  (display + dir) -> render/toDir round-trip byte-equality. The plan
  position is identity-symmetric (parse/render accept 99/100/101) but toDir
  drops it (filename-surface dimension only).
- m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs
  99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps
  the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of
  the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry.
- m1.3 leading-zero 007 -> not-canonical rejection at every position/form.
- m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing
  slug (never stored, never mis-read as a plan) and rejects a line
  terminator; toDir's allow-list sanitizer collapses each to a safe
  [a-z0-9-] token or rejects sanitize-to-empty.
- m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test);
  an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE;
  an abs-path string is not a bracket id; an abs-path dir slug is dropped to
  a clean tuple.
- m1.6 whitespace-only -> not-a-bracket-phase-id.
- m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral):
  garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes.

No accept-not-reject case is a src bug: parse never STORES an abusive slug
(dropped from the identity tuple) and toDir independently re-sanitizes on
emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by
parse is the documented positional design (toDir drops the plan; the
read-token caps it) — divergence pinned, not papered over.

Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both
its parse-side and emit-side), ran -> 8 distinct named failures, reverted ->
66/66 green. Confirms every new group executes and can fail.

Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar +
continuation-grammar-parity + collision-characterization + phase-id family
175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green
except one pre-existing, unrelated env failure (npm-integrity-gate: a live
npm-audit advisory in the production dep tree — reproduces with this change
stashed; no package.json/lock change here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 12:49:53 -04:00
Rezolv
155c08facf docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase (#2574)
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase

These two commands never parse --validate (silent no-op); the flag is
real only for /gsd-quick. Remove the false flag-table rows and CLI
examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/
ko-KR/pt-BR mirrors), and correct the manager.flags.execute example
from --validate to --cross-ai (a flag execute-phase actually parses).
/gsd-quick's real --validate docs are left untouched.

Ref #2197

* docs(#2197): add changeset for --validate docs removal

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-07-24 12:47:53 -04:00
Tom Boucher
3b15a1e3cc fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos (#2597)
* test(#2576): add failing regression for padded resolves_phase compare

* fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos

* chore(#2576): backfill changeset pr to 2597

* test(#2576): drop fast-check property test (cross-platform-fragile on Windows CI)
2026-07-24 12:10:32 -04:00
Tom Boucher
be3bf97eff docs(#2584): ADR-1239 Codex-binding amendment + dispatch.isolation capability (Phase 0) (#2600)
* docs(#2584): add ADR-1239 Codex-binding amendment + dispatch.isolation capability

* chore(#2584): backfill changeset PR number (#2600)
2026-07-24 10:46:26 -04:00
Tom Boucher
a5180d96a3 fix: add regression-test-presence gate + missing tests for #2429/#2279 (#2563)
lint-fix-has-regression-test.cjs: new gate that fails if a fix(#NNNN)
or feat(#NNNN) commit has zero behavioral test files (*.test.cjs,
excluding auto-generated fixtures/baselines) in its diff. Wired into
lint:ci so it runs before PR creation.

Missing regression tests added:
- #2429: codex local scope does not set $HOME/.agents skills home;
  global scope does (tests/runtime-artifact-layout.test.cjs)
- #2279: map-codebase instructions say to overwrite existing dates,
  not just replace [YYYY-MM-DD] placeholders (tests/commands.test.cjs)
2026-07-23 12:13:12 -04:00
Tom Boucher
7d298d6d4d fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too (#2561)
* test(#2474): update dispatch gate test for dual-gate behavior

The #2772 test asserted the gate reads USE_WORKTREES_FOR_PLAN only.
Update to accept the dual-gate (USE_WORKTREES + USE_WORKTREES_FOR_PLAN).

* fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too

The per-plan dispatch condition checked only USE_WORKTREES_FOR_PLAN
(submodule-derived), ignoring the project-level USE_WORKTREES flag.
Add USE_WORKTREES to the gate. Net-negative edit: compress two
nearby prose lines to offset the added shell condition (93353 bytes,
down from 93368).

Closes #2474

* docs(#2474): backfill changeset PR number (2561)

* fix: merge coverage gate into single-process check (#2474)

The test:coverage:unit script chained two c8 invocations with &&:
the first ran tests and wrote coverage data to .nyc_output/, the
second read that data for per-file branch checks. On fast CI runners
(ubuntu/24), the second process started before the filesystem flushed
the first process's writes — a classic TOCTOU race that caused
intermittent coverage gate failures.

Replace the two-process chain with a single c8 invocation that
generates both text and json-summary reports, followed by a Node
script (scripts/check-coverage-gate.cjs) that reads the JSON summary
once and checks both overall and per-file thresholds. No filesystem
race is possible because the JSON report is fully written before the
check script reads it.
2026-07-23 09:14:02 -04:00
Tom Boucher
ae8526cd50 fix(#2429): scope Codex skills home override to --global only (#2553)
* fix(#2429): scope Codex skills home override to --global only

The skills-kind home override (redirecting skills to $HOME/.agents) was
applied regardless of scope. Gate it behind scope === 'global' so
--local installs keep skills project-local under the config directory.

Closes #2429

* docs(#2429): backfill changeset PR number (2553)
2026-07-23 07:38:42 -04:00
Tom Boucher
2bcfaa2e27 fix(#2400): warn on planned-phase no-op + sync progress.total_plans (#2552)
* fix(#2400): warn on planned-phase no-op + sync progress.total_plans

Bug A: When STATE.md Current Position has no recognized labels (narrative
prose), emit a warning field so the workflow detects the no-op instead
of continuing with stale state.

Bug B: Sync progress.total_plans in the YAML frontmatter when a plan
count is provided, preventing contradictory state between frontmatter
(0) and body (actual count). This writes the explicitly-provided count,
not a re-derivation from disk (#500 safe).

Closes #2400

* docs(#2400): backfill changeset PR number (2552)
2026-07-23 07:38:21 -04:00
Tom Boucher
1482dc5ce0 fix(#2366): scope parseCoverageMatrix to recognized coverage tables (#2551)
* test(#2366): regression tests for parseCoverageMatrix scoping bugs

Bug 1: summary table outside matrix not parsed as data
Bug 2: multi-section matrix with repeated headers parses correctly
Bug 3: markdown emphasis on decision cell is stripped

* fix(#2366): scope parseCoverageMatrix to recognized coverage tables

Replace latching sawHeader with contextual inMatrix tracking that
resets on non-pipe lines, preventing summary tables from being parsed
as data (bug 1). Allow multiple headers for multi-section matrices
(bug 2). Strip markdown emphasis from decision cells before validation
(bug 3).

Closes #2366

* fix(#2366): update representative-corpus test to expect correct behavior

The test previously documented the known-buggy parseCoverageMatrix behavior.
Now that the fix is in place, test against the expected correct output
(expectedBlock, expectedCounts, expectedErrorCount) instead of the
currentBuggyOutput snapshot.

* docs(#2366): backfill changeset PR number (2551)
2026-07-23 07:38:00 -04:00
Tom Boucher
0ad3c5dd14 fix(#2279): refresh date stamps on map-codebase Update runs (#2550)
* fix(#2279): reword date stamping to overwrite existing dates on Update runs

The map-codebase agent and workflow instructions only said to replace
[YYYY-MM-DD] placeholders, but Update-path files already contain concrete
dates from the prior run. Reword to SET the date stamps unconditionally,
overwriting whatever date is already there.

Closes #2279

* docs(#2279): backfill changeset PR number (2550)
2026-07-23 07:37:37 -04:00
Tom Boucher
200daa456c fix(#2269): add --files to three unscoped query commit call sites (#2549)
* test(#2269): regression test for query commit --files scoping

Verify secure-phase.md, validate-phase.md, and next.md all pass --files
to their query commit calls.

* fix(#2269): add --files to three unscoped query commit call sites

secure-phase.md, validate-phase.md, and next.md were the only 3 of 65
query commit call sites that omitted --files, causing blanket staging
of .planning/ and committing unrelated files. Add --files with the
specific artifact path to each.

Closes #2269

* docs(#2269): backfill changeset PR number (2549)
2026-07-23 07:37:15 -04:00
Tom Boucher
77bf21b3a6 fix(#1995): widen worktree branch regex to accept agent-<id> namespace (#2548)
* test(#1995): regression test for agent-<id> branch namespace

Add failing-first tests proving that normalizeCleanupManifestEntry and
planWorktreeRecordAgent reject Claude Code's current agent-<id> isolation
branches (only worktree-agent-<id> is accepted). Boundary tests cover both
namespaces plus rejection cases.

* fix(#1995): widen worktree branch regex to accept agent-<id> namespace

Claude Code's isolation="worktree" branch naming changed from
worktree-agent-<id> to agent-<id>. Widen the regex in all 7 locations
from ^worktree-agent-[A-Za-z0-9._/-]+$ to ^(worktree-)?agent-[A-Za-z0-9._/-]+$
so both namespaces are accepted. Introduce a shared WORKTREE_AGENT_BRANCH_RE
constant in src/worktree-safety.cts to prevent future drift.

Closes #1995

* fix(#1995): update workflow guards, test assertions, and baselines

Widen the branch-check regex in execute-phase.md and execute-plan.md.
Update all test assertions that checked for ^worktree-agent- to expect
the widened ^(worktree-)?agent- pattern. Regenerate golden-install-parity
fixtures, agent-size-baseline, and workflow-size-baseline.

Closes #1995

* fix(#1995): update extractCwdGuardBash sanity check for widened regex

The e2e test's sanity check verified the extracted bash block contained
'worktree-agent-'. After widening to '(worktree-)?agent-', update the
check to match the new pattern.

* fix(#1995): widen missed workflow-guard branch check + changeset + lint fixes

- hooks/gsd-workflow-guard.js: widen startsWith('worktree-agent-') to
  /^(worktree-)?agent-/ regex — same defect class, was missed in prior commit
- tests/worktree.test.cjs: fix indentation regression from prior edit
- Add .changeset/1995-worktree-agent-branch-namespace.md (pr:0 placeholder)

Found by orthogonal code review (Step 4).

* fix(#1995): regenerate golden + size baselines for workflow-guard change

* docs(#1995): backfill changeset PR number (2548)
2026-07-23 07:36:53 -04:00
Tom Boucher
4fc89497d0 fix(#2505): set USERPROFILE in kimi-variant test env + add issue refs on allow-test-rule exemptions (#2545)
The kimi-variant-disambiguation test set only HOME in the spawnSync env,
but os.homedir() on Windows resolves USERPROFILE — so the installer never
found the probe config files and the 'variant mismatch' warning never
fired. Every other installer test in the repo sets both HOME and
USERPROFILE (agent-skills, augment-upgrades, antigravity-upgrades, etc.);
this one was newly written for Phase 5 and missed the pattern.

Also: both new test files carried allow-test-rule exemptions without the
'see #NNN' issue reference required by ADR-456, failing lint-allow-test-rule-refs.
2026-07-22 20:58:29 -04:00
Tom Boucher
aa0f7dee99 fix(#2460): pi before_provider_request fail-opens without explicit model_profile_overrides (#2499)
* fix(#2460): pi before_provider_request fail-opens without explicit override

pi/gsd.cjs's buildBeforeProviderRequestHandler unconditionally rewrote
payload.model to the built-in pi/sonnet tier default (claude-sonnet-5)
via resolveTierEntry's catalog fallback. For any pi user on a non-Anthropic
provider (kimi-coding, zai, openrouter, openai-codex, minimax, ...), this
silently broke every request: pi's chosen model was replaced with one the
active provider did not know.

The fix inspects model_profile_overrides.pi[tier] explicitly BEFORE calling
resolveTierEntry (which falls back to the built-in catalog and would mask
the 'user did not opt in' signal). When the user has not set an override
(or set it to null), the handler returns undefined — fail-open — and pi's
chosen model flows through untouched. Only an explicit opt-in via
model_profile_overrides.pi[tier] steers.

Tests:
- the ACTUALLY-REGISTERED handler fail-opens when no override configured
  (was: 'steers to default-tier model-catalog pi id' — encoded the bug).
- new test reproducing the reporter's exact repro (model: 'k3' → undefined).
- override path: explicit model_profile_overrides.pi.sonnet config steers
  to the user-configured model id.
- defensive: explicit null override also fail-opens.

Per the reporter's suggested fix #1 of #2460.

* test(#2460): regen pi golden parity + install tree fixtures

pi/gsd.cjs changed → pi install hash changed → regenerate the parity
+ install-tree fixtures via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.

* fix(#2460): treat empty-string override as fail-open (M1 review)

Per code-review M1 + security M1: an explicit empty-string override
(`{ pi: { sonnet: "" } }`) silently bypassed the fail-open guard because
the check was `=== undefined || === null` only. resolveTierEntry's falsy
`if (userRaw)` then fell back to the built-in catalog and rewrote
payload.model to claude-sonnet-5 — re-introducing the exact bug this PR
fixes, via a degenerate config shape.

Fix: widen the guard to also reject `''`. The test now exercises both
null and '' in a loop, asserting fail-open for both.

* fix(#2460): clear hono/@hono/node-server moderate advisories via npm override

GHSA-v422-hmwv-36x6-class advisories (3 moderate) appeared during this
PR's session:
- @hono/node-server <2.0.5 (path traversal on Windows via encoded paths)
- hono 4.3.3 - 4.12.26 (API Gateway v1 adapter drops distinct repeated
  request header values)
- both transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol
  /sdk@1.29.0

The npm audit fix re-resolved hono to 4.12.31 (within the existing ^4.11.4
range declared by MCP SDK), clearing the hono advisory without an override.
The @hono/node-server advisory cannot be re-resolved the same way: MCP SDK
pins @hono/node-server@^1.19.9, and the fix requires 2.0.5+. There is no
MCP SDK release that allows @hono/node-server@2.x (latest 1.29.0 is the
most recent), and bumping @anthropic-ai/claude-agent-sdk to 0.3.x does
not help (its peerDependency is still @modelcontextprotocol/sdk@^1.29.0).

The override (sibling to the existing 'qs' and 'body-parser' entries) is
therefore the only available tool — distinct from the body-parser case in
re-resolution.

Verified: npm audit --omit=dev reports 0/0/0/0 advisories.

Also: regenerated the pi golden-install-parity fixture (the pi/gsd.cjs
change in this PR altered the pi install hash).

* docs(changeset): add Fixed fragment for #2460 PR

The two-otters-jog.md changeset was created earlier but lost during the
cherry-pick detour to fix #2454 PR 1's npm advisory cascade. Recreating
here with PR number 2499 backfilled (no placeholder cycle needed).
2026-07-22 16:05:41 -04:00
Tom Boucher
d579daa3ed docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum

* fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value

* docs(changeset): backfill PR #2538 for Phase 6 (#2512)
2026-07-22 14:30:26 -04:00
Tom Boucher
28f79e6e22 feat(#2505): Phase 5 — Kimi variant install-time disambiguation (#2535)
* feat(#2513): Phase 5 — Kimi variant install-time disambiguation (descriptions + mismatch warning)

* fix #2513: use helpers.cleanup (Windows-EBUSY retry budget) not raw fs.rmSync

* fix #2513: spawnSync captures stderr (execFileSync drops it on exit-0)

* docs(changeset): backfill PR #2535 for Phase 5 (#2513)
2026-07-22 13:49:31 -04:00
Tom Boucher
f654c24a3e feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505)

* fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping

* fix(#2508): remove leftover old-preamble lines (keep prose variant only)

* fix #2508: prose-only reference file

* test #2508: regen golden install parity after workflow preamble additions

* fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines

* docs(changeset): backfill PR #2525 for Phase 4 (#2508)
2026-07-22 10:22:42 -04:00