Commit Graph

261 Commits

Author SHA1 Message Date
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
cbd180c5cd test(#3147): bound the lint/changeset/docs cluster onto the process seam (#3181)
* test(#3147): bound the lint/changeset/docs cluster onto the process seam

Migrates 69 unbounded sync spawn sites across 24 files. Allowlist 73 to 49.

Two shared helpers move: tests/helpers/graphify.cjs (6 importing suites) and
tests/fixtures/index.cjs, whose three quoted-argument shell strings became
single argv elements rather than whitespace splits.

changeset-lint's throw-native git() helper routes to gitOrThrow; migrating it
to bare runGit would have silently swallowed a failure that is loud today.
ingest-docs goes the other way -- its catch never rethrew, it degraded failure
into data every call site asserts on, so throwIfFailed would have thrown where
the original returned. The design doc said otherwise and was corrected.

tsconfig-noemit runs a real tsc --noEmit and takes a bespoke 180000ms per the
ensure-runtime-build precedent, not the 30000ms build-hooks norm -- that norm
is for a file copy, and sizing against a label rather than the work is the
same error in the opposite direction.

* test(#3147): add toLegacyResult and settle review findings

The seam exposed a throwing adapter (throwIfFailed) but no non-throwing one,
so eight files independently re-derived the same unwrap back to the legacy
{status, stdout, stderr} shape. That is the third time this epic produced N
copies of one mechanism -- seven throw wrappers in Wave 1, fifty-two timeout
constants in Wave 2, eight result adapters here. The pattern is that whenever
the seam does not expose a mechanism, every suite re-derives it.

toLegacyResult now sits beside throwIfFailed, with its own tests.

Two sites are deliberately NOT converted: changeset-cli's runRender and
runRenderIn return {status, report, stderr} from parsed JSON and never a raw
stdout, so they are a different shape family. lint-legacy-dir-name keeps its
local GUARD_TIMEOUT_MS: 30000 matches the build norm numerically but bounds a
lint probe, not hooks bundling, and importing it would encode a coincidence
as a relationship.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:45:51 -04:00
Tom Boucher
3fac6e629f test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam

Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73.

Timeouts are sized from evidence already in the tree rather than a house
default, because this wave spawns installers rather than git plumbing and an
undersized bound does not catch a hang -- it manufactures CI flake, which is
worse, since a flake gets re-run instead of investigated. install.test.cjs
records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while
another lane passed the same commit in 12.7s, so full installs are bound at
120000ms against that recorded incident.

Also adds an auditable escape to the guard's timeout ceiling. The 600000ms
cap was set in #3143 from partial evidence, but fragment-single-edit-
propagation carries a documented, load-tested 900000ms bound on a run that
chains a full build plus eight generators -- the guard would have rejected a
correct timeout the moment that file left the allowlist. A value above the
ceiling is now permitted only with an inline allow-spawn-timeout-ceiling
marker carrying a non-empty reason. It raises the ceiling; it never waives
the requirement for a bound, which is asserted directly.

install-shared.cjs keeps its hand-rolled assert rather than routing through
throwIfFailed: its message embeds both streams, and throwIfFailed carries
only a trimmed stderr. The message now also names the outcome, so a bounded
timeout reads as such across its 38 importers instead of as
expected null to equal 0.

* test(#3145): extract class-norm timeouts and correct the build-hooks sizing

A pre-PR review found 52 copies of four class-norm timeout constants across
this wave. These are not per-suite fixture bindings -- they are shared facts
about how long a class of subprocess takes, derived from a recorded bench
incident. That norm already moved once (60000 to 120000 after a real
ETIMEDOUT), and 52 copies would have drifted the next time it moved.

Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and
converts the copies. A site that genuinely differs -- a real tsc compile, or
regen:derived -- keeps its own local constant with its own justification.

Also corrects a misclassification: scripts/build-hooks.js was sized as a
build at 120000 in twelve places and 60000 in another, but it compiles and
bundles nothing. Its own header says no bundling needed; it copies pre-built
files and syntax-checks them with vm. Three different values bounded one
script; now there is one.

* test(#3145): fix red CI — lint self-match and a Windows chunk overrun

Two failures on PR 3176.

lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The
fixture exists to prove an unrelated marker does NOT suppress the rule, so it
carries that marker's literal text as test data. Split via concatenation, the
same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own
self-match problem. The explanatory comment needed the same treatment.

The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped
seven minutes before the kill, so this was an overrun rather than a slow
chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs
regen:derived bounded at 900000ms, which is larger than the whole chunk
budget, so the chunk killer always fires first and it can never complete
there. Both the test and that bound predate this change; modifying the file
pulled it into the Windows targeted set and exposed it. Skipped on Windows
with the reason recorded; the Linux lanes cover it. The 900000 bound and its
ceiling marker are unchanged -- they are correct.

* test(#3145): refresh the stale test-timings cost table

The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs
packs chunks by measured duration from tests/test-timings.json, and an
unknown file falls back to the table's median weight -- advisory by design,
but it silently underweights exactly the files that matter.

Four of the failing chunk's 22 files were absent from the table, including
the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s
(it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s.
Both were weighted as average, so the chunk's total weight read 53.68 against
a budget of 60 and the packer produced a single chunk.

Regenerated from a passing full-suite run, per the remedy the script itself
documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since
gen-test-timings.cjs replaces the table wholesale rather than merging.

Proven against the real packer: the same 22 files now weigh 103.91 and split
into two chunks. No logic, budget, or timeout was changed; raising a budget
to make a red gate pass is not a fix.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:18:18 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
Tom Boucher
1d208e5af6 test(#3144): bound the git/worktree cluster onto the process seam (#3152)
* test(#3144): bound the git/worktree cluster onto the process seam

Migrates 180 unbounded sync spawn sites across 19 files. Every previously
unbounded call now carries an explicit timeout with a comment giving the
number and why.

The migration is not a callee swap. execSync and execFileSync throw on a
non-zero exit and the seam never does, so each site was classified first:
sites that rely on the throw route to gitOrThrow, and sites that already read
.status to detect an EXPECTED non-zero -- an intended cherry-pick conflict, a
rev-parse outside a repo driving a skip -- route to the never-throwing runGit
instead, which would otherwise throw on exactly the exit being probed for.

Two same-named git() helpers in worktree-cleanup.test.cjs have different
return contracts, one trimmed and one raw; both are preserved rather than
unified.

Collapses five hand-rolled throw wrappers onto one throwIfFailed in
git-fixture.cjs, which gitOrThrow now also uses so the shape cannot drift.

Allowlist drops 139 to 120; BASELINE lowered to match.

* test(#3144): fix pre-PR review findings

Documents throwIfFailed in the CONTEXT.md glossary and CONTRIBUTING.md --
it became the shared throw mechanism without either doc naming it.

Routes the sixth and seventh hand-rolled copies of the throw shape through
throwIfFailed (worktree-baseref-install, worktree-safety-reap); the first
consolidation missed both.

Converts ci-rebase-check's 8 fixture-setup calls from unchecked runGit to
gitOrThrow so a failed setup step aborts where it fails rather than
surfacing later as a confusing failure against the wrong subject.

Adds 12 direct unit tests for throwIfFailed, which until now was only
exercised transitively.

Splits verify.test.cjs's non-git grep/sed bound off GIT_TIMEOUT_MS.

---------

Co-authored-by: sim <sim@local>
2026-08-07 10:58:34 -04:00
Tom Boucher
2afe17bbdb test(#3143): add the no-unbounded-spawn guard and throw-preserving git fixture (#3150)
* test(#3143): add no-unbounded-spawn guard and throw-preserving git fixture

Adds the ESLint rule local/no-unbounded-spawn, wired into the tests/**/*.cjs
block, plus an allowlist that only ratchets down: a listed file with zero
violations reports its own entry as stale.

The rule resolves renamed destructures and chained requires rather than
matching literal callee names -- both forms exist in the suite today and a
name-only matcher leaves them permanently invisible. It resolves an options
object held in a single-write const, which is what keeps process-seam.cjs,
the bounded reference implementation, from flagging itself.

timeout: 0 and anything above the 600000ms ceiling are rejected as only
nominally bounded.

Adds tests/helpers/git-fixture.cjs so a migrated execSync call site keeps
its throw-on-non-zero contract; process-seam.cjs is unchanged.

* test(#3143): prove the allowlist guards can actually fail

Extracts the D4/D6/D7/D8 checks into pure helpers and drives each against a
synthetic fixture carrying an injected violation. Without this the suite only
proved that today's clean data passes, which a deleted check would also
satisfy.

* fix(#3143): close two ceiling and alias escapes found in review

Nested arithmetic bypassed the ceiling entirely: the numeric evaluator only
resolved a flat literal, so `timeout: 60 * 60 * 1000` (3600000ms, six times
the ceiling) fell through to trusted and reported nothing. The evaluator now
recurses through arithmetic and unary signs with a depth cap.

Alias resolution was traversal-order dependent, not scope dependent: a call
textually above its own require destructure saw an empty alias map and
reported clean. The map is now built in a Program pre-pass.

Also: an explicit timeoutMs:undefined no longer overwrites the git fixture
default via spread, adds the missing seam-routed rule test, and de-duplicates
the repeated try/catch in the fixture tests.

---------

Co-authored-by: sim <sim@local>
2026-08-07 09:43:36 -04:00
Tom Boucher
0e6fa2e2cf enhance(#3118): close the dead injectables and the shell projection follow-on — Wave 4 (#3124)
* test(#3118): failing-first coverage for the dead injectables and the shell projection

Adds the counter-tests Wave 4 closes against, before any fix:

- antigravityWatermark had zero test references. The four existing tests
  that look like watermark coverage hand the fallback a literal mark and
  never call the producer, so nothing pinned whether a real run's mark is
  correct. Covers all six branches plus the non-object cache classes.
- Pins the fail-open: a transcript read that throws reports lines:0,
  indistinguishable from a genuinely empty transcript, and the consumer
  then replays a previous run's review as this run's.
- Pins the export-line escaping across the repair, persist and win32 bash
  lanes, including the parity assertion that they must not diverge.
- sliceCurrentPositionSection: empty-vs-absent, fenced heading, second
  occurrence, H3, CRLF.
- Proves deps.progressProvider is inert by supplying a throwing stub to
  all ten transition intents.

Verification through the remote runner only.

Refs #3118

* fix(#3118): distinguish an unreadable transcript from an empty one

antigravityWatermark's final read can throw on a transcript that
indisputably exists. It returned lines:0, which is the same value a
genuinely empty transcript produces, so the caller could not tell the
two apart.

antigravityTranscriptFallback derives its skip from that count. A mark
of {convId:'c1', lines:0} for a conversation that pre-dates the run
makes it skip nothing and return the last PLANNER_RESPONSE in a
transcript written before this run started — a previous review
presented as this one's, which is exactly what the function's own
'never stale' docstring promises cannot happen.

The unreadable case now sets unreadable:true and the fallback declines
for a same-conv-id unreadable mark. An absent or empty transcript is
untouched: those genuinely have zero prior lines.

* fix(#3118): escape the export line for the file it lands in, not the echo

Three lanes emit export PATH="<dir>:$PATH". repair escaped it with
escapePosixDoubleQuoted; persist and the win32 Git Bash lane escaped it
with escapeSingleQuotedShellLiteral instead.

The single-quoting is correct for the echo, so nothing runs when the
user pastes the command. But the bytes appended to ~/.bashrc are the
export line itself, and inside double quotes in an rc file a $(...) or
a backtick in the directory name is command substitution that runs on
every new shell. Those characters are legal in a path on both POSIX and
Windows, so the path was reachable.

projectPathExportLine is now the single source of that line and escapes
for its final rc-file context; each lane still applies its own transport
escaping on top. fish keeps the single-quote escaper — its value really
does stay single-quoted.

The cmd.exe lane interpolated into a cmd double-quoted string with no
cmd-level escaping, so a quote closed the region and &cmd& ran. A quote
is reserved on Windows and cannot appear in a real path, so there is no
correct command to suggest: the win32 lanes now fail closed for one.

Metacharacter-free paths render byte-identically on every lane.

* fix(#3118): drop a stray carriage return and a deps field nobody reads

locateCurrentPosition subtracted a fixed one byte to exclude the newline
before the next heading, which assumes LF. On a CRLF document the slice
kept an unpaired trailing carriage return. It now walks back over the
newline and over a preceding carriage return if there is one.

StateTransitionDeps also required a progressProvider that 33 sites
supplied and no site ever called. A required field nothing reads widens
the module's interface without changing its implementation, which is the
shape epic #3051 cites as its reason for refusing blanket injection.
Removed along with the ProgressRecord alias that existed only as its
return type; state-document.cts's unrelated interface of the same name
is untouched.

* fix(#3118): stop an empty span duplicating bytes, and name the empty results

Three findings from the isolated review pass.

locateCurrentPosition could return end < start when the section was
empty and the next heading followed with no blank line between. Every
mutator splices with slice(0,start) + body + slice(end), so an inverted
span duplicated the region between them — a blank line silently
inserted into STATE.md on every transition, two bytes on CRLF. The span
is now clamped, and an empty section is a zero-length span, which is
what it always meant.

The win32 fail-closed path left the installer printing 'Add it with one
of:' with nothing under it. An empty shellActions folded two different
facts together, so projectPathActionProjection now carries a frozen
PATH_ACTION_REASON and the installer branches on it. Two empty results
with different causes staying distinguishable is the subject of the
epic this belongs to.

fish_add_path parses a leading dash as an option, so a directory named
-v printed 'No paths to add' instead of being added. Verified against
fish 4.8.1: the end-of-options separator fixes it.

Replaces the console-prose test the second fix first arrived with — a
regex over captured stdout is what CONTRIBUTING prohibits, and the
typed reason is the surface it asks for instead.

* fix(#3118): escape TOML control characters, and stop a test name overstating

Five findings from the two review axes.

escapeTomlDoubleQuotedString escaped only backslash and quote. TOML
basic strings also require U+0000-U+0008, U+000A-U+001F and U+007F to be
escaped, so a value carrying a raw newline or NUL wrote a config.toml no
parser accepts — rejecting the whole file, not just that value. Four of
its call sites write real config. Tab stays raw; the grammar exempts it.

The byte-identity test claimed every lane was unchanged for an ordinary
path, which is false: fish now takes the end-of-options separator on
every path, not only hostile ones. Renamed, and the one intended delta
now has its own named test instead of hiding inside a claim that read
as broader than it was.

Also: exact-equality assertions in place of substring checks that could
pass on a subtly wrong escape, newline and null-byte cases for all five
quoting primitives, and a temp dir registered with t.after so it is
removed when an assertion fails.

* docs(#3118): add the changeset fragments

* fix(#3118): degrade instead of throwing on a null conversation cache

A cache file whose whole content is the literal null — what a truncated
or zeroed write leaves behind — made both antigravityWatermark and
antigravityTranscriptFallback throw. JSON.parse('null') succeeds, so the
try/catch wrapping the parse never fired, and resolveConvId then called
hasOwnProperty on null.

Both functions advertise the opposite; the existing test next to them is
named 'a missing cache or transcript degrades to empty, never throws'.
Parsing successfully is not the same fact as the payload being usable,
and a guard that only wraps the parse cannot tell them apart.

resolveConvId is now total for any non-object input, so one guard covers
both callers. Caught by the null case in this wave's own cache matrix.

* test(#3118): correct a stale fish expectation and a parity comparison

The pre-existing 'POSIX persist mode escapes single quotes' test pinned
fish_add_path without the end-of-options separator this wave adds, so it
asserted behavior that is no longer correct. A repo-wide scan found one
such hardcoded expectation; every other site derives its expectation
from the projection.

The new parity test compared the token from a POSIX path against the
win32 lane, which posix-normalizes its input first — two different
inputs, so the tokens differed for a reason that had nothing to do with
the parity it claims to check. It now derives the win32 expectation from
the same input the lane receives.

* docs(#3118): reword a comment the injection scanner reads as an instruction

The scanner pattern act\s+as\s+(?:a|an|the)\s+ carries no word
boundary, so 'the same fact as the payload' matched on the tail of
'fact'. Reworded per the documented remedy for this collision.

The missing boundary is a scanner defect rather than a prose problem —
any contributor writing 'fact as the' trips it — but the pattern is
gate plumbing, which the sibling epic owns, so it is surfaced rather
than changed here.

* chore(#3118): backfill changeset pr number to 3124

* chore(#3118): backfill changeset pr number to 3124

* fix(#2784): make the negation scan single-pass and index it correctly

Three defects in the negation suppression added by #3127, all in one
block, none of which had a test.

The pair scan was verbs.some(nouns.some(...)) with a slice and a split
per pair, so it grew cubically with clause length: 1.1ms before that PR
and 8462ms after, on 800 verb+noun pairs in one clause. api-coverage's
property test generates documents large enough to reach the runner's
600s file cap, which is why it hangs as 'fail 0, cancelled 1' rather
than failing an assertion. Every (verb, noun) window is a subset of the
single widest one, so one scan of that window answers the same question
in a linear pass. Verified equivalent against the old predicate over
20,000 generated clauses.

Both checks also subtracted clause.start from offsets that collectTerm-
Matches already returns clause-local. The first clause on a line has
start 0 so it worked there and nowhere else: later clauses went
negative, and slice reads a negative index from the end, so suppression
silently examined unrelated text.

The comment claimed 'without any API integration' was suppressed. It is
not — the qualifier sits outside the two-word lookback and the noun
precedes the verb. Widening the window would trade a false positive
that costs one declaration line for a false negative that slips a real
integration past a blocking gate, so the behavior stands and the
comment now says so. Pinned by a test.

The qualifier sets were also rebuilt for every line of every document.
2026-08-06 23:57:05 -04:00
Tom Boucher
5fd5c81042 test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it

Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync,
each returning a typed discriminated union
{ outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }.
Every call is timeout-bounded; there is no unbounded path.

runGsdTools becomes an adapter over the seam. Its legacy
{ success, output, error, exitCode } shape and retry-once-on-kill behaviour
are preserved byte-identically, so none of its 136 caller files change.

Outcome discrimination was corrected against probed runtime behaviour rather
than assumption: a timeout and a maxBuffer overflow are identical on both
status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs
ENOBUFS). Overflow is therefore classified before timeout. This fixes a live
defect — the previous isKilled() treated an overflow as a kill, retried it for
a second full 60s run, and then reported "host OOM or scheduler contention"
for a child that had merely printed too much.

Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs,
which brought 31 previously unlinted shared helpers under the same rules their
sibling test files already obey, and fixes the 5 violations that surfaced —
including a bare npm invocation without shell:true in
tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now
routed through the existing portable runNpm helper.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): migrate every local spawn wrapper onto the process seam

Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions
with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name,
parameter list, return shape and post-processing (JSON parse, ANSI strip, env
sanitising, field extraction) — only the spawn mechanism changes, so no test
assertion moves.

The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather
than a fourth primitive; it is explicit rather than inferred from the file
extension, because guessing an interpreter from a path fails silently when a
script's name does not match its shebang.

Seven wrappers were previously unbounded and now carry an explicit timeout
sized to what each actually runs, not the seam default. Two of those seven
(gsd-write-guard, lint-docs-command-form) were absent from the issue's
inventory entirely and were found by scanning after the migration.

Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md
reference section covering the three primitives, the discriminated union, and
the two rules the seam enforces.

Scope disclosure recorded in the phase design notes: the issue scoped three
identifier names. A scan for local helpers that spawn AND return the spawn
result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded
direct git call sites. This change bounds 25 of those. The remaining surface
is the same defect class and is NOT closed by this PR.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify an externally-killed child as KILLED, not EXITED

Blocker found in this branch's own diff, independently confirmed by an
isolated reviewer.

A child killed by an external signal — a genuine bench OOM kill — makes
spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field.
The seam's "no error implies EXITED" rule therefore classified it as a clean
exit, and runGsdTools returned { success: false, exitCode: 1 } without
retrying. That silently defeated the #969 kill-discrimination for precisely
the case it was built for: the old isKilled() fired on `signal != null`,
retried once, then threw a labelled resource-starvation error. A real OOM
would have been reported as an ordinary assertion failure.

Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes
the adapter retry on TIMED_OUT or KILLED — reproducing the old
`killed || signal != null || code === 'ETIMEDOUT'` condition exactly.
SPAWN_FAILED still does not retry (matching the old behaviour, where signal
was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate
divergence: the old code retried it because signal was SIGTERM, burning a
second 60s run on a child that had merely printed too much.

All five outcomes verified against the live runtime rather than assumed:
SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT),
>1MB stdout -> buffer_overflow (ENOBUFS).

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): address standards-review findings on this branch

Three findings from the standards axis of the review, all in this branch's
own diff.

The CONTEXT.md glossary entry this branch introduced was already stale on the
branch's own last commit: it enumerated a 4-member OUTCOME while the code had
5, because the KILLED fix did not update it. That is precisely the drift the
"module changes update Domain-terms" gate exists to catch, so the entry now
lists all five and explains KILLED.

api-coverage-gate-e2e compared an outcome against the raw string 'exited'
rather than OUTCOME.EXITED, the only such outlier; the enum is now imported
and used. A sweep for the other four outcome literals found no further
comparison sites.

Three call sites hand the literal bash flag '-c' to the seam's first
parameter, which the JSDoc described as an absolute script path. Rather than
add a fourth primitive, the contract is corrected to match reality: the
parameter is renamed `target` and documented as the first argv element handed
to the interpreter — normally a script path, but for an interpreter invoked
with an inline program it may be that interpreter's own flag. No behaviour
change.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): assert the cross-platform timeout contract, not the macOS one

The remote runner failed on both Linux lanes (node 22 and node 24, identical)
while the same tests passed locally on macOS. Two assertions encoded a
platform-specific behaviour as a cross-platform guarantee.

When spawnSync times out, macOS preserves the child's partial stdout/stderr;
Linux discards it and returns empty strings. Verified on node v26.5.1 both
ways. The seam passes through whatever spawnSync hands it and cannot
manufacture output that was discarded, so the production code was correct —
the tests were wrong.

Both tests now assert the guarantee the seam actually makes on every
platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being
strings rather than undefined or a Buffer. The partial-content assertions are
retained behind an explicit process.platform === 'darwin' guard so the macOS
coverage is not lost, and the first test is renamed to say what it now
guarantees.

This is the failure mode the remote matrix exists to catch: local macOS
verification would have shipped it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout

Windows CI caught two defects the Linux matrix could not.

tests/context-predicates-query.test.cjs passes a 32K-char argv value. On
Windows that exceeds the argv limit and spawnSync fails with code
ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise,
status === null implies TIMED_OUT" — swallowed it, so the adapter retried a
spawn that can never succeed and then threw the resource-starvation error. The
old isKilled() returned false for that shape and returned an ordinary failure
result.

TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set.
Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG,
E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform
whose timeout errno differs classified correctly, so the greedy catch-all is no
longer needed.

The second defect is a contract regression I introduced and had claimed
otherwise. That same test asserts `typeof r.exitCode === 'number'`, and
toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the
assertion failed on type. The old code returned `err.status ?? 1` on every
non-retried failure path. The adapter now returns 1 again for both, and the
comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is
retracted: the seam keeps the richer truth (exitCode null plus a distinct
outcome), the legacy adapter keeps the old numeric contract its callers
actually depend on.

Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT ->
SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL ->
KILLED; clean exit -> EXITED.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 01:06:39 -04:00
Tom Boucher
da062c0e0d chore(#2996): inventory the workflow fragment tree as its own manifest families (#3061)
* feat(#2996): inventory the workflow fragment tree as its own families

Epic #1671 Phase 6.5, the epic's last deliverable.

47 step files across 15 workflows and 13 mode files were invisible to
docs/INVENTORY-MANIFEST.json. Not through a missed row — through construction:
buildManifest walks each family with a flat readdirSync + isFile() and never
recurses, so nothing under gsd-core/workflows/<wf>/ could ever appear. modes/
has been invisible that way since #717 without any gate firing, which is the
evidence that this is a generator gap rather than someone forgetting a row.

Two new families, workflow_steps and workflow_modes, keyed by
<workflow>/<subdir>/<file> rather than a bare basename. That is deliberate: two
workflows may each own a regression-gate.md, and a step file may share a name
with a top-level workflow. The manifest is compared by JSON equality, so a
basename collision would silently drop an entry and read as "up to date".
Recursion is bounded at exactly one named subdirectory, and a limit+1 test pins
that bound so it cannot quietly become a general walk.

tests/inventory-manifest-sync.test.cjs carried its OWN duplicate copy of the
FAMILIES table — the DEFECT.GENERATIVE-FIX divergence class. Adding a family to
the generator alone would have left that test verifying six of eight families
while still reporting green. The table now lives once in the generator and is
imported, so the two surfaces cannot drift; runMain is guarded behind
require.main so importing does not execute the CLI.

The per-file roster stays in the generated manifest rather than being copied
into INVENTORY.md: 60 hand-maintained rows in lockstep with a generated artifact
is precisely the drift this file exists to catch.

CONTEXT.md's RULESET.MANIFEST-CANONICAL-KEY and DEFECT.INVENTORY-DRIFT both said
"six families" and now say eight, with the two key shapes and the import rule
recorded. The non-shipping example index was regenerated for the same edits.

Note on scope: this issue also asked for a one-fragment-edit proof. That landed
independently as PR #3046 and is not rebuilt here.

Refs #2996

* fix(#2996): correct a fabricated roster and an inert coverage pragma

Isolated review returned one blocker and three lesser findings. All four were
real; all four are fixed.

BLOCKER — docs/INVENTORY.md claimed the workflow_modes roster was
"discuss-phase, sketch". There is no gsd-core/workflows/sketch/ and never has
been; the second member is `help` (4 mode files), exactly as the manifest
generated by this same diff already listed. A doc contradicting the manifest it
describes, in the PR whose whole purpose is closing doc/reality drift. The
adjacent hand-maintained "15 workflows" count is also removed: an unenforced
number in a table cell is the same staleness class this file exists to catch,
and no test guards table-cell counts.

MAJOR — the CLI entry guard carried `/* istanbul ignore next */`, which excludes
nothing here. This repo measures coverage with c8 (test:coverage:scripts-floor,
55% floor over scripts/**/*.cjs), and c8/v8-to-istanbul honors only
`/* c8 ignore next */`. The pragma looked like it was doing something and was
not — the same failure shape as a marker that looks like working gating.

MINOR — collectNested called statSync/readdirSync unguarded, so a dangling
symlink or an EACCES directory under any workflow's steps/ would throw uncaught
and red the manifest gate for the entire repo. An entry that cannot be statted
is, for inventory purposes, not a countable file — the same disposition as "not
a directory". Row 13c pins the behavior with a real dangling symlink.

Refs #2996

* chore(#2996): backfill changeset pr number to 3061

* test(#2996): guard the dangling-symlink row on Windows

fs.symlinkSync throws EPERM on Windows without elevation or Developer Mode, so
row 13c would red the Windows lane. Guarded with the repo's idiom — a
process.platform check plus a genuine t.skip() carrying its reason, never a bare
return, which node:test counts as a PASS and would hide the gap.

Worth recording why this was not caught here: CI classified this PR's diff as
inert (no bin/, gsd-core/, or src/ changes), so the full test matrix was SKIPPED
entirely — the 'full test (${{ matrix.os }}, ...)' job shows as skipping with
its matrix expression unexpanded. The Windows lane never ran. It would have
fired on the next PR that does touch core code, in someone else's change.

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:51:12 -04:00
Tom Boucher
ed360cd99f chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point

Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers
only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim
into every runtime — and agent text is loaded into a subagent's context on every
dispatch.

The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op
for agents: agents never traverse that function. Agent content is read for
emission at five independent points, and the obvious chokepoint
`stageAgentsForProfile` short-circuits on the DEFAULT `full` profile
(`skills === '*'` returns the real unstaged directory), so a hook placed there is
dead code on most installs.

Composition now happens at two call sites instead of five parallel surfaces:
`stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind`
routed through it via an identity converter) and the inline agent loop in
bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` ->
`.windsurf/` regex can never reach inside a marker attribute — the ordering
#2930 established for workflows.

`installCodexConfig` was the fifth read point: Codex embeds each agent's prompt
into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it;
the exhaustive per-runtime emission sweep found it. That is why the new guard is
behavioral rather than structural — a sixth read point fails the sweep without
anyone remembering to extend a list.

tests/agent-fragments-emission.install.test.cjs spawns a real installer for every
runtime at every agent-bearing scope, derived from RUNTIME_META and the
capability registry at run time so a new runtime cannot be silently
under-covered. It asserts markers are absent AND the `when="always"` body is
retained, so marker-absence cannot be satisfied by dropping content. An
identity-composer negative control proves the assertion can fail.

Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope
paths; red before the wiring on claude(global+local), zcode(global+local),
kimi, codex and opencode.

Refs #2995

* chore(#2995): give the tightest agents headroom and correct the design lock

Epic #1671 Phase 6.4, second half.

`agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE
cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now
extract reference material to `gsd-core/references/` behind an @-reference — the
documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy:

  gsd-verifier  49,140 -> 46,371 B   headroom    12 -> 2,781
  gsd-debugger  57,197 -> 48,851 B   headroom   147 -> 8,493

Byte accounting proves no content was lost: the combined agent+reference delta
is exactly the new files' headers plus the agents' slim replacement blocks. Each
agent keeps its routing table and a one-line summary per entry, so it degrades
gracefully on a runtime that does not inline @-references.

`agents/gsd-planner.md` is untouched and still passes both char guards
(49,130 < 49,152); it needed no change, so it took none.

The other nine LARGE/XL agents carry NO gsd:section markers, and that is
deliberate, not deferred. `when=` selection is read from
gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives
from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key,
no per-agent init entry point. An agent atom therefore fails admission gate (2)
("a fact the init seam demonstrably computes at a real entry point") and would
evaluate false forever while looking like working gating. Marking agents would
manufacture exactly the silent-inertness rot the frozen vocabulary exists to
prevent.

ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage
found against what actually merged:

  - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR
    amendment, which that bullet's own rule forbids. Recorded now.
  - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to
    "the LARGE/XL rollout phase". Five shipped; this one is permanently
    rejected, and that disposition lived only in a merged PR body.
  - Phase 6.4's own finding: emission extends to agents/, gating does not.

CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still
listed the original 4-atom vocabulary and described when= as "not yet acted on",
and Section Manifest Module still described InvocationFacts as
{waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract.

Inventory manifest regenerated AFTER build:lib per the documented ordering
landmine; 19 install-tree fixtures pick up the two new references.

Refs #2995

* chore(#2995): correct the compose-site count and mark the raw stager

Self-review found two comment defects in the prior commit. The agentsKind
comment claimed composition lands at TWO call sites; it is three, since
installCodexConfig's per-agent .toml writer was added after that comment was
written. And stageAgentsForProfile is now production-dead — both callers route
through the composing stager — while staying exported and unit-tested, which
makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged
source directory under the default profile, so a future caller would silently
reintroduce the marker-shipping path. Its JSDoc now says so.

* test(#2995): guard the marker-documenting-doc class for agents

Widening the composer's scope to agents/ makes reachable the exact class #2930
narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an
unfenced example is indistinguishable from a real marker, so the composer drops
that line from the emitted artifact.

Three rows. A fenced example must compose byte-identically. No shipped agent may
carry a marker outside a fence — asserted by parsing every real agent and
requiring zero explicit sections, which is what makes the fence protection
load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED
marker IS parsed as a real marker, so if that ever stops being true the second
row is guarding nothing.

Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it
had no production caller, which is false — bin/install.js's _stageAgents still
calls it, and its consumers compose before writing. Corrected to state the
invariant instead. And a let/const nit in the emission sweep.

* fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture

The first remote run came back red with three failures. Both root causes were
mine.

1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally
   contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status
   vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent
   no longer had it.

   Byte accounting said no content was lost, and byte-wise that was true — but a
   contract required those tokens to live IN THE AGENT. That is ADR-1671:66's
   flexReserve floor stated concretely: a load-bearing fragment must not be
   trimmed out of its host, and "the bytes still exist somewhere" is not the
   test. The two status tables are restored to the agent and deliberately
   mirrored in the reference with a note saying so, so the procedure there still
   reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083,
   rather than the 2,781 the first attempt claimed.

2. Row 12b of the new marker-documentation guard asserted that an unfenced
   marker example parses as a real marker, and threw instead:
   "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The
   fixture had put the OPEN marker inline mid-sentence, so it was correctly not
   recognised as an open while the close, on its own line, was.

   That is a real refinement of the hazard this guard exists for: only a marker
   on its OWN line is mis-parsed — which is exactly how a documentation example
   is normally written. Row 12b now uses a whole-line marker, and a new row 12c
   pins the inline case as explicitly NOT a marker.

No test was weakened to accommodate the change; the change was corrected to
satisfy the tests.

Refs #2995

* chore(#2995): backfill changeset pr number to 3058

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:10:31 -04:00
Tom Boucher
4eb8e3648c fix(#3050): consolidate the spawn-timeout predicate and propagate the unresolved-root reason (#3060)
* chore(#3050): changeset and review artifacts for the follow-up

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3050): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 17:21:24 -04:00
Tom Boucher
ef823ca9d9 fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking

Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct,
transitive (2 and 3 hop), and diamond dependents of a halted plan across
both independent "which plans are incomplete" readers (phase-plan-index's
cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative
case (an unrelated decoupled plan stays runnable), and a parity check that
the two readers agree. Uses only modules that already exist at this
commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs)
so the test file loads and runs cleanly on a fresh clone of this exact
commit. These fail against current behavior: neither reader has any
concept of a halted plan or a blocked_by/runnable view yet.

* fix(#2830): a halted plan no longer leaves its dependents on the runnable work list

A plan that reaches a designed stop still writes a SUMMARY, so both
"which plans are incomplete" readers saw it as an ordinary completion and
reported its dependents as ordinary runnable work — never checking
whether an upstream plan had halted rather than finished.

- New `status: halted` frontmatter value, documented in all four SUMMARY
  templates alongside the existing `status: complete`.
- New shared src/plan-dependency-graph.cts: a single computeHaltPropagation
  pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and
  phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50
  dependent symbols across 5 command routers) now call, so the
  two-implementation divergence that caused this bug cannot recur. It
  accepts an optional precomputedOrder so cmdPhasePlanIndex — which already
  runs Kahn's algorithm in computeDependencyLevels for wave assignment —
  passes that order straight through instead of a second traversal;
  searchPhaseInDir (no prior traversal) lets the module derive its own.
  The two small duplicated predicates each reader would otherwise carry
  (is this status "halted"?, which summary file matches which plan id?)
  are centralized in the same module as isHaltedStatus/buildSummaryFileIndex.
- Additive fields only: `halted`/`blocked_by`/`runnable` on
  cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/
  `runnable_plans` on searchPhaseInDir's result. The pre-existing
  `incomplete`/`incomplete_plans` fields are unchanged in meaning and
  membership.
- execute-phase.md's discover_and_group_plans step now also skips any
  plan whose `blocked_by` is non-empty, reporting it by name with its
  blocking chain, in addition to (not instead of) the existing
  has_summary skip rule.

Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the
prior commit) with direct unit coverage of computeHaltPropagation
(including the precomputedOrder call shape) and a fast-check property
test — both only possible once this commit's new module exists.

Closes #2830

* fix(#2830): surface the halt-aware view from init execute-phase

The adopted work made phase-locator compute halted_plans / blocked_by /
runnable_plans, but cmdInitExecutePhase builds its output by explicitly
enumerating fields, so all three were computed and then silently dropped at
the exact consumer the issue names as regressed.

Forwards them additively -- incomplete_plans and incomplete_count keep their
name, type and semantics byte-for-byte -- and adds the same three empty
defaults to the roadmap-only fallback so the shape is consistent in both
branches. Covered by a new test that drives the real CLI end to end rather
than the locator function, since the locator already worked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect

Three review findings, all fixed:

- BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in
  the Kahn pass, so they were excluded from the topological order, never visited
  by the forward pass, and vanished from blocked_by entirely -- i.e. reported as
  runnable. The wave-grouping reader hard-fails on a cycle so it never hit this,
  but the phase-location reader does not, so init execute-phase offered a plan
  depending directly on a halted plan. Reproduced, then fixed in the shared
  engine so every consumer is safe regardless of pre-checks: a node absent from
  the order is now blocked with a deterministic, non-empty named cause. A plan
  silently missing from both blocked_by and runnable is the exact disappearance
  this issue exists to prevent.

- MAJOR (isolated adversarial). All four summary templates showed the field as
  an inline comment on the value line. Frontmatter parsing does not strip
  trailing comments, so an executor copying the templates' own presentation
  wrote a halt that parsed as a non-halted string, silently reproducing the
  original bug. Guidance moved off the value line, and the halt predicate now
  tolerates an unquoted trailing comment.

- HARD standards violation. A test regex-matched child-process stderr prose for
  /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal
  plus a differential assertion (same fixture without the cycle edge must
  succeed), so it stays cycle-specific without matching prose.

Also folds the duplicated read-summary-and-check-halted wrapper out of both
readers into the shared module -- centralizing only the predicate left the exact
two-copies-that-drift pattern the module exists to prevent -- and commits the
artifact-types documentation for the new status value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2830): stop the property generator hanging the whole suite

The remote runner did not fail -- it hung. Two containers sat in this file for
31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner
passes --test-timeout=0, so nothing ever reaps it: this would have hung CI
indefinitely, not reported a failure.

Root cause: the DAG generator built edges by rejection --

  from: fc.integer({ min: 0, max: n - 1 })
  to:   fc.integer({ min: 0, max: n - 1 })
  .filter(({ from, to }) => from < to)

With n === 1 both integers are forced to 0, so the predicate is unsatisfiable
and fast-check retries value generation forever. n is drawn from 1..12 and
fast-check biases toward boundary values, so n === 1 is reached almost at once.

This also explains why the failing-first run completed normally while the fixed
run hung: before the fix the graph module did not exist, so the property test
threw on import and never reached generation. It only starts hanging once the
code under test works.

Generates the DAG by construction instead -- `to` is drawn strictly above
`from`, with the degenerate single-node case short-circuited to an empty edge
list -- so no rejection sampling is involved. Switches the import to the shared
fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds
a bounded regression guard that samples the arbitrary directly, so a future
reintroduction fails loudly instead of hanging.

Verified: the file now completes in 2 seconds, 29 tests started and 29 finished,
zero failures, against an indefinite hang before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): restore the depends_on display contract and acknowledge the workflow growth

Full-suite run surfaced two things the focused harnesses could not.

1. Regression of a pinned pre-existing contract (#3785). A refactor routed the
   EMITTED depends_on field through the new dependency resolver, which also
   consults the canonical-prefix map. The original consulted the plan map only,
   so a short canonical prefix passed through verbatim -- '24-01' stayed
   '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a
   DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with
   a comment recording why it must not use the resolver; full resolution is
   still used for the wave DAG and halt propagation, which is what needs it.

2. The workflow file grew 518 bytes without an acknowledgment, from the
   halt-aware skip rule and the widened parse contract. Acknowledged.

Note on where the acknowledgment landed: the guidance is to add a NEW fragment,
but execute-phase.md is already named by an existing fragment and the linter
hard-fails when two ack sources name the same path. Appending to the owning
fragment, following its own established multi-PR pattern, was the only
lint-clean option.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2830): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 06:47:46 -04:00
Dennis Kim
d3ddcaba1c fix(#2785): implement missing gate predicate evaluators (#2816)
* fix(#2785): implement missing gate predicate evaluators

* fix(#2785): gate predicate numerical coercion

* fix(#2785): address evaluator review findings

* fix(#2785): use safe frontmatter read seam

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:16:56 -04:00
Tom Boucher
fd07e1a357 fix(#2850): resolve the active workstream in the statusline GSD-state segment (#3012)
* test(#2850): add failing-first tests for workstream statusline state

readGsdState only ever reads the flat .planning/STATE.md via a directory
walk-up; it has no path for .planning/workstreams/<ws>/STATE.md and never
consults GSD_WORKSTREAM or the stored active-workstream pointer, so the
GSD-state segment silently disappears in workstream mode. These tests
prove the RED before the fix lands.

Uses shared saveSessionEnv/restoreSessionEnv/clearSessionEnv helpers now
added to tests/helpers.cjs (single source of truth for the session-env-var
save/clear/restore pattern also used by tests/active-workstream-store.unit.test.cjs,
which is updated here to consume the same shared helpers instead of its own
local, already-diverged copy).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2850): resolve active workstream in the statusline

readGsdState only ever walked up looking for a flat .planning/STATE.md; it
had no branch for .planning/workstreams/<ws>/STATE.md and never consulted
GSD_WORKSTREAM or the stored active-workstream pointer, so the GSD-state
segment silently vanished in workstream-mode projects with no root
STATE.md (exit 0, no diagnostic).

Reuses the existing CLI>env>store resolution seam (resolveActiveWorkstream,
active-workstream-store.cts) and the existing mode-detection/path-building
seam (listAvailableWorkstreams/planningPaths, planning-workspace.cts)
rather than re-implementing either inline. When workstream mode is
detected but nothing resolves, readGsdState now returns a
{noActiveWorkstream:true} sentinel that formatGsdState/formatGsdStateCompact
render as "no active workstream" -- observable, never silent emptiness.
Flat-mode behavior and the case where a resolved workstream has no
STATE.md yet are both unchanged.

Adds active-workstream-store.cts's peekActiveWorkstream: a read-only
sibling of getActiveWorkstream. resolveActiveWorkstream's default store
lookup self-heals a stale/invalid pointer by deleting it
(adapter.clear()) -- correct for a command, but not for a renderer
invoked once per prompt, which must never mutate persistent, possibly
cross-session state as a side effect of drawing a screen. The statusline
now injects peekActiveWorkstream via resolveActiveWorkstream's own
getStored override, keeping the env>store precedence itself fully reused
while removing only the store tier's write side effect. This satisfies
the issue's AC4 ("the fix is purely additive to what's displayed").

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2850): backfill changeset PR number to 3012

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 21:44:04 -04:00
Tom Boucher
a987cf2731 chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init

Extends the init bundle with a typed per-invocation section manifest so an
invocation loads only the branch guidance it will actually take.

The three flag/state-gated branches in execute-phase.md move into their own
step files; the parent keeps its gsd:section markers wrapping a one-line
on-demand reference, so each section's prose lives in exactly one file and
the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator
derives the shipped section manifest from those markers, and a new pure
evaluator maps invocation facts to applicable section ids.

The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser
(Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary
and the predicate map stay exhaustively in sync.

Closes #2932

* fix(#2932): fail closed on prototype-chain when values

An isolated adversarial review found WHEN_PREDICATES[section.when] was a
bracket lookup on a plain-prototype object, so inherited Object.prototype
members resolved as predicates: "constructor"/"toString"/"valueOf"/
"hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and
"__proto__" threw an untyped TypeError carrying no .reason. Both violate
the module's documented fail-closed contract, and the manifest is read from
disk at run time so it cannot be assumed trustworthy.

Builds the predicate map on a null prototype and guards the lookup with an
explicit Object.hasOwn check. Adds table-driven coverage for nine
Object.prototype-shaped keys asserting the TYPED reason (asserting only
that it throws would still pass while broken) plus a fast-check property
injecting a hostile value at an arbitrary document position.

* test(#2932): retarget execute-phase step assertions at extracted step files

* fix(#2932): emit typed reasons for generator lib-load and write failures

* fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures

* chore(#2932): backfill changeset pr number to 2987

---------

Co-authored-by: sim <sim@local>
2026-08-02 12:34:41 -04:00
Adnan
137f3fbb9c fix(#2562): scope workstream progress/status to the current milestone (#2588)
* fix(#2562): scope workstream progress/status to the current milestone

`workstream progress` / `workstream status` / `workstream list` share one
derivation that could report a workstream's CURRENT milestone as
"milestone complete" / 100% while phases in that milestone were unstarted,
in progress, or failing verification. Three coupled defects:

1. The shipped signal was project-lifetime, not milestone-scoped:
   workstreamMilestoneShipped() returned true if ANY *-ROADMAP.md snapshot
   existed or "SHIPPED" appeared anywhere in ROADMAP.md. Every prior shipped
   milestone leaves a permanent collapsed <summary>✅ … SHIPPED</summary>
   block, so any post-v1.0 workstream was pinned to "milestone complete"
   forever (over-correction from #1913).
2. The denominator dropped declared-but-unscaffolded phases, and completed
   PRIOR-milestone phase directories inflated the numerator, letting
   progress_percent round to 100 while real work remained.
3. Phase completeness ignored the VERIFICATION verdict — SUMMARY >= PLAN
   count alone marked a phase complete even with a human_needed verdict.

Fix: derive both numerator and denominator from artifacts scoped to the
current milestone. The current version comes from the workstream STATE.md
`milestone:` field (ROADMAP in-progress markers can be stale); the ROADMAP
`## Progress` table maps every phase — including dirless ones — to its
milestone, and the matching set is both the denominator and the directory
membership filter. The shipped signal now requires the CURRENT version's
archived ROADMAP snapshot (REQUIREMENTS snapshots are not accepted; they can
be written at milestone start) or the current milestone's own line marked
shipped. Phases with an explicit failing verdict (gaps_found/human_needed)
count as in_progress; missing/unknown/stale are left untouched so
verifier-disabled projects do not regress to never-complete.

Greenfield roadmaps with no versioned Progress table, and projects whose
current version cannot be determined, keep the prior behaviour.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2562): add changeset for workstream milestone-scoping fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(#2562): parse the Progress table via findTableWithColumns

The ad-hoc pipe-table regex tripped the local/no-adhoc-markdown-parsing
ESLint rule. Use the canonical markdown-table helper instead: the
milestone-grouped RoadmapProgress variant is located by its required
`Phase` + `Milestone` columns and cells are addressed by column NAME,
so the parser tolerates column reordering and injected columns. The
`flat` variant (no Milestone column) yields no attribution, which is
the intended fallback to legacy counting.

Behaviour is unchanged: verified against a real multi-workstream project
(same status/percent/phase and plan counts before and after).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2562): count table-only phases in the unscoped denominator

Addresses the reporter's repro detail: a phase declared as a `## Progress`
table row with no `### Phase N` heading is missed by countRoadmapPhases
EVEN WHEN other headings exist — the heading regex counts 1 for a
"1 heading + 1 table-only" roadmap — not just in the zero-heading fallback
path. Milestone scoping did not cover this, because a flat Progress table
(no Milestone column) carries no per-phase attribution, so greenfield and
single-milestone projects kept the old heading-only denominator and the
declared phase silently vanished from it.

When milestone scoping cannot engage, the denominator is now the union of
the Progress table's declared phase numbers and the phase directories, so
neither source can shrink it. Verified against the reporter's minimal
fixture (phase 1: 1 PLAN + 1 SUMMARY + gaps_found; phase 2: table row only,
no heading, no dir), which now reports 0/2 at 0% across all four
table/STATE permutations instead of 1/1 at 100%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2562): attribute dir-only sub-phases to their parent's milestone

A sub-phase directory inserted mid-milestone (e.g. `30.1-…` under a
table-declared phase 30) usually has no ROADMAP Progress-table row, so it
had no milestone attribution and was scoped out of the rollup entirely —
its completed work was invisible and it could never hold the percentage
below 100.

It now inherits its parent phase's milestone and joins BOTH sides of the
calculation. Both sides is the load-bearing part: adding it to the
numerator alone would let completed_phases exceed a denominator that never
counted it, cap back to 100% via Math.min, and reintroduce exactly the
defect this issue reports. A regression test pins that failure mode (all
declared phases complete + an in-progress dir-only sub-phase → 75%, not
100%).

Attribution is deliberately one-directional: a sub-phase counts only when
its PARENT is in the current milestone, so a follow-up created in a later
milestone under an older parent is excluded rather than misattributed —
conservative (under-count) rather than falsely inflating.

Verified on a real project: the reported workstream moves from 2/6 (33%)
to 3/7 (43%), the 3/7 being the honest figure — a completed sub-phase that
was previously invisible now counts, and so does its plan total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2562): describe the denominator + sub-phase fixes in the changeset

The fragment was written at the first commit and only covered the three
original defects. Bring it up to date with what actually ships: the
table-only-phase denominator union (heading-only counting dropped a
declared phase even when other headings existed) and sub-phase milestone
inheritance across both sides of the calculation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(#2562): promote the canonical phase-key surface to phase-id

state.cts kept `phaseKeyFromToken`/`phaseKeyFromDir` private, so every other
module that had to compare two independently-derived phase references — a
ROADMAP table cell against a phase directory, say — wrote its own regex. That
is the defect class #2562 reports: a padded `01` and an unpadded `1-slug` land
in different key spaces and the comparison silently yields nothing.

Move the pair to the phase-id owner module and add `phaseKeyFromProse` (for
ROADMAP/STATE prose, markdown emphasis stripped) and `parentPhaseKey` (a
sub-phase's parent). state.cts imports them; its call sites are unchanged.

* fix(#2562): own milestone-shipped detection and accept a workstream scope

Three changes to the module that owns milestone parsing, so its consumers stop
reimplementing it:

- `isMilestoneShippedInRoadmap(content, version)` answers "does the ROADMAP mark
  THIS milestone shipped" from heading and `<summary>` lines only. A bullet that
  merely names the version (`- [x] 03-01: ship the v2.0 login endpoint`) is prose
  about a phase, not a milestone verdict. The version token is boundary-matched
  with `(?![\w.-])` — `\b` does not bound it, since `.` is a non-word character,
  so a shipped `v2.0.1` heading would otherwise close `v2.0`.
- `extractCurrentMilestone` and `getMilestonePhaseFilter` take an optional
  trailing workstream name and thread it to `planningDir(cwd, ws)`. A caller
  iterating workstreams cannot set `GSD_WORKSTREAM` per iteration, which is what
  the existing resolution falls back to. Omitted, resolution is unchanged.
- `getMilestonePhaseFilter` exposes `versionScoped`, true only when the phase set
  really is one milestone's. On an unversioned roadmap `phaseCount` spans the
  project's lifetime and must not be read as a current-milestone denominator.

The closed/active milestone-marker patterns were kept in three byte-identical
copies; they are hoisted to module scope as one `isClosedMilestoneHeading`.

* fix(#2562): derive membership and denominator from one phase-key space

The milestone scoping added earlier in this PR derived the ROADMAP table key and
the phase-directory key with two different regexes, and dropped rows it could not
attribute. Each of those was another way to reproduce the symptom this issue
reports — a rollup contradicting its own `phases[]` listing:

- a padded `| 01. … |` row never matched a `1-slug` directory (and a bespoke
  `^0*(\d+…)` never matched `PROJ-05-…` at all), so phases fell out of the
  milestone entirely and the percentage collapsed or pinned;
- a blank or malformed Milestone cell deleted the phase from BOTH sides, letting
  an unstarted phase vanish and the remainder round to 100%;
- shipped detection scanned bullets, so any checkmarked line naming the version
  closed the milestone;
- the numerator counted per-directory while the denominator counted distinct
  phases, so a stale same-numbered directory (Bug #2445's scenario) pushed
  `completed_phases` past the denominator, where `Math.min` capped it to 100%
  and hid the unstarted phase.

Both sides now key off the phase-id owner module (`phaseKeyFromDir` /
`phaseKeyFromProse`), directory membership additionally consults
`getMilestonePhaseFilter` when that filter is genuinely version-scoped, and the
denominator is the union of the roadmap's declarations with the member
directories' keys — so `completed_phases <= denominator` holds by construction.
The Builder asserts it and throws; the `Math.min` cap survives only on the legacy
unscoped path, where the denominator is a heading count that cannot bound the
numerator. An unattributable row degrades over-inclusively (kept, never dropped),
matching the degrade direction roadmap-parser already commits to.

* test(#2562): boundary coverage for each milestone-scoping reproduction

One test per way the scoping could still report "milestone complete"/100% while
phases are incomplete: zero-padded rows vs padded dirs (and the mirror),
project-code-prefixed dirs, a blank/malformed Milestone cell, a checkmarked
bullet naming the version, a shipped `v2.0.1` heading against a current `v2.0`,
and a stale same-numbered directory. Plus the current milestone's own shipped
heading (the signal must survive the boundary fix), the Builder's
numerator-above-denominator throw, a parity check that every non-`passed`
verifier status blocks completeness, and a guard that scoping reads the
workstream's ROADMAP rather than the project root's.

Reverting only `src/` reddens six of them.

* docs(#2562): record the milestone-scoped semantics and its consumer impact

CONTEXT.md: the Workstream Inventory Module's completion fields now describe the
current milestone, not the workstream's lifetime; phase-id owns the canonical
phase-key surface; roadmap-parser owns milestone shipped/active classification
and takes an optional workstream scope.

Changeset: name the behaviour change explicitly — `roadmap_phase_count`,
`completed_phases` and `progress_percent` change meaning with no schema signal,
and `getOtherActiveWorkstreamInventories` filters on the derived status, so
consumers see real movement.

* fix(#2562): collapse every zero-padding spelling to one phase key

A property test over the key surface — table cell and directory decorated
INDEPENDENTLY, which is the point — found a divergence neither review named:
`padStart(2, '0')` is a no-op once the input is already ≥2 characters, so `5`
normalised to `05` while `005` stayed `005`. A `| 5. … |` row and a `005-slug`
directory therefore never compared equal, which is the same failure mode as the
padded-vs-unpadded blocker, one level down.

The strip belongs in `phaseKeyFromToken`, not in `normalizePhaseName`: applying
it to the latter regressed multi-decimal leading-zero plan IDs (`001.10-PLAN.md`
capture + wave assignment), which rely on its verbatim rendering. Confining it
to the key surface fixes the comparison and leaves rendering untouched.

Also tightens `isMilestoneShippedInRoadmap`'s patterns to anchored,
complementary character classes so an untrusted ROADMAP cannot drive
backtracking, and makes the project-code test discriminating — it previously
passed pre-fix, because an unresolvable key collapsed scoping to the whole
roadmap and happened to land on the same number. It now carries a
prior-milestone directory that a collapse would wrongly admit.

* fix(#2562): prefer the milestone-attributing Progress table; pin the seams

Three gaps the earlier self-check missed:

- Both RoadmapProgress variants carry a `Plans Complete` column, so probing it
  first picked a FLAT table appearing earlier in the document over the
  milestone-grouped one that actually carries the attribution. Every row came
  back unattributed, was treated as current-milestone, and silently re-admitted
  prior-milestone phases. The attributing shape is probed first; flipping the
  order reddens the new test.
- `lint-phase-id-drift` exempts phase-id.cts by design, so it is silent on
  `phaseKeyFromToken`'s own segment strip by construction — not evidence. Its
  interaction with `stripProjectCodePrefix` (which runs AFTER) is pinned across
  project codes and hyphenated ids, including the pre-existing `M1-46-6` vs
  `M1-46-6-rs` asymmetry, which is `extractPhaseToken`'s #2043/#2232 slug-word
  rule and not something to "fix" by accident.
- `listWorkstreamInventories` loops every workstream with no try/catch, so a
  REACHABLE Builder-invariant throw would take down `workstream list`/`status`/
  `progress` for all of them. The invariant test only exercised the pure Builder
  with hand-built inputs. A test now drives `inspectWorkstream` over every
  adversarial shape at once (prior-milestone dirs, three colliding duplicates,
  a dirless declaration, an unattributed row, a project-code prefix, a dir-only
  sub-phase) and asserts it does not throw and the invariant holds — so the
  throw stays a contract assertion for external callers, not a runtime path.

Also covers the active-marker-wins rule (`## v2.0 — 🚧 IN PROGRESS … ✅` must not
mark shipped), which nothing exercised.

* fix(#2562): scope a declared-but-empty current milestone instead of falling back to history

The review's open MAJOR. `STATE.md`'s `milestone:` field updates the moment
`/gsd-new-milestone` writes the heading, while the `## Progress` table and phase
sections land later. In that window nothing attributes a phase to the current
milestone, `scoped` went false, and the fallback counted the project's ENTIRE
phase history as both numerator and denominator — a milestone with zero work
done reported 100% off its predecessors'. That is #2562's own symptom reached by
a different precondition, and none of the 16 tests covered it.

Reproduced first, four ROADMAP shapes, at `inspectWorkstream` rather than the
Builder — the Builder takes the scoping decision as an input, so a Builder-level
test proves it honours a flag, not that the derivation sets it. Three of the
four reported 2/2 100% with no phase of the current milestone begun.

Which signal witnesses the state depends on the ROADMAP's shape, and no single
one covers all three:

- `## v3.0` exists but declares no phases. `getMilestonePhaseFilter` DOES locate
  the section and sets `versionScoped`, then the zero-phase pass-all degrade
  resets it to false — erasing the only evidence the milestone exists. Neither
  existing flag survives that path, so this adds `versionSectionFound`, set
  beside `versionScoped` and deliberately preserved through the degrade.
- No section for this version at all, in a roadmap that versions its others —
  the existing `missingExplicitVersion`, already exposed and tested.
- Unversioned headings, but a Progress table attributing every row elsewhere:
  neither filter flag fires and the table is the only witness.

A ROADMAP that attributes NO versions anywhere matches none of them, which is
the point. Its rows parse with `version: null`, land in `currentMilestoneKeys`,
and never reach the new branch. `readCurrentMilestoneVersion` returns a non-null
version for very nearly every project (`getMilestoneInfo` defaults to `v1.0`),
so keying off `currentVersion` alone would have zeroed out every free-form
legacy project — the condition looks fussy for that reason. A test pins it.

Within an empty milestone, membership inverts: a directory belongs unless
another milestone's row claims it. Excluding everything would have dropped a
phase scaffolded before the roadmap caught up from BOTH sides of the rollup, and
hiding real work is the same class of defect as inventing it — this codebase
degrades over-inclusive, never under.

Scoping is now stated by the caller (`milestoneScoped`) rather than inferred
from `currentMilestonePhaseCount > 0`. That inference was the root cause: it
cannot represent a milestone that is scoped AND legitimately zero-phase, so the
Builder read "no phases yet" as "no scoping" and reopened the whole-history
path. The count-derived value stays the default for callers that say nothing.

A regression test also pins that a zero denominator does not trip the Builder's
`completed_phases <= denominator` throw, since `listWorkstreamInventories` has
no try/catch and a crash on every freshly-declared milestone would be worse than
a wrong percentage.

The changeset and CONTEXT.md no longer claim membership is derived in "ONE" /
"a SINGLE" phase-key space. `getMilestonePhaseFilter` still runs its own
`normalizePhaseIdSegments`; the signals are OR'd so a divergence can only widen
membership, but two normalisers coexist and the docs now say so.

* fix(#2562): cross-validate the shipped marker against the milestone's artifacts

`status: "milestone complete"` was asserted from the shipped marker alone, so a
single payload could report it beside `progress_percent: 67` — this issue's own
symptom, reached through `status` rather than the percentage.

The marker is now a claim checked against the milestone's own artifacts, and the
two signals are checked at DIFFERENT strengths because one check cannot serve
both. A `heading` marker (operator-typed, live ROADMAP) is refused on a short
completion ratio, which also catches phases declared but never scaffolded. A
`snapshot` marker is NOT ratio-gated: `milestone complete` moves the milestone's
phase dirs into `milestones/<version>-phases/` (milestone.cts:755-762) while
copying — never truncating — the live ROADMAP (:671-674), so a correctly
archived milestone reads 0/N by construction and a ratio gate would strip
`milestone complete` from every archived milestone. It is refused instead when
an in-milestone phase dir is still live and unfinished, reachable because
`milestone complete` does not advance STATE's `milestone:` field
(state-transition.cts:1335 vs :1224). `legacy` stays ungated — only reachable
when scoping is off.

A refused marker does not fall through to a STATE field claiming the same thing;
against contradicting artifacts neither source may report completion. The
refusal surfaces as `milestone_shipped_unverified` rather than staying silent,
distinct from `status_conflict` (derived-vs-field only).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* test(#2562): pin both marker strengths, the archived guard, and the owner modules

workstream-inventory: four tests, all four red against the prior src and green
with it. A live-ROADMAP SHIPPED heading over an incomplete milestone is refused;
an archived snapshot SURVIVES its phase dirs being moved out (the regression the
obvious single ratio-gate would cause — swapping the snapshot branch to that
gate reddens this AND the pre-existing `CURRENT-version snapshot marks the
milestone complete` at :321); an archived snapshot is refused once a phase is
reopened under it; and a refused marker is not re-asserted by a STATE field
claiming the same.

roadmap-parser: `isMilestoneShippedInRoadmap` gets unit coverage at its owner
module rather than only through the inventory that consumes it, plus two
characterisation tests for `getMilestonePhaseFilter`'s legacy call surface —
omitting the new trailing `ws` param is indistinguishable from `undefined`/`null`,
and the `GSD_WORKSTREAM` env fallback still resolves. These characterise the
call surface; they do not stand in for coverage of its individual callers.

phase-id: the `phaseKeyFrom*` / `parentPhaseKey` one-key-space contract, incl. a
property that padding a directory number never changes its key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* docs(#2562): record the two-strength shipped cross-check + milestone_shipped_unverified

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* fix(#2562): refuse an archived snapshot on a DIRTY archive, not just a live-unfinished dir

The snapshot arm checked `liveIncompletePhases > 0`, which misses the shape
@davesienkowski reproduced: a COMPLETE live dir beside a phase declared in the
Progress table with no directory. Nothing is live-and-unfinished, the marker
sails through, and `cmdWorkstreamProgress` returns
`{"status":"milestone complete","progress_percent":50}` — the reported symptom
verbatim, from one payload. Reproduced at 483a3ba30 before changing anything.

His diagnosis is the right one and better than mine: an in-milestone directory
outliving the archive means the archive is not CLEAN, and once that is true the
completion ratio is meaningful again. So the check is the conjunction — any live
in-milestone dir AND `completedPhases < effectivePhaseCount`. That strictly
subsumes the old predicate (an incomplete member dir is in the denominator and
not the numerator, so the ratio is always short when one exists) and leaves the
clean-archive guard green, since a clean archive has no live dirs at all.

Also corrects the module comment: the `scoped &&` prefix ungates all three
signals, not just `legacy`. That is correct behaviour — unscoped, the
denominator is the whole-roadmap count and membership is everything, so there is
no current-milestone artifact set to check a current-milestone claim against —
but the comment claimed otherwise. And the `milestone.cts` citations were ~28
lines stale after the rebase; they are now :700-702 (copy) and :783-790 (move),
re-verified against this head.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* fix(#2562): project milestone_shipped_unverified from list, status and progress

The inventory carried the field and every renderer dropped it — `workstream.cts`
was not in this PR's diff at all — so at the CLI a refused marker looked exactly
like no marker: a fallback `status` and nothing saying one was seen and rejected.
That is the silent collapse this issue is about, reintroduced one layer up, and
it made the changeset's "visible rather than silent" claim false at every
surface.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* test(#2562): pin the dirty-archive shape and the CLI projection

Five tests, all five red against the prior src and green with it.

The reviewer's repro at the builder: an archived snapshot with a COMPLETE live
dir beside a dirless declared phase must be refused, and status must not
contradict the percentage.

Four at the CLI via runGsdTools, the surface that was dropping the field rather
than the builder that already had it: `workstream progress`/`status`/`list` each
project `milestone_shipped_unverified: true` for that workstream, and a clean
archive still reports `false` with `status: "milestone complete"`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* docs(#2562): correct the snapshot check, the scoped-only caveat and the CLI claim

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 23:06:43 -04:00
0xdhx
c61dd49d95 enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes

Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks
(decision: 'block', exit 2) a whole-file Write collapsing a curated
.planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md,
STATE.md) below 40% of its on-disk line count. Files under 40 lines
are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message)
bypasses for legitimate milestone resets.

Fix 3 of #973 — the only defense independent of per-agent tool config.
Registered on the Claude plugin surface (hooks.json), settings-json
runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi
spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and
INVENTORY regenerated; regression tests negative-controlled (16/16
RED with the hook absent, 16/16 GREEN with it present).

* chore(#2255): backfill changeset pr number to 2301

* enhance(#2255): address review — fail-closed reads, typed block output, registration, property test

Review fixes for trek-e's CHANGES_REQUESTED on PR #2301:

- Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES
  (no-shipping-drift test).
- Blocker 3: update the always-on hook enumerations in ADR-766 and
  CONTEXT.md from six to seven.
- Major 4: fail CLOSED on non-ENOENT read errors — only a missing file
  (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a
  typed readError field and the override still honored. Tested, with a
  negative control against the pre-fix hook.
- Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES
  budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the
  floor; sub-floor always exempt), boundary examples pinned.
- Major 6: block output now carries typed oldLines/newLines/
  overrideEnvVar fields; tests assert on those instead of regexing the
  free-form reason string.
- Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS
  bypass on macOS/Windows); limit+1 boundary tests added for both the
  floor and the ratio.

* enhance(#2255): engage the write guard on Kimi's native payload shape

The guard shipped with Claude-vocabulary checks (tool_name 'Write',
tool_input.file_path), which #2304 showed leaves a guard dormant on
Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli
forwards its native payload verbatim — tool_name 'WriteFile' (bare or
module-qualified) and tool_input.path per its tool schemas
(src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown
name, and exited 0.

Apply the same per-guard normalization PR #2326 gives the three
sibling guards (name + field mapping, inlined — hook scripts stage as
standalone files), and write the block reason to stderr as well as
stdout JSON: Kimi feeds stderr, not stdout, back to the model on
exit 2, so a stdout-only reason blocks without telling the model why
or naming the documented override.

Regression tests pipe Kimi-shaped payloads (engage, qualified-name,
stderr-reason) plus exemption pins (StrReplaceFile stays out of scope
by design; non-curated paths pass) — verified red against the pre-fix
guard, green after.

* enhance(#2255): rebase onto next; regenerate golden-parity fixtures

* enhance(#2255): wire the escape hatch into complete-milestone's reorganize step

Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP
reorganize — the tree's only legitimate milestone reset and the exact caller
GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the
rewrite through a shell write with the hatch set on the command (a hook
inherits the runtime env, so a bare Write cannot carry a per-step override),
and a binding test derives the env var name from the guard's typed output and
asserts (a) the workflow step sets it and (b) the guard passes the identical
catastrophic payload under it — so the next complete-milestone.md edit cannot
silently re-break the wiring.

* enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload

Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string
reconstruction were unreachable-by-effect — the guard exits 0 for any
tool_name !== 'Write', so nothing ever read the fields they set, leaving
guaranteed-surviving mutants against the Stryker bar. The map now carries
only WriteFile -> 'Write'; the StrReplaceFile exemption test message states
the fall-through it actually exercises.

* enhance(#2255): review minors — American spellings; writeSync before exit(2)

Minor 1: normalised/normalise -> American house style. Minor 2: the two
block paths wrote stdout+stderr via async pipe writes then exit(2) —
async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block
payload durable.

* enhance(#2255): assert stderr equals the typed reason, not raw prose

Minor 3: the last raw-text match in the suite pinned override-name prose on
stderr. The contract is "stderr carries the reason Kimi feeds back" — now
asserted as stderr non-empty and byte-equal to the parsed stdout.reason.

* enhance(#2255): bind the write-guard's Kimi normalization into the parity test

Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with
nothing binding it. This extends PR #2326's kimi-guard-normalization-parity
test (same path and helpers, authored as a superset so either merge order
resolves cleanly): sibling byte-parity is existence-gated zero-or-all —
trivially green until #2326 lands, full-strength after — and the write-guard
copy is bound semantically (map is the value-inverse of convertKimiToolName;
the Kimi name for Write must map, or the guard is dormant on Kimi; the
path -> file_path half must be present). Byte-parity is deliberately not
asserted for this copy: it legitimately omits the Edit-class mapping
(Major 1 — dead code in a Write-only guard).

* enhance(#2255): refresh golden-parity fixtures for revised guard + workflow

* chore(#2255): regenerate golden fixtures after rebase onto next

The committed fixture hashes were generated against a tree predating
next's latest 11 commits, which independently modified the same
install-parity surface. Rebased onto next and regenerated with
`npm run gen:golden`.

Verified: against upstream/next the regenerated fixtures differ by
exactly this PR's own entries -- hooks/gsd-write-guard.js (new),
hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and
gsd-core/workflows/complete-milestone.md. No unrelated drift.

* fix(#2255): regenerate workflow size baseline for complete-milestone

`complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2
review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step,
but tests/workflow-size-baseline.json was never regenerated. The
per-file workflow baseline test (issue #1074) failed on
ubuntu-latest/22 and both macOS shard 1/3 jobs.

The growth is justified: it is the escape-hatch binding requested in
review round 2 (the guard must not hard-block the tree's only
legitimate milestone reset), not incidental bloat.

Regenerated via `npm run size:baseline`; the diff is exactly the one
entry.

* chore(#2255): regenerate golden fixtures and size baseline after rebase onto next

* enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes

Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert
(no PreToolUse hook exists on Bash in this family; the write succeeded by
dodging the guard, not by the override firing) and the protection was prose.
The hatch is now a transport code consults: complete-milestone's reorganize
step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the
Write tool as the sanctioned path, and the guard — at the block point only —
verifies the sentinel is fresh (15 min) and names the pending target, then
CONSUMES it and allows that one write. Path-bound + single-use + freshness
keep it from becoming a standing unlock. The env var remains as the
interactive transport, where it can actually reach the hook.

Regression tests written first (negative control: 3 failed pre-fix): the
armed-sentinel Write passes and consumes; stale does not exempt; a token for
a different file neither exempts nor is consumed; the binding test now takes
the sentinel name from the guard's typed output (overrideSentinel), asserts
the step arms it, and asserts the step no longer routes the rewrite around
Write via a shell pipe.

Also in this commit, same file:
- m2: block emission is exception-safe — emitBlock() wraps both writeSync
  sites in their own try/catch that still exits 2, so an EPIPE can no longer
  convert fail-closed into the outer catch's fail-open.
- Header discloses the two reviewed design limits (cumulative sequential
  shrink; lexical match vs symlinked paths) per round-5 scoping.

* docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4)

USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the
runtime-hooks-surface registration comment, and the changeset now describe
both hatches — the single-use sentinel for workflow steps and the env var
for interactive use — instead of implying a per-step env can reach a hook.
The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)`
parenthetical the repo's fragments use (round-5 m4).

* chore(#2255): regenerate derived families on the rebased tree (full sweep)

Full generator sweep after rebasing onto next @ the body-parser-patched
lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every
regen delta verified to be either a PR-owned entry (gsd-write-guard.js,
complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's
committed value for entries our arbitrary-side conflict resolution had left
stale (all 18 runtime fixtures checked mechanically).

* test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync

The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY
retry budget; the sentinel disarm now rides it like every other teardown.

* chore(#2255): regenerate derived families after rebase onto next

Full sweep on the rebased tree (build -> gen-inventory-manifest ->
gen:golden -> size:baseline). Every delta is either a PR-owned entry
(hooks/gsd-write-guard.js, its registration surfaces
hooks/managed-hooks-registry.cjs and the two plugin buses,
gsd-core/workflows/complete-milestone.md) or exact convergence to
next's committed value across all 18 runtime fixtures.

* chore(#2255): regenerate derived families after rebase onto next @ a5180d96

Rebase onto current `next` (a5180d96) resolved 12 conflicting
golden-install-parity fixtures; all regenerated via the full generator
sweep (build, gen:golden, size:baseline) rather than a single generator.

`lint:generated-sync` reports every generated artifact in sync. All 45
differing fixture keys and the single workflow-size-baseline entry map
to files this PR actually touches; no foreign drift.

* fix(#2255): remove the stale unguarded reorganize_roadmap step (round-8 blocker)

complete-milestone.md carried a second ROADMAP-collapsing step,
`reorganize_roadmap`, distinct from the sentinel-armed
`reorganize_roadmap_and_delete_originals` this PR wired. It is a vestige
of the pre-archive-then-reorganize design: it sits BEFORE
archive_milestone, so executing it as written would collapse ROADMAP.md
before the archive snapshots the full phase detail — and its Write is
exactly the shape gsd-write-guard hard-blocks, with no hatch armed. The
file's own success criteria describe only one reorganize outcome
(Backlog-preserving, overwrite-in-place — the later step's properties),
and archive_milestone points forward to "the reorganize step".

Removed rather than wired, per the round-8 review's confirm-and-remove
option. A new binding test asserts the sentinel-armed step is the ONLY
reorganize step in the workflow, so an unguarded collapse step cannot be
silently reintroduced (negative-controlled: fails against the pre-fix
tree). Golden-parity fixtures and the size baseline regenerate for the
shrunk file; every changed fixture key is complete-milestone.md's own.

* test(#2255): document why the read-error injection is a path collision, not an fs monkeypatch

Round-8 nit: the non-ENOENT tests inject via a directory-at-target-path
collision instead of the repo's fs-method monkeypatch pattern. That is
deliberate, not drift — runHook exercises the hook as a spawnSync child
process, so an in-process fs.readFileSync patch (the pattern the cited
siblings use on require'd, in-process code) can never reach the code
under test. Record the reasoning at the injection site.

* chore(#2255): regenerate derived families after rebase onto next @ 0d08c320

Rebase onto current next (0d08c320) for the CONFLICTING/DIRTY state. All 32
conflicts were generated artifacts (19 golden-install-parity, 12 install-tree,
workflow-size-baseline); resolved arbitrarily and regenerated via a full
generator sweep (build, gen:golden, size:baseline, gen-inventory-manifest)
rather than hand-merged. No source conflicts.

Regen diff verified against the PR's changed-file set: 7 distinct differing
keys, all PR-owned (gsd-write-guard.js, managed-hooks-registry.cjs,
plugins/gsd-core.js, complete-milestone.md, and their .kimi mirrors).
lint:generated-sync clean.

* chore(#2255): regenerate derived families after rebase onto next @ 9138271b

Conflict set was 20 paths, every one a generated artifact, zero source
conflicts — resolved arbitrarily during the replay and regenerated here,
per the maintainer's round-9 recipe (never hand-merged).

Generator sweep (not just gen:golden): npm run build, gen:golden,
size:baseline, gen-inventory-manifest, gen:registry. INVENTORY-MANIFEST
came back byte-identical, so the merged value was already correct.

Regen diff verified == PR-touched entries: every differing leaf key
attributes to a file this PR changes (complete-milestone.md, the guard,
managed-hooks-registry.cjs, both plugin buses) — derived from the
changed-file set, not a feature-name filter.

* fix(#2255): port the #2595 Kimi-payload hardening the base landed for the sibling guards

Self-found via the round-10 base sweep: a8b40fa5 (#2547/PR #2595) hardened
gsd-prompt/read/worktree-path-guard's normalizeKimiPayload after this PR was
authored, and this guard carried all three pre-fix shapes:

- `path` now wins outright over a model-supplied `file_path` — kimi-cli's
  WriteFile schema carries no `file_path`, so a payload pairing a curated
  `path` with a spurious `file_path: ""` had the guard reading '' and
  exiting 0 while kimi-cli wrote to `path` (one-key bypass, no crash needed).
- null/primitive payloads pass through deliberately instead of throwing into
  the fail-open catch.
- the `file_path` read is typed (`typeof === 'string'`), closing the
  crash-to-allow on `[]`/`{}` — this is what the base-added
  kimi-guard-typed-payload-reads scan flags.

Negative-controlled: the shadowing payload exits 0 against the pre-fix guard
and 2 against this one.

* fix(#2255): round 9 Major 2 — stop printing the sentinel recipe in the agent-facing denial

The block message ended by telling the blocked agent exactly how to arm the
single-use sentinel — handing the #973 actor a mechanical self-authorization
in the moment it is blocked. The recipe is gone from the prose; the
corrective advice and the env override's name stay (the latter is a #2255
acceptance criterion, and a per-step env prefix cannot reach a hook anyway),
and the typed overrideSentinel field stays for the binding tests. The hatch
remains documented in USER-GUIDE.md and complete-milestone.md, where humans
and the workflow engine read.

* fix(#2255): round 9 Minors 1-2 — realpath-resolve the target before the curated match; disclose the /i Linux cost

Minor 1: a Write to a non-curated path that symlinks into a curated file was
not matched while writeFileSync followed the link — the target is now
realpath-resolved before the curated match (ENOENT keeps the lexical
resolution so new-file Writes still pass; any other realpath error falls
through to the read, which fails closed). Negative-controlled: the symlink
payload exits 0 against the pre-fix guard, 2 against this one. Test skips on
win32, where symlink creation needs privilege.

Minor 2: the header's design-limits block now names the unconditional /i
cost on case-sensitive Linux (a genuinely distinct .planning/roadmap.md is
also treated as curated) next to the stateless limit, and drops the closed
symlink limit.

* test(#2255): round 9 Minors 3-4 — CRLF counting pin + a passing Write leaves a fresh sentinel unburned

Minor 3: countLines' split('\n') is CRLF-safe for a count (the \r rides
along), confirmed by trace in the review — this pins it against this repo's
recurring CRLF regressions, on both sides of the compare and at the 40%
boundary.

Minor 4: consumeSentinelFor runs only after the ratio check would block, so
a within-tolerance Write never burns the workflow's token — true by
construction, previously un-asserted.

* fix(#2255): round 9 Major 3 — correct the stale env-var line in archive_milestone's summary

complete-milestone.md's "After archival" bullet still said the reorganize
happens "under GSD_ALLOW_PLANNING_SHRINK=1" — the wording from the round-2
design this PR's own history rejected in round 5 (a per-step env var cannot
reach a hook; setting it in a Bash step silently does nothing). It now points
at the sentinel mechanics the reorganize step actually documents, matching
that step and USER-GUIDE.md.

* docs(#2255): round 9 Major 1 — user-facing docs state the stateless per-Write limit

The changeset and USER-GUIDE described the guard as covering "catastrophically
shrinks" with no caveat, while the stateless design was disclosed only in the
hook header — an operator reading the shipped docs would conclude iterative
erosion is covered. Both surfaces now state the per-Write comparison and the
erosion non-goal explicitly, in line with what the guard does.

* fix(#2255): round 10 Blocker 1 — arm the hatch in the roadmapper's curated writes

gsd-roadmapper Step 7 Writes both .planning/ROADMAP.md and .planning/STATE.md
wholesale, and /gsd:new-milestone spawns it against the OUTGOING milestone's
files. new-milestone's `phases.clear` archives phase DIRECTORIES, never
ROADMAP.md, so nothing compacts it first, and no ordering rule forces
/gsd:complete-milestone to run before /gsd:new-milestone.

Measured against the shipped hook at the #973 file size (292 lines): a new
4-phase roadmap lands at 18.2% and an 8-phase one at 31.8% — both blocked; only
a 12-phase replacement (45.5%) clears. So the collapse is reachable on a
first-party flow, which is the failure mode #2255 para 4 says the guard must not
have.

Arms the single-use sentinel immediately before each Write, one per target
(the token is path-bound and single-use, so one arming cannot cover both), and
gates each arming on `[ -f ]` so the /gsd:new-project path — where the guard
already exempts via ENOENT — strands no unconsumed token.

Binding test derives the sentinel name from the guard's own typed output, so a
rename on either side fails the suite instead of silently unwiring the hatch.
Negative-controlled: 2 assertions fail against the pre-fix roadmapper.

* docs(#2255): round 10 Major 2 — scope the claim to the guarantee that ships

The sentinel is a plain file, so an agent that would reason past an advisory
can arm one with a single Bash call it is already permitted to make. The
"only defense independent of per-agent tool config" framing outran that, and
the changeset was on its way into CHANGELOG.md.

Retitles the claim on all three surfaces (changeset, guard header, USER-GUIDE)
to what the guard actually delivers: it blocks accidental and single-shot
collapse and is not a defense against a determined agent; what it converts is
"ignore a sentence" into "take one deliberate, path-bound, single-use,
auditable action".

Pinned by test on the DURABLE surfaces only — the guard header and USER-GUIDE.
The changeset fragment is deliberately not pinned: it is consumed at release,
so a test reading it would start failing the moment the release lands. The
bound-statement assertion normalizes comment markers and whitespace first, so
it pins the claim rather than the paragraph's line wrapping.

Negative-controlled: both assertions fail against the pre-fix surfaces.

* test(#2255): acknowledge the roadmapper growth from the round 10 Blocker 1 wiring

The emitted-attribution gate (#2719/#2767) flags gsd-roadmapper.md growing 1130
bytes without an acknowledgment. The growth is the Blocker 1 sentinel wiring
plus the rationale a future editor needs to keep it, so it gets an ack fragment
rather than a silencing regen — the gate's own message is explicit that there is
nothing left to regenerate.

Fragment is PR-scoped (2301-…) per the gate's naming instruction, and uses the
plain-string reason form the shipped fragments use.

Verified against the TRUE upstream tip, not the fork's origin/next: a stale
origin made this same gate report unrelated phantom drift (1 emitted path + 6
grown files + 5 stale acks) that vanishes when GSD_EMITTED_BASE is pinned.

* test(#2255): renumber the roadmapper PROSE_ALLOWLIST pin after the Step 7 wiring

CI red on shard 2/3, all four platforms. The #2751 gate keys PROSE_ALLOWLIST on
{file, line}; the Blocker 1 wiring added 18 lines above the allowlisted
parenthetical in agents/gsd-roadmapper.md, moving it 624 -> 642. Both halves of
the gate then fired: the moved line reads as a new offender, and the stale
entry no longer matches anything.

Line content at 642 is byte-identical to what the entry describes — a
descriptive "e.g." naming SDK queries a user could run — so this is a
renumber, not a re-classification.

Swept the defect class rather than the instance: agents/gsd-roadmapper.md is
the only line-pinned reference to any file this round changed.

Negative-controlled: both assertions fail against the un-renumbered allowlist.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:19:49 -04:00
0xdhx
cc3ee301a7 fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root

installSharedHooksBundle wrote `{"type":"commonjs"}` over
<configRoot>/package.json unconditionally — no existence check, no merge,
no backup — on every install and every /gsd-update re-install. On the 11
affected runtimes that file is often user-owned; on OpenCode and Kilo it is
the documented place to declare local-plugin npm dependencies, so a user's
name/type/dependencies/scripts were destroyed on each run.

The uninstall path already read the file and unlinked it only on an exact
content match. That asymmetry was the defect: the discipline existed in the
codebase, it just was not applied on the write side.

Move the marker into the directories GSD creates and fills with its own .js
files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the
nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop
writing the config root entirely. New src/commonjs-marker.cts owns the marker
string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed
on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker,
so install and uninstall cannot drift apart again.

Nothing else depended on the config-root marker: package identity is baked at
build time (#378/#498) and version resolution prefers gsd-core/VERSION and
already tolerates a missing root package.json (#1383) — Codex has installed
without one all along. A package.json in plugins/ or extensions/ is inert to
plugin discovery, which globs *.{ts,js} only (see installer-migration 006).

Uninstall retires the pre-fix config-root marker, so upgrading users are
cleaned up on removal, and still never touches a file it did not write.

* fix(#2544): point the changeset fragment at the filed PR

The fragment's `pr:` field is only knowable after `gh pr create` returns.

* fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set

bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the
linted source), so it belongs in the ADR-457 ignore list like its siblings.
Clears the lint-tests no-var failure and the repo-invariants
"linted xor ignored" migration-state test.

* fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root

The UPGRADE 1 test still asserted the pre-#2544 marker location
(~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the
directory GSD itself creates — matching the updated golden-install-parity
and install-tree fixtures. Also asserts the root marker is NOT written.

* fix(#2544): make the CommonJS marker write path non-fatal

Review round 2, Major 3 + Minor 1 + the stagedHooks nit.

ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site
caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole
install with a raw stack trace. Every other marker interaction in the module is
best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker
swallows read failures — and this was the write path, i.e. the one most likely
to fail on a locked-down config dir. It now returns a new 'failed' outcome and
both call sites warn and continue.

Sibling found while sweeping for the same defect class: fs.mkdirSync sat
OUTSIDE the try block, so an unwritable parent threw past the guard entirely.
Creating the directory is the same environmental hazard as writing into it, so
it moved inside.

Also in this file:

- The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks
  alone. stagedHooks is computed from the SOURCE listing before the copy loop,
  so it stays true when the copies land but verifyInstalled() then fails —
  marking a hooks/ GSD did not successfully populate claims an ownership the
  install did not earn.
- The uninstall rmdir of the native plugin dir is gated on GSD having actually
  removed something from it. Hoisting it out of the adapter-exists guard (so
  the marker-only case could prune) had silently widened it into deleting a
  user-created but empty plugins/ or extensions/ dir — the same "don't touch
  territory GSD didn't fill" principle this issue is about, inverted.
- Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the
  same call site that writes its replacement. That path is outside kimi's
  configDir, so installer-migration 007 structurally cannot reach it.

* fix(#2544): retire the stale config-root marker via installer-migration 007

Review round 2, Major 1 — the PR's headline claim was false for existing
installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale
{"type":"commonjs"} at the config root, so their config root stayed pinned to
CommonJS and their dependency manifest stayed gone until they uninstalled.

The migration is unusual in one way, and it is the part worth reviewing: the
config-root marker was never recorded in gsd-file-manifest.json (writeManifest
records hooks/, agents/, commands/, scripts/ and the native plugin, never a root
package.json), so classifyArtifact answers 'unknown' for it and the planner's
own guard downgrades a remove-managed on an 'unknown' classification to
preserve-user. 007 therefore supplies the "purpose-built detector for an old
GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions —
exact content match, the same predicate removeCommonJsMarker has always used —
and declares the resulting classification on the action. A package.json with any
other content is left untouched, and there is deliberately no backup-and-remove
branch: a non-matching file here is not a patched GSD artifact, it is somebody
else's file.

Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`:
validateStringArray requires the field to be non-empty WHEN PRESENT, while the
runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan
time and the migration never runs. The metadata test pins this.

Kimi is a deliberate carve-out, named in the migration's own header: its marker
lived at ~/.kimi, outside kimi's configDir, and migration relPaths are
structurally confined to configDir. It is retired by the installer instead.

Registration: shipped-migrations table, .gitignore for the emitted .cjs, the
EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not
copied from migration 006 by rote — 006 needs no entry because it imports
nothing, while 007 imports node builtins, so tsc emits its __importDefault
helper and the `var` in it trips no-var. This is the same lint gate that made
round 1 red.

* test(#2544): fault-injection and multi-runtime marker coverage

Review round 2, Major 2 + Minors 4 and 5.

Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and
the suite had no fs monkeypatching at all. Every branch now covered is one whose
doc comment claims it as the module's safety posture:

- classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule),
  with an ENOENT control alongside it so the test discriminates rather than
  just asserting one side
- classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never
  downgrades to the permissive answer) — the fixture's bytes are exactly GSD's
  marker, so the test fails if the code ever answers on content it could not read
- a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already
  covered with a real symlink, the directory case needs no injection at all)
- the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx'
- the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the
  mkdirSync that used to sit outside the guard
- removeCommonJsMarker unlink throw -> false

These save and restore fs methods in `finally` rather than using chmod 0o000,
which does not fault under root and would pass vacuously in root Docker and CI.

Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both
kimi locations now have behavioral coverage, install and uninstall, each paired
with a user-authored-file case proving GSD leaves it alone.

Minor 5 — the stagedHooks gate had no assertion behind its stated reason.
A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime
that declares skipSharedHooksInstall and asserted to stay marker-free, with its
user content intact.

Also regression-tests the uninstall rmdir gate from the previous commit: an
empty plugin dir GSD removed nothing from must survive.

* docs(#2544): correct stale marker prose, register the module, document the trade-off

Review round 2, Minors 2, 3 and 6.

Minor 2 — six files asserted the installed ROOT ships the synthetic marker.
None was load-bearing (all three walk-up consumers are VERSION-first with
try/catch and the marker never carried a `version`), but ADR-457:52 is the
rationale for keeping a generated module, so a future reader would mis-derive
the constraint from it. Each site is corrected to what is now true: the
installed tree carries no package.json with a .name at all, because the only
ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own
directories.

Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js
and the platform-gate test both described `require('../package.json').name`
resolving to undefined; post-#2544 that require does not resolve at all, so the
history is kept accurate and the present-tense claim corrected rather than just
moved. And src/runtime-artifact-conversion.cts described the no-root-package.json
case as Codex-only — it is now every runtime, which strengthens that comment's
own argument for lazy resolution. The generated .cjs sibling needs no edit: it
is gitignored build output, not a tracked file.

Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer
module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added,
including the fail-closed posture and the never-throws contract.

Minor 6 — the plugins//extensions/ marker shadows the config root for all .js
siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is
exactly what #2544's Fix section prescribed and it is disclosed in the PR body,
but the PR body is not documentation. It now lives in the OpenCode section of
docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than
a pure improvement, with the .ts mitigation and a fallback for ESM plugins.

* test(#2544): attribute the CommonJS marker in the emitted-provenance rules

The differential emitted-attribution gate (#2723, landed on `next` after this
branch was cut) went red on the macOS shards once this PR rebased onto it. Two
distinct causes, both real gaps rather than noise:

1. `plugins/package.json` and `extensions/package.json` matched NO rule — the
   `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an
   unattributed emitted family.
2. `hooks/package.json` fell through to `hooks-built`, which attributes an
   emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no
   `hooks/package.json` in the repo, so it resolved to a nonexistent path.

Cause 2 is exactly the failure already documented three lines above it for
Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix
follows that precedent rather than inventing one: `package.json` is excluded
from `hooks-built` the same way, and a dedicated `commonjs-marker` rule
attributes the family across all four roots it can appear in (both hooks roots
plus `plugins`/`extensions`) to the sources that actually emit it.

Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for
a one-off ripple and goes stale by design — the gate fails a stale ack precisely
so it cannot pre-clear the next change on that path. These markers are a
permanent part of the emitted tree from #2544 onward, so they need standing
attribution.

Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3
provenance errors + 12 unattributed paths before, 35/35 green after.

* fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker

#2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned
in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via
dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's
module on the two properties that matter:

  - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a
    DANGLING one, so a dangling package.json symlink classified as absent and
    the write went straight through it. Demonstrated: against the pre-fix copy,
    ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json
    symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory.
  - create: a plain writeFileSync leaves the classify->write window open, where
    commonjs-marker creates with flag:'wx' (O_EXCL).

Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's
own docstring already claimed was the single place these rules are enforced.
Exported signatures are unchanged (still boolean), so bin/install.js and the
#2717 tests are unaffected.

The new subtest is the only coverage that fails if the duplicate is ever
reintroduced — the two implementations agree on every non-adversarial input, so
the existing suites pass against both.

* test(#2544): pin the stagedHooks gate on zcode, not windsurf

The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall
kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the
marker was correctly absent.

#2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via
dedicated paths and get the marker beside those scripts. Measured on this tree,
windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning
behaviour that is now wrong, not the gate it was written for.

ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin
surface to spawn hooks, so GSD stages no .js there by either route (measured: 0
staged, no marker). The property under test is unchanged — a user-created hooks/
directory GSD never fills stays marker-free.

* test(#2544): use the shared cleanup helper in the migration test

Addresses the review's Major 1. The suppression's stated reason — "no helpers
import available" — was not correct: tests/helpers.cjs exports cleanup, and the
other test file added in this same PR imports it (tests/commonjs-marker.test.cjs).

The local reimplementation dropped two protections that are live on this repo's
windows-latest lane: the CWD guard (Windows cannot remove a directory that is the
current working directory) and the 20 x 250ms retry budget that absorbs the
deferred-scan handle Windows Defender holds on newly-written files.

Local function and suppression both removed; local/no-raw-rmsync-in-tests now
passes without one.

* test(#2544): expect hooks/package.json for the #2717 runtimes

The fresh-install contract table predates #2717, which stages cursor/windsurf/
codex .js hooks via dedicated paths and writes the CommonJS marker beside them.
All three therefore now receive hooks/package.json legitimately.

Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with
the marker; cline/copilot/trae/zcode stage none and get none, so their contracts
are unchanged.

* fix(#2544): gate the #2717 marker writes on having staged something

The three dedicated marker writers #2717 added ran unconditionally. Each one
mkdirs hooks/ up front and stages its scripts conditionally on the source
existing, so with an absent or empty hook source they created a directory,
filled it with nothing, and marked it as GSD's anyway.

That is the same write-into-someone-else's-territory this issue is about, and
installSharedHooksBundle already guards the identical case with `stagedHooks`.
The dedicated paths now carry the matching gate:

  - cursor / windsurf: `installedScripts.size > 0`
  - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that
    hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY
    entry landed.

Covered for cursor and windsurf by driving each writer against a src tree whose
hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its
trigger state needs a package tree where hooks/dist exists but holds none of the
allowlist, which is not constructible from a real checkout.

* test(#2544): scope the commonjs-marker sources per root

The rule declared one flat source list for every marker root, so
`extensions/package.json` was attributed to runtime-hooks-surface.cts (which
never writes there) and `.kimi/hooks/package.json` to install-engine.cts.

That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source,
so a flat list containing bin/install.js let any change anywhere in that
13k-line file authorise marker drift for every root — the blanket escape hatch
this file's own agents-verbatim comment refuses for exactly the same reason.

Sources are now derived per root from ctx.rel. Note the rule ctx is
`{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent
every path down one branch silently.

* test(#2544): state precisely what the zcode assertion pins

The comment claimed the test pinned installSharedHooksBundle's `stagedHooks`
gate. It does not, and neither did the windsurf version it replaced: zcode
declares skipSharedHooksInstall, so the outer guard skips that helper entirely
and the gate is never evaluated. The test passes on the runtime exclusion.

What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays
marker-free — is still worth having, and is what the review asked for. The two
`staging zero hook scripts` tests are the ones that pin a real staged-nothing
gate. Comment corrected rather than left implying coverage that is not there.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:00:23 -04:00
Tom Boucher
640eaee16e chore(#2930): fragmentize execute-phase.md and prove per-runtime composed emission (#2972)
* feat(#2930): fragmentize plan-phase.md workflow into per-runtime-composed sections

Adds src/workflow-fragments.cts (in-file <!-- gsd:section --> marker
parser/composer, ADR-1671 epic #1671 Phase 3), wires it into
bin/install.js's copyWithPathReplacement emission path, and pilots the
marker grammar on gsd-core/workflows/plan-phase.md.

Bookkeeping ripple for the new src/*.cts module: .gitignore,
eslint.config.mjs, docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json,
and a CONTEXT.md glossary entry. Amends ADR-1671 with open questions 1
and 2 resolutions and records the closed when= applicability grammar.
Adds docs/reference/workflow-fragments.md and an ARCHITECTURE.md
section documenting the marker authoring model.

* fix(#2930): put allow-test-rule issue ref on the same line as the marker

lint-allow-test-rule-refs.cjs requires the #NNN issue reference on the
same source line as `allow-test-rule:`; it was one line below and read
as an unreferenced novel exemption.

* docs(#2930): link the orphaned gate-predicates reference from the docs index

Found while adding the workflow-fragments reference doc: docs/reference/gate-predicates.md
shipped without an entry in docs/README.md, so it was unreachable from the docs index.
Fixed inline rather than deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): scope composition to workflows, add typed failure reasons

Review findings from two orthogonal passes:

- Scope composeWorkflow to gsd-core/workflows/ only. It previously ran on
  every .md the installer copied, so a future agent/command/reference doc
  documenting the marker syntax with an unfenced example would have been
  mis-parsed and silently stripped — a lossy drop the phase forbids.
- Add a frozen REASON enum; failures attach a typed .reason and tests assert
  on it instead of matching free-form message text (CONTRIBUTING.md:635-694).
- Derive the property generator's when= values from WHEN_VOCABULARY instead
  of duplicating them (DEFECT.GENERATIVE-FIX).
- Add adversarial parser fixtures: Unicode headings, NUL, U+FFFD, BOM,
  fence-within-fence, tilde and indented fences, lone-CR marker line.
- Document why --mvp is structurally unmarkable: its content is interleaved,
  not sectioned, so the whole-line grammar cannot reach it.

Also fixes two stale tests on this branch, each reproduced on the unmodified
tree before correction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): retarget the pilot from plan-phase to execute-phase

The full remote matrix went red on both Linux lanes. Root cause was ours:
tests/phase6-capstone-conformance.test.cjs holds a PRE_PHASE6 ceiling of
94519 bytes for plan-phase.md, asserting an ADR-857 Phase-6 completion
property. That is a third size gate beyond the tier caps and the
differential ratchet, and it left plan-phase.md just 36 bytes of headroom
rather than the 3821 computed from the XL cap. The 330 marker bytes
overran it by 294.

Raising the ceiling is not an option: it is a red line certifying another
ADR's completion. plan-phase.md is reverted to byte-identical origin/next
and the pilot moves to execute-phase.md, which has 728 bytes of headroom
under its own ceiling and lands at 93147 with 3 marker pairs.

The vocabulary narrows to the atoms actually used: always, flag:--wave,
state:gap-closure-phase, state:has-prior-phases.

Recorded in the ADR: every branch the epic names lives in plan-phase.md,
which cannot be fragmentized until caps move from source to emitted bytes.
That is direct evidence for the epic's premise and may reorder phases 3-4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2930): backfill changeset PR number (#2972)

* fix(#2930): make the emission install tests portable on Windows

The windows-latest lane went red on three tests in the new install suite;
Linux was green. Both causes were in the test harness, not the module.

Root normalization: the opencode converter always embeds the install root
forward-slashed, but the tests stripped it with the native-separator string
from mkdtemp. On Windows that never matched, so the root leaked through
unstripped — and because the real and stub install roots have different
prefix lengths, that length difference landed directly in the byte-delta
assertion (344 observed vs 275 expected). Normalize both text and root to
one separator form before stripping.

@-ref resolution: the helper stripped only the @~/ and @$HOME/ forms, so a
Windows absolute ref (@C:/Users/...) fell through and was joined onto the
root, producing ...\@C:\Users\... Strip the @ first, then detect
absoluteness from the token's own shape (POSIX, drive-letter, or UNC) with
no platform branching, so every OS takes the same path.

Neither assertion was weakened; the exact-equality byte check is the point
of the test and still holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2930): document every REASON member and guard the doc/enum parity

Code review found the reference doc's 'Fails closed' list covering 10 of the
11 frozen REASON members — MALFORMED_ATTRIBUTES (parseAttrs rejects malformed
key="value" syntax) had no bullet, and it is distinct from
UNRECOGNIZED_ATTRIBUTE, which is valid syntax with an unknown key.

Two parallel surfaces sharing one constant with nothing asserting they agree is
the DEFECT.GENERATIVE-FIX class, so the same commit adds the parity assertion:
the test derives the enum side from the built module and the doc side by parsing
the reference page, keyed on the reason IDENTIFIER rather than prose so a
reworded bullet does not break it, and reports set differences in both
directions by name.

Proven non-vacuous: removing the MALFORMED_ATTRIBUTES bullet turns the suite
red naming that exact member; restoring it returns 44/44.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 12:12:20 -04:00
Tom Boucher
9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
90771ddf02 enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable

ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.

Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.

The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.

Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.

Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.

Closes #2904

* fix(#2904): bound and sanitize untrusted registry `interactions` strings

Review findings from the pre-PR passes.

Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.

`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.

Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.

Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.

* chore(#2904): backfill changeset PR number (#2912)
2026-07-31 08:11:57 -04:00
Tom Boucher
4bd6fb066b chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam

The no-adhoc-markdown-parsing rule matched only a negated class whose sole
member was a pipe ([^|]), so the stricter and more common [^|\n] spelling
evaded it entirely -- src/state-document.cts hand-rolled exactly that shape
and linted clean. Widen the fingerprint to any negated class excluding a
pipe, which is the ADR-2143 section 7 prohibition as written.

With the rule fixed, state-document.cts goes red. Replace tableRowPattern
with locateFieldRow: a line scan using the markdown-table seam's
splitTableRow for cell semantics, returning the value cell's byte range, and
splice that range instead of running a whole-document content.replace. An
edit now physically cannot cross a row boundary (section 4).

Behavior is frozen -- stateReplaceField has 79 dependents across 5 command
processes. Characterization tests lock all 14 table-branch rows plus CRLF,
extract round-trip and the withFallback caller shape; a fast-check property
asserts every non-target line stays byte-identical.

Refs #2880, epic #2143

* fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint

Isolated adversarial review found four defects in the first commit.

1. locateFieldRow split lines on \n only. JS treats a lone \r as a line
   terminator, so the regex it replaced matched rows separated by bare CR.
   "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now
   CR, LF and CRLF are all terminators, byte offsets unchanged.

2. The field name was normalised with trim().toLowerCase(). The old regex
   embedded it verbatim, so its whitespace had to be absorbed by the row's
   own padding -- and because the group is a literal-character match rather
   than a whitespace class, a tab-padded cell does not accept a
   space-padded name. Replaced with an offset-aligned search reproducing
   the original backtracking exactly.

3. The widened fingerprint regex had two unbounded [^\]]* around an
   optional and ran quadratically over every regex source in every linted
   file: 256000 chars took 23 seconds. Replaced with a single-pass scanner
   that never rescans; the same input is now ~1ms.

4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|].
   Narrowed to a class excluding the pipe plus only \n, \r or \t.

Differential fuzz against origin/next: 20000 cases, 0 mismatches.

Refs #2880, epic #2143

* test(#2880): drop wall-clock assertion from the ReDoS regression guard

local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md
bans timing assertions outright as flaky. The 256000-char input stays as
the regression guard for the quadratic scan; correctness of the verdict is
what is asserted. If the quadratic path returns, the test stops completing
and surfaces as a suite timeout rather than a silent pass.

Also adds the changeset fragment for #2880.

Refs #2880

* fix(#2880): spec-correct case folding, property tests, naming

Code-review findings.

The field-name comparison used toLowerCase(). The regex it replaced used
/i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a
non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII
K where the old code returned null. Replaced with spec-correct
Canonicalize, including the multi-character uppercase case (eszett -> SS),
which a naive uppercase comparison also gets wrong.

Added the fast-check property tests CLAUDE.md requires for parsers: one
for the negated-class scanner, one for the field-name fold semantics, each
against an independent reference implementation. Both reference impls
failed on first run against real bugs, so neither property is vacuous.

Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a
duplicated comment to a cross-reference.

Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the
harness proven to discriminate the KELVIN case.

Refs #2880

* chore(#2880): backfill changeset PR number (#2889)

* docs(#2890): correct the local ESLint plugin path in CONTEXT.md

CONTEXT.md named the local AST-rule plugin directory as
scripts/eslint-rules/, which does not exist. The real location is
eslint-rules/ at the repo root -- what eslint.config.mjs actually
imports -- and CONTEXT.md's own later entry already says so
explicitly, so the file disagreed with itself.

Found by a line-by-line audit of all 1036 lines against the live
graph; this was the only confirmed inaccuracy.

Closes #2890

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:55:03 -04:00
Tom Boucher
7372d99a26 enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales

The reviewer lane roster was hand-enumerated across five documentation
surfaces and three workflow files that had drifted apart: --kimi-code was
missing from all four translated COMMANDS.md mirrors, --coderabbit from
every workflow forwarding list, and --antigravity from FEATURES.md.

Adds checkReviewerDocsParity, a second pure gate deliberately separate from
checkReviewerLaneParity so a stale doc cannot make the runtime checker look
red. Workflows now derive their flag lists from a new review-lane flags
query instead of hand-enumerating them, which also retires the unanchored
grep that matched --agy inside --antigravity.

Documents the previously absent reviewer body and hostBehaviors field in
the capability manifest reference.

Closes #2800
Closes #2781
Closes #2272

* fix(#2800): key the docs parity table arm on first-cell position

Review found the flag arm was file-scoped, so the forwarding row that lists
every flag in its third cell satisfied it on its own. Deleting a lane's own
reviewer-table row -- the #2781 regression this gate exists to prevent --
therefore passed undetected.

Arm 4 keys on the FIRST table cell, which separates a lane row from the
forwarding row structurally and in every locale. Regression test included.

* fix(#2800): shape-filter the flags subcommand output

All three consumers read review-lane flags through an unquoted command
substitution so the output word-splits into loop items. Phase 2 admits
third-party overlay lanes, so an overlay flag containing whitespace would
inject a second loop item and one containing a glob would expand against
the cwd. Emit only well-formed flags so neither reaches the shell.

* fix(#2800): remove the regex length ceiling and count only prose mentions

Review found two real defects in the docs parity gate.

The never-throws contract was false: building a RegExp from a declared flag
or section title throws SyntaxError past ~100k chars, and Phase 2 admits
overlay lanes whose declared strings are untrusted in length. Every one of
these matches is literal, so String.includes replaces the regex outright,
which also deletes escapeLiteral and the llama.cpp escaping it existed for.

Arm 1 was context-blind: a flag mentioned only inside a fenced example or a
commented-out row counted as documented. Both are stripped before matching.

Also advertises all 13 lane flags in the argument-hint and corrects a stale
eleven-lane count in the slug grammar note.

* test(#2800): repoint the convergence suite off deleted workflow text

The derived flag loop deleted the literal per-flag grep lines four tests
matched on. Two of those failed loudly. The behavioral and property tests
failed SILENTLY instead: their end marker no longer resolved, so the parse
block extracted empty and both passed vacuously, and the property test's
gsd_run stub had a no-op default that hid it.

All now share one extractor and execute the real deployed block through a
gsd_run shim backed by the actual binary. The whitelist assertions become an
anti-parity check: re-adding a hand-written flag list must fail.

Also repairs two vacuous cases in the docs parity suite. The unreadable-doc
test called its own mock rather than the reader, and the integration test
bounded nothing, so a doc losing its marker would have been silently skipped
and still passed green.

* fix(#2800): run the derived flag loop after the launcher preamble

The remote matrix caught a real runtime bug, not a test artifact. In
autonomous.md and plan-review-convergence.md the launcher preamble that
defines gsd_run lives in a separate, LATER bash fence than the derived loop.
Each fence is its own shell, so gsd_run was undefined where the loop ran:
the command substitution yielded nothing and zero reviewer flags would have
been forwarded. Worse than the drift this epic fixes, and silent.

The whole CONVERGENCE_ARGS construction moves as one unit, because the
--max-cycles append sits between the loop and the preamble and would
otherwise have run against an uninitialized variable and then been dropped
by the relocated initializer.

Also documents all 13 lane flags in help/modes/full.md, which the repo gates
bidirectionally against each command's argument-hint.

* test(#2800): repoint the two converge suites off deleted flag literals

Both asserted workflow.includes('--codex') against the hand-enumerated list
the derived loop removed. They now assert the derivation itself, keep --all
and --text (convergence controls, still literal), and add an anti-parity
guard so re-adding a hardcoded list fails.

The lost pass-through proof is replaced with a real one: every flag the
tests used to hardcode is asserted present in the actual roster emitted by
the binary, which is the property the old assertion was protecting.

* test(#2800): acknowledge the workflow byte growth from the derived flag loop

* chore(#2800): backfill changeset pr number to 2882

* fix(#2800): strip HTML comments to a fixed point in the parity gate

CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the
single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick
construction smuggles a commented-out row past the gate and it counts as
documented. Not an injection risk here since nothing is rendered, but it is
the exact false pass this helper exists to prevent.

Strips to a fixed point, then treats any surviving opener as unterminated so
the multi-line branch closes it on a later line. Terminates because every
pass strictly shortens the string.

* test(#2800): pin the comment-smuggling regression with a real reproducer

The obvious fixture for this class does not reproduce it: <!--<!---->-->
leaves a dangling --> rather than a live <!--, and is caught either way, so
it would have passed with and without the fix. The join-trick construction
(<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely
regresses on the single-pass strip and is what the test now uses.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:14:13 -04:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00
Tom Boucher
8b44a0da43 chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract

Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as
the declared contract for all 11 cross-AI reviewer lanes, and the
DEFECT.GENERATIVE-FIX parity assertion the roster has never had.

The lane contract lived in three unrelated surfaces — the roster, ~640
lines of hand-authored per-CLI bash in invoke_reviewers, and the
write_reviews section headings — so cross-cutting fixes landed per-leg
(#2494 and #2605 were the same empty-output defect filed twice).

- src/review-lane-descriptor.cts: frozen table declaring per lane the
  slug, flags, probe, invoke shape, timeout floor, empty-output policy,
  REVIEWS.md section, evidence class, required binaries, prompt-budget
  key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so
  Phase 2 harvests the shape with no translation layer. It declares;
  it does not execute — invoke_reviewers iterates in Phase 5b.
- checkReviewerLaneParity: bidirectional parity across descriptor,
  roster, invoke_reviewers legs and write_reviews sections. Forward-only
  would miss the failure it exists to catch (#2718 added a leg, #2781
  was the drift). ADR-1517 instance headings are exempt per D8.
- Legs carry an explicit <!-- reviewer-lane: slug --> marker; five
  non-lane bold labels share the bold-then-fence shape a heuristic
  matcher would key on.
- ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an
  error in both the core module and the workflow prose that mirrors it.
  A code-only change would be unobservable — the module has no
  production caller; the workflow narrates the policy. Discovery paths
  (--all, review.default_reviewers) stay lenient.
- Fixes the qwen leg, the last one discarding stderr to /dev/null.

Two ADR-2782 D2 vocabulary widenings were forced by surveying the
shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and
outputChannel 'file-arg' (Codex writes via -o and discards stdout,
#1698). Both are additive and closed; Phase 2 owns the validator.

Closes #2690

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): make the parity checker total and pin the lane slug grammar

Findings from the orthogonal review passes.

Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2
"verbatim" while diverging in three undisclosed ways, which is the
translation layer Phase 2 was supposed to be spared:
- `transport` moves from `invoke.transport` to the LANE level, a sibling
  of `probe`/`invoke`, exactly as D1's manifest example places it. The
  nested form read better as a TS discriminated union; the union is now
  discriminated at the lane level instead, which costs nothing.
- The header and the CONTEXT.md glossary now enumerate all FOUR
  widenings (adding `outputArg` and `flags[]`), not two.

Standards axis — CLAUDE.md requires a fast-check property test for a
parser, and `checkReviewerLaneParity` parses markdown for markers and
headings. Adding one found two real defects that the hand-written
matrix missed:
- NOT TOTAL: a malformed descriptor entry threw on `lane.flags`
  iteration, contradicting the module's own "never throws" claim. Every
  field is now narrowed from `unknown` at the trust boundary and
  reported as MALFORMED_LANE / INVALID_SLUG. This matters because
  Phase 2 feeds this function third-party overlay data, and a parity
  gate that crashes is indistinguishable from one never run.
- SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a
  slug outside that class was unmatchable — its marker could be present
  and correct and the scan would still report LEG_MARKER_MISSING
  forever. LANE_SLUG_RE now pins the grammar and a violating slug is
  reported INVALID_SLUG. A loud named violation beats a silent miss.

Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371):
seeding from the module's own matchers could only produce documents
those matchers already recognize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): register the new bin/lib module in the ESLint ignore list

The remote runner caught this; lint:ci did not, because the invariant
lives in the test suite rather than the lint chain:

  tests/repo-invariants.test.cjs
  "each bin/lib/*.cjs is linted xor ignored according to migration state"
  -> tsc-generated bin/lib modules not yet added to ESLint ignore list:
     review-lane-descriptor.cjs

Adding a src/*.cts module ripples to six surfaces (.gitignore, the
ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md
glossary, the capability/inventory manifests, and any size baseline).
The other five were covered; this was the miss.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced

Building the Phase 1 descriptor table against all eleven shipped legs is
the first time every lane's contract was written in one place, and it
surfaced four cases the ADR's original survey did not cover. Amending
the design lock rather than diverging from it, so Phase 2 (#2795)
implements the manifest validator against the amended vocabulary instead
of rediscovering the gaps.

All four are additive widenings of closed enums; no decision reverses:

- D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it
  reviews the working-tree diff.
- D2 outputChannel gains `file-arg` — the ADR called a file-writing lane
  a shape a real CLI *could* take; codex already is one, writing via
  -o/--output-last-message and discarding stdout (#1698).
- D2 gains `outputArg`, required iff file-arg — knowing the review lands
  in a file is useless without the argument naming it.
- D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes —
  antigravity is selected by both --antigravity and --agy, which a
  single-valued field cannot express.

This is the same evidence path that produced the openai-http transport:
the vocabulary widens on a lane that exists, under review, never on
speculation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2794): backfill changeset pr number to 2820

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:31:32 -04:00
Tom Boucher
9624167eec fix(#2810): accept the documented effortSurface axis on EoS registry entries (#2813)
* fix(#2810): accept the documented effortSurface axis on EoS registry entries

The EoS registry schema required an exact eight-key `interactions.axes`
object, while `docs/registries/README.md` and `CONTEXT.md` both documented
nine keys including `effortSurface`. An entry that faithfully mirrored its
upstream descriptor's `effortSurface` key was rejected outright.

`effortSurface` reached the runtime-descriptor vocabulary through ADR-1239
amendment #2481 (`HOST_INTEGRATION_AXES`), but the registry's hand-maintained
copy of that vocabulary never picked it up. The runtime-descriptor surface is
guarded by tests/host-integration-validator-parity.test.cjs; the registry copy
had no equivalent guard, which is what let the two drift.

Accept `effortSurface` as an OPTIONAL ninth axis validated against the
canonical ['argv','none'] rather than a required one: registry entries mirror
their upstream registry/eos-entry.json byte-for-byte, so requiring it would
retroactively invalidate every entry published before the amendment.

Adds tests/registry-axes-parity.test.cjs, which asserts that every key shared
between the registry vocabulary and HOST_INTEGRATION_AXES has an identical
enum array, plus limit-1/limit/limit+1 boundary coverage on the axes key set.

Closes #2810

* test(#2810): fail when a canonical axis is added but never mirrored

The enum-equality assertion compares only keys the registry and
HOST_INTEGRATION_AXES already share, so it is blind to the exact drift that
produced #2810: a new canonical axis appears and the registry copy is never
told. Verified by simulation — mutating an enum is caught, adding a new
canonical key is not.

Assert instead that every HOST_INTEGRATION_AXES key is either modeled by the
registry or named in an explicit NOT_MODELLED allowlist (subagentToolkit and
isolation, both dispatch sub-fields the registry collapses into its free-form
dispatch summary). Adding a canonical axis now fails until someone decides
which bucket it belongs in. The allowlist is itself guarded against going
stale.

Refs #2810

* fix(#2810): harden the axis value lookup with the CodeQL barrier pattern

Both orthogonal reviews flagged the same line: `AXES[key] !== undefined`
is not an own-property test, and the bracket reads are shaped like a
prototype-pollution sink even though the unknown-key gate above provably
makes them unreachable.

Switch the presence test to `Object.hasOwn` and add the repo's inline
literal guards (`capability-state.cts:146-155`, "Prototype-pollution guard
(inline literal, CodeQL barrier)"), which CodeQL can follow where it cannot
follow the `.includes()` filter that actually does the work.

Behavior is unchanged — re-verified all five axes key-count shapes plus a
genuine own `__proto__` property built through JSON.parse (the shape a
third-party registry PR would submit): it is rejected as an unknown key and
Object.prototype is untouched.

Refs #2810

* chore(#2810): backfill changeset PR number
2026-07-29 07:00:35 -04:00
Tom Boucher
1e3c995e6f fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it

Every input to `diffEmitted` is base-relative -- `baseline` vs `current`,
`changedPaths` from `git diff base...HEAD` -- except the ack set, which
was read absolutely, from the working tree only. A differential machine
consulting a non-differential input.

So `staleAcks` asks exactly one question, "did a delta consume you?", and
that cannot distinguish an ack that never explained anything (an
authoring mistake) from one whose ripple is now absorbed into the base
(the ack's SUCCESS condition). After merge an ack is in the second state
but reports as the first.

The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request
events only, so a push to `next` falls through to origin/next -- the very
commit under test. Both sides build identical content, no deltas remain,
and every live ack is reported stale. PR #2768 acked a deliberate 40866
-> 42020 byte growth, was green on its own lane, and reddened `next` the
moment it merged. It also reds every PR branching off the poisoned base,
and since publish-emitted-baseline is gated on the test job, it blocked
baseline publication too.

Give the ack the base side it was missing. `diffEmitted` now takes
`baseAck` -- the same document at the base ref, via `readAckFileAtRef`.
An entry already present there is SPENT: it may no longer consume a delta
and is never reported stale, only surfaced as `spentAcks` for tidying. An
entry new or reworded in this diff stays live, and if nothing consumes it
that genuinely fails, with blame on the author who just wrote it.

This closes a hazard the IMPLEMENTATION named but could not prevent -- a
leftover ack silently pre-clearing the next ripple on its path. (ADR-2719
§3 asserted only that TOUCHING the file is the alarm; its residual-risk
list never covered pre-clearing, and §3 now carries an amendment.)
Verified against the two-PR laundering sequence -- land an innocuous ack,
then change the artifact -- which passed silently before and now fails on
both the hash pass and the size ratchet.

Three things the design has to get right, each of which was wrong first:

  - A read failure on the base document THROWS; only absence-at-the-ref
    returns null. Returning null on error LOOKS armed (every entry stays
    live) but a live entry's defining power is that it CONSUMES a delta,
    so null is armed on the staleness axis and DISARMED on consumption --
    silently the whole pre-#2789 gate. `git show` cannot tell absence
    from fault, so absence is established with `ls-tree`.
  - Re-arming a spent ack costs actual PROSE. Internal whitespace and the
    zero-width family collapse, and `runtime` is not compared: a doubled
    space, an invisible character, or a decorative field would otherwise
    re-arm an ack whose justification still describes the previous
    ripple, showing a reviewer nothing.
  - `baseAck` is REQUIRED once an ack declares entries -- omission is an
    error, not a silent "inherit nothing" -- so a dropped argument fails
    loudly instead of quietly restoring this bug with the suite green.

Because a corrupt document ON THE BASE is expensive (the loud base-side
failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs
blocks one from landing. It is standalone rather than importing parseAck
-- scripts/ ships in the npm package and tests/ does not -- so a parity
test runs both surfaces over one corpus and fails on divergence; it
caught one immediately, a `null` document, now classed as policy rather
than schema. Deadlock is separately foreclosed: a tree carrying no ack
never reads the base, so the PR that DELETES a corrupt file still lands.

`readAckFileAtRef` takes an injected git runner so all four branches are
tested deterministically; it never executes in the remote runner, where
the real-tree test skips for want of a base ref. It also refuses an
option-shaped ref, since execFileSync's array form stops shell
metacharacters but not git's own option parsing.

Rejected: skipping the differential when base == HEAD. It treats the
symptom, costs real coverage on the push-to-next lane, and does nothing
about the downstream PRs the same flaw was reddening.

Deletes the now-spent tests/emitted-drift-ack.json, and updates the
CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a
LIVE entry is, and a spent one is inert.

Closes #2789

* chore(#2789): backfill changeset PR number
2026-07-28 21:28:10 -04:00
Tom Boucher
e276cc7f00 enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate

check-glossary-refs asserts that every backticked tests/ token in
CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719
section 3) is absent on a healthy next BY DESIGN — it appears only
inside a PR that needs it, which is what makes touching it the alarm.

It passed before only by accident of backtick pairing: CONTEXT.md's
RULESET entries are themselves backtick-wrapped and contain backticks,
so the token happened to fall outside a code span. Any edit that
shifted the parity exposed it. A gate that passes by luck is not
passing.

The exemption is exact, not a prefix hole: a sibling missing tests/
path still fails, and a test locks that.

* feat(#2778): make the size-ratchet failure name its own remedy

The growth branch stated a requirement and withheld the means of
satisfying it: no ack file named, no schema, no key format, and no
do-not-regenerate line — so the likeliest guess was to hunt for a
baseline that #2724 deleted. Observed live on #2543.

All remediation now comes from one frozen REMEDIATION export whose
example document is rendered from ACK_VERSION, so the taught schema
cannot drift from the schema parseAck accepts. A round-trip test feeds
the printed document back through parseAck.

The report is now built as a typed IR (buildReport) that formatReport
renders, so tests assert on structure rather than prose, per
CONTRIBUTING.md's raw-text-matching rule.

Two defects found and fixed inline while building:
- diffEmitted's validation early-return omitted newFileCapExceeded
  while formatReport reads its length, so the branch that reports a
  failed git diff threw a TypeError instead of naming the problem.
- Printing one complete ack document per failing branch made each read
  as the whole file, so pasting the second over the first silently lost
  an acknowledgment. One document now covers the whole report.

Closes #2778

* chore(#2778): backfill changeset pr number to 2780
2026-07-28 18:26:55 -04:00
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Tom Boucher
9138271b5f test(#2723): differential emitted-attribution check, dual-run beside the golden (#2737)
* test(#2723): differential emitted-attribution check, dual-run beside the golden

Phase 3 of #2719. The conservation law itself, running BESIDE
golden-install-parity.test.cjs -- both green, fixtures untouched.

Every emitted path whose hash moved between next HEAD and PR HEAD must be
attributable, through the Phase 2 table, to a path the PR actually changed.
Unattributable deltas fail with the paths NAMED. The only way through is a
committed acknowledgment, never a flag -- a contributor facing a red gate
sets a flag, which is what UPDATE_GOLDEN=1 is today.

The central decision is that the law is a PURE function (no fs, git,
installer, or clock), with I/O confined to a separate resolver. The naive
one-big-integration-test shape would need ~38 installer spawns per assertion,
so #2723's four failing-first criteria would not in practice have been
written -- which is exactly how a phase ships promised-but-not-built. Pure,
they are millisecond table tests, and the Stryker gate can actually bite.

Buckets are conserved: every moved path lands in exactly one of
attributed | unattributable | acked, property-tested at 400 runs. A path the
provenance table cannot resolve surfaces as an error, never a silent skip.

Asymmetries that are deliberate, each with a test:
- an ADDED emitted key is a ripple too, not just a modified one
- synthesized paths are exempt; code-derived ones are NOT (Phase 2 refused to
  mark them exempt precisely because exempt means permanently blind)
- shrinkage needs no ack; growth does. Gating shrinkage would punish exactly
  what the size ratchet wants
- a STALE ack is a hard failure -- an ack outliving its ripple pre-clears the
  next one on that path
- a failed `git diff` is an explicit error, never an empty changedPaths set;
  reading it as "nothing changed" would make everything unattributable and
  produce a failure storm that reads like a real finding
- prefix sources are SEGMENT-aware, so `agents/` does not attribute
  `agentsfoo/x.md`

Baseline is cached, not committed, keyed on the next sha. A stale key is
refused rather than used: absence fails loudly and gets fixed, whereas
staleness produces a confident wrong answer. An explicitly pointed-at
GSD_EMITTED_BASELINE that is stale is a hard stop; a stale cache falls
through to the in-job build. No baseline-unavailable path returns -- ADR-2719
section 6 names that trap, since in node:test a bare return is a PASS.

Both the conservation property and the staleness gate were mutation-verified
(injecting a swallowed key fails 9 tests; disabling the staleness comparison
fails 5).

Fixtures, generators, the merge driver and the ADR status are untouched --
those are Phase 4 (#2724).

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): compute stale acks once, after the size pass

Self-review defect found while the reviewers were running. `staleAcks` was
computed twice: once between the hash pass and the size pass, then again
after. Only the second value was returned, so the first was dead code -- and
the dead one was placed where it would have been WRONG.

An acknowledgment can be consumed by either a hash move or a size growth.
Computing staleness before the size pass reports a legitimate growth ack as
stale, which is a false failure that pushes a contributor to delete the very
ack that is doing its job.

Now computed once, after both passes, with a regression test. Verified by
mutation: restoring the early computation fails 2 tests.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): wire the attribution check to the real tree, not just synthetic input

An isolated reviewer caught that the first cut was INTERFACE-ONLY: nothing
read the ack file from disk, nothing shelled git, nothing built real
manifests. Every test was true of hand-built inputs and none of the repo, so
the acceptance criterion "both this check and golden-install-parity green on
the same tree" was trivially true rather than meaningfully true. That is the
promised-but-not-built failure this epic keeps finding in its predecessors,
recurring one phase later for the wiring itself. Taken, not argued.

Adds tests/helpers/emitted-runtime.cjs -- the only module that touches git,
disk, or the installer -- and an integration test that runs the same pure law
against reality:

- CURRENT side: 19 real installer spawns via runMinimalInstall +
  buildParityManifest, the same machinery the golden harness uses.
- BASELINE side: `git show origin/next:<fixture>`. That is next's RECORDED
  emitted state and it costs nothing. Deliberately NOT the working-tree
  fixtures, which are whatever this PR's author regenerated -- comparing
  against those would be vacuous. Phase 4 deletes the fixtures and swaps in
  resolveBaseline's cache path, already implemented and tested.
- changed paths from real `git diff --name-only origin/next...HEAD`, with the
  git subprocess bounded at 30s per CLAUDE.md's unbounded-subprocess rule.
- the real tests/emitted-drift-ack.json (absent is legal; present-but-empty
  or unparseable throws rather than being read as absent).

Verified it can actually fail: an uncommitted edit to a shipped workflow
moves emitted output but never appears in the committed diff, and the check
names all 18 affected emitted paths with the message format ADR-2719 §1
specifies. Restores clean.

Also from review:
- readAckFile now has a real test exercising the SUT across absent / valid /
  empty / unparseable / unreadable. The previous test asserted fs behaviour
  rather than SUT behaviour, because no SUT ack-reading path existed yet.
- formatReport's sampleLimit gains true limit-1/limit/limit+1 coverage at
  19/20/21. A test was previously NAMED "(limit+1)" while testing no numeric
  limit at all, which is worse than no coverage because it reads as covered.

Windows uses an explicit t.skip (install output is platform-specific there,
mirroring the golden harness) -- never a bare return, which node:test scores
as a PASS.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): cover the claude-local manifest family in the real-tree check

Isolated adversarial review, MAJOR. The real-tree wiring enumerated
Object.keys(RUNTIME_META) -- 18 entries -- while the emitted manifest set has
19 families. The 19th is claude-local: claude is the reference host and the
only runtime with a distinct LOCAL "legacy flat-commands" layout
(commands/gsd-*.md + agents/gsd-*.md at project scope), which
golden-install-parity.test.cjs guards with a hand-coded test outside its
RUNTIME_META loop (#2086).

The family was dropped from BOTH sides, so the test's own self-check
(current.length === baseline.length) passed vacuously at 18 === 18. A PR
changing Claude's local-scope output would have failed the golden while this
check reported ok -- and that disagreement is precisely what the dual-run
window is designed to surface as a provenance-table hole. A wiring omission
masquerading as one is the worst available failure here, because it would
have been read as evidence about Phase 2 rather than a bug in Phase 3.

Fixed by deriving MANIFEST_FAMILIES explicitly (18 global + claude-local at
local scope) instead of inferring the set from RUNTIME_META.

The self-check is also repaired: it now asserts both sides against the
INDEPENDENT EXPECTED_MANIFEST_COUNT from the Phase 2 table, and asserts
claude-local specifically. Comparing the two sides to each other can never
catch a family missing from both -- the assertion has to come from outside.
Verified by mutation: removing claude-local again fails the test.

Also from the same review:
- sourceSatisfiedBy returns the matched source string, so an empty-string
  source would return '' and the caller's `if (hit)` would silently discard a
  real match. Unreachable today (every rule source is a non-empty template)
  but a footgun for the next rule author; now `!== null`.
- the purity fixture used a single-element changedPaths array, so an in-place
  sort would have been invisible. Now three elements in unsorted order.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2723): resolve the base ref tolerantly instead of hard-requiring origin/next

The first matrix run failed on both linux lanes:

  differential attribution over the real tree
  cannot resolve origin/next (Command failed: git rev-parse origin/next)

Not a flake, and not an environment excuse -- a real defect in this diff. The
gsd-test runner shallow-clones and merges base+head, so no origin/* remote-
tracking refs exist in the container. My own fail-loud path fired correctly;
what was wrong was hard-depending on that ref existing. GitHub Actions has the
same shape by default, which is exactly why changeset-required.yml carries an
explicit `git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"`.

Now resolved through an ordered candidate list -- GSD_EMITTED_BASE (explicit
lane override), then origin/$GITHUB_BASE_REF and $GITHUB_BASE_REF, then
origin/next and next -- de-duplicated, each verified with
`rev-parse --verify <ref>^{commit}`.

When NO candidate resolves the test takes an explicit t.skip() naming every
ref it tried and stating that the gate did not run here. That is the
ADR-2719 section 6 distinction: t.skip is REPORTED as skipped, whereas a bare
return is scored as a PASS. Hard-failing was the other option and is wrong --
it would make the suite permanently red wherever a base ref cannot exist by
construction, which is a statement about the checkout, not a propagation
finding.

The candidate ordering is pinned by a unit test rather than left implicit,
since the ordering IS the fix.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 23:22:37 -04:00
Tom Boucher
1f6822ccba test(#2722): emitted-artifact provenance table with a totality guard (#2735)
* test(#2722): emitted-artifact provenance table with a totality guard

Adds the declarative emitted-path -> source-path table that ADR-2719 §2
specifies, plus the totality guard that keeps it honest. Phase 2 of #2719.

Every emitted path across all 19 committed golden-parity manifests (8,524
paths) must match exactly one rule. Zero matches, two matches, and a rule
matching nothing are all hard failures, so a new emitted family fails the
build loudly instead of passing through unattributed.

The measured surface is larger than #2722 estimated from claude.json alone
(26 top-level families across 19 runtimes, not 13), which is itself what the
totality guard exists to surface. It resolves to 19 rules.

Building the table caught three false attributions that were total but
resolved to repo files that do not exist -- Copilot's `<name>.agent.md`
rename, Kimi's code-literal `agents/gsd.{yaml,md}` root agent, and Copilot's
`hooks/gsd-session.json` registration. The "every attributed source exists"
test is therefore a first-class gate, not a nicety.

Notable correctness decisions:
- Emitted shapes are hard-coded; deriving them from the installer would make
  the guard tautological (it would follow any installer change silently).
  Only source paths read a first-party descriptor, and only where the
  descriptor is the sole declaration (hostBehaviors.nativePlugin.source).
- Emitted skills attribute to commands/gsd/*.md, NOT the repo skills/ dir --
  that directory is generated from commands/gsd by gen-plugin-skills.cjs, so
  attributing to it would be false attribution that still passes totality.
- Attribution is keyed on (rel, runtime): plugins/gsd-core.js has different
  sources for opencode and kilo.
- Rule order carries no semantics (property-tested), since exactly-one
  matching is enforced rather than first-match-wins.

Nothing here reads a git diff, builds a live manifest, or touches a fixture;
the differential check, drift-ack file and size ratchet are #2723.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* docs(#2722): record the delivered provenance table in the CONTEXT.md glossary

The `### Emitted Artifact Provenance` entry landed in #2721 describing the
table as future work. Phase 2 delivers it, so the glossary now records what
actually exists and the invariants #2723 must preserve:

- where the table lives, its rule count, and that it is total over all 8,524
  emitted paths across the 19 manifests
- dead-rule detection, so table rot is loud in both directions
- the corrected surface measurement (26 families, not the 13 estimated from
  claude.json alone)
- the two invariants #2723 inherits: shapes hard-coded (deriving them would
  make the guard tautological), and attribution keyed on (rel, runtime)
- the skills/ false-attribution trap, and that totality does NOT catch a
  wrong-source rule — the source-existence assertion is what does

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2722): close five review findings on the provenance table

Two orthogonal review passes plus an isolated adversarial reviewer returned
findings at minor..major. All fixed; no blockers were raised.

Standards axis (CONTEXT.md:456, RULESET.TESTS.guard-toplevel-readFileSync):
- module-level loadManifests() threw at require time before any test()
  registered, turning a missing fixture dir into an opaque crash instead of
  one named failure. Now a memoized lazy accessor.
- matchRules and assertTotality each carried their own copy of the matching
  loop; assertTotality now calls matchRules. That is the #2266 divergence
  class, and two copies could let the guard and the attributor disagree.
- named the corpus stride constant; dropped an inline require.

Spec axis:
- the CONTEXT.md glossary carried a "26 families" figure that is not
  reproducible from the code and that no test pinned -- a hand-maintained
  number in permanent canon, i.e. exactly the silent drift this epic exists
  to end. All volatile counts are now removed from the glossary, with the
  reason stated inline: the guard recomputes them every run, so they belong
  in a failure message, not in prose. No test was added to pin the count,
  because that would rebuild the brittle committed number we are deleting.

Isolated adversarial review:
- `.+` tail captures let a `..` segment reach a constructed source path that
  resolves outside the repo. Not live-exploitable (fixtures are committed and
  the only consumer is an existsSync probe) but Phase 3 feeds these strings
  into a diff-consuming check, so assertSafeRelPath now fails closed once, in
  matchRules, rather than per-rule.
- attributeEmittedPath's ambiguous-match branch was never exercised; only
  assertTotality's parallel path was. Now tested directly.
- sampleLimit's truncation branch had no limit-1/limit/limit+1 coverage.
- the fast-check property could not fail for the reason it was named for.

That last one took two attempts and is the one worth reading. The property
hand-rolled its shuffled side from the per-rule matchOne primitive, which is
order-independent by construction, so it held for reasons unrelated to the
shipped matchRules. Routing it through the real matchRules was still not
enough: on an unambiguous table, first-match-wins and collect-all return
identical results for every path (measured: 0 of 190 corpus paths differ).
Order can only matter where more than one rule matches, so the property now
also asserts that an intentionally ambiguous table reports BOTH hits as a set
under every permutation. Verified by mutation -- injecting a `break` into
matchRules makes it fail, and restoring makes it pass.

Enabling all of the above: matchRules and attributeEmittedPath now take an
injectable rules table, so tests can drive the real code path instead of
re-implementing it by hand.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:51:51 -04:00
Tom Boucher
80778e2674 fix(#1881): report an unreadable ROADMAP instead of reading it as absent (#2729)
* test(#1882): stage one file per commit in the base-ref ancestry fixture

CI failed on ubuntu-24 inside this test's setup loop, before any code under test
ran: at commit 32 of 60 the index referenced a blob whose object write had not
landed -- "invalid object ... for 'base-31.txt' / Error building trees".

The loop staged with `git add .`, which re-stages every file already in the tree.
Across 60 iterations that rehashes O(n squared) blobs -- roughly 1,800 stagings
and 60 full index rewrites to add 60 one-line files -- and that churn is what the
object store failed under. Each commit only ever adds a single new file, so
staging that one path is equivalent and removes the redundant work entirely.
Verified the loop still builds the intended history: 61 commits, git fsck clean.

The fixture already carries a note from an earlier fix in this epic recording
that it passed on ubuntu-22 and windows-24 and failed on ubuntu-24 for the same
commit. That was a different stage -- fetch versus diff -- but the same lane and
the same brittleness, so this is the second time this fixture's cost has surfaced
as a red build rather than as a test failure.

Not caused by this PR's change, which touches two configuration lists and cannot
reach a scratch git repository in tmpdir. Fixed here rather than deferred,
because the run surfaced it.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1881): prove an unreadable ROADMAP is indistinguishable from an absent one

Failing-first. Encodes the issue's runtime repro: an unreadable ROADMAP.md makes
getRoadmapPhaseInternal return the same null it returns for "phase not found",
and getMilestoneInfo return the same {v1.0, milestone} it returns for a project
with no roadmap at all -- so a permission or I/O fault reads as a brand-new
project.

Half these cases exist to hold the opposite line. getMilestoneInfo has no
existsSync guard, so platformReadSync's null-for-ENOENT is converted to a
synthetic Error carrying no errno, and that lands in the SAME catch as a real
EACCES. Reporting unconditionally there would flag every project without a
ROADMAP.md -- every brand-new project -- as corrupt. The absent case, the
errno-less error, a non-string errno, unparseable content and a genuinely missing
phase are all pinned silent.

One case guards a decision rather than behaviour: an unreadable STATE.md alone
must stay silent, because the inner catch that swallows it is deliberate and
documented under the #2245 audit as an optional enhancement falling back to
ROADMAP-only heuristics.

Two more pin the invariant ADR-1411 names explicitly -- neither function may
throw, because src/state.cts removed its own defensive try/catch on the strength
of that guarantee.

Assertions are on the frozen reason enum and the emission counter, never on
diagnostic prose. Faults are injected by overriding the platformReadSync seam and
restoring in t.after(), never chmod 0o000, which root bypasses.

Adds the ROADMAP_UNREADABLE reason to the shared vocabulary as scaffolding; no
call site emits it yet, which is what makes these tests red.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1881): report an unreadable ROADMAP instead of reading it as absent

getRoadmapPhaseInternal returned null for a read failure exactly as it does for
"phase not found", and getMilestoneInfo returned {v1.0, milestone} exactly as it
does for a project with no roadmap -- so a permission or I/O fault presented as a
brand-new project and workflows synthesised a blank phase or skipped requirement
extraction with no signal.

Both return values are preserved exactly, per ADR-1411's amendment: continuity is
correct, the silence was the defect. Each catch now reports through the shared
unusable-input seam that shipped with #1882 rather than a second copy of the same
mechanism.

The discriminator is the errno, and it is load-bearing in the silent direction.
getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT
is converted into a synthetic Error with no code that lands in the same catch as
a real EACCES. Reporting unconditionally there would flag every project without a
ROADMAP.md -- every brand-new project -- as corrupt. A genuine read fault always
carries an errno; absence never does. The parse is regex over text and cannot
throw, so nothing else reaches these catches.

Neither function gains a throw. ADR-1411 names this explicitly: src/state.cts
removed its defensive try/catch around getMilestoneInfo under the #2245 audit
because it never throws, and two tests pin that. The inner STATE.md catch stays
untouched and silent -- its fallback to ROADMAP-only heuristics is a deliberate,
documented optional-enhancement path, not a fault.

Where the fix belongs was the design question. platformReadSync does not leak: it
keeps absent and unusable as two channels, exactly as an abstraction should. Both
callers re-collapsed that distinction, so the fix is caller-side and the
projection seam -- with roughly ninety other dependents -- is untouched.

Closes #1881

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1881): admit the roadmap reason to the locked vocabulary

The seam documents adding a reason as three coordinated changes -- the enum
entry, the emitting call site, and the test that locks Object.keys(...).sort().
This PR made the first two and the lock caught the third, which is the whole
point of pinning the key set rather than asserting each value exists.

The roadmap suite no longer re-locks the full set. Two complete locks would mean
two files to update every time a later phase adds a reason, and #1883 and #1884
are both going to. The canonical lock stays in the seam's own suite; the roadmap
suite asserts only the value it introduces.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1881): resolve the roadmap path inside the try, not outside it

Naming the file in the diagnostic required the resolved path in the catch, and
the obvious way to get it was to hoist `path.join(planningDir(cwd), 'ROADMAP.md')`
above the try. planningDir throws a plain Error for an invalid GSD_WORKSTREAM or
GSD_PROJECT segment -- one containing a slash, backslash or `..` -- so hoisting it
let that throw escape uncaught.

That broke the exact invariant ADR-1411 names as this file's hazard: src/state.cts
removed its defensive try/catch around getMilestoneInfo under the #2245 audit
because that function never throws. Of its callers only archivePhaseDirectories
wraps it; cmdInitExecutePhase, cmdInitNewMilestone, cmdInitMilestoneOp,
cmdInitManager, cmdInitProgress, cmdProgressRender and cmdStats all call it bare,
so a workstream name with a slash in it crashed the CLI outright instead of
degrading.

The previous commit asserted "neither function gains a throw -- two tests pin
that". That was false. Both tests inject faults through platformReadSync only and
never through planningDir, so neither could have exercised the path that broke.
The guarantee was claimed, not demonstrated.

The path is now declared before the try and resolved inside it, so the catch can
still name the file when there is one, and a path that never resolved reports
nothing and returns the sentinel unchanged. The two test names are narrowed to
what they actually prove -- that a failing READ does not throw -- and a new case
injects the planningDir failure directly, which is what would have caught this.

getRoadmapPhaseInternal carried the same hazard, resolving the path outside its
try since before this branch. It is fixed the same way rather than left: ADR-227
is explicit that throwing breaks pipeline continuity, this read path already
degrades to null for every other failure, and a PR whose purpose is hardening
this invariant is the wrong place to leave the sibling crashing.

Behaviour otherwise unchanged and re-verified: healthy lookups, EACCES reporting
on both functions, absent-roadmap silence, and the errno discriminator all
unaffected.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1881): backfill changeset pr number to 2729

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 20:17:09 -04:00
Tom Boucher
a613caaeef enhance(#2721): regenerating merge driver, regen:derived, and a name for the emitted-artifact family (#2730)
* test(#2721): failing-first suite for the gsd-regen driver and CONTEXT.md parity

Tests precede the implementation per the TDD gate. The driver module does not
exist yet, so tests/git-merge-regen-driver.test.cjs fails at require time; the
contributor-standards parity assertions fail against next as it stands today,
where the standards doc names two CONTEXT.md headings that have never existed.

Refs #2721

* feat(#2721): add the gsd-regen merge driver and regen:derived

The golden parity manifests and the two size baselines are pure functions of
the source tree, so their only correct merge is "recompute" -- something git's
ours/theirs interface cannot express. 140 of 143 conflicted-file instances
across the open PR queue are these files.

The driver deliberately does NOT regenerate. Four probes established that at
merge-driver time neither the working tree nor the index reflects the merge:
both hold the ours side, a file added by theirs does not exist yet, and
MERGE_HEAD is unwritten. Git also invokes the driver once per conflicted path
(20 here). A regenerating driver would therefore read the ours-side tree and
emit a plausible-but-wrong hash manifest -- worse than a conflict, because a
conflict is visible. So it accepts %A, runs zero subprocesses, records the
resolved paths, and prints one notice pointing at npm run regen:derived.
Staleness stays caught where it already was, by golden-install-parity in CI.

Every failure path degrades toward today's behaviour (a normal conflict).
install-tree is deliberately excluded per ADR-2719 section 7.

Also folded in, per the no-defer rule: workflow-size.cjs claimed .md files have
no eol=lf in .gitattributes; git check-attr shows eol: lf, set by .gitattributes
line 2 since #1088.

Refs #2721

* docs(#2721): document regen:derived and the gsd-regen merge driver

Adds the how-to a contributor actually reaches for when the generated parity
manifests or size baselines conflict, in both places they would look: the
merge-conflict path in CONTRIBUTING.md and the full guide in TESTING-SUITES.md,
including what the driver deliberately does not do (it does not clear GitHub's
CONFLICTING label, and it does not regenerate mid-merge).

Also scopes the new contributor-standards parity assertion to the doc's own
CONTEXT.md section. Its first run flagged `## Decision`, `## Consequences` and
`## Standards followed`, which the doc attributes to an ADR body and a PR body
rather than to CONTEXT.md -- a doc-wide extractor would have demanded CONTEXT.md
grow headings that do not belong to it.

Refs #2721

* fix(#2721): stop passing %P to the merge driver — shell injection

The isolated adversarial review found, and I independently reproduced, local
arbitrary command execution.

Git does not invoke a merge driver with an argv array. It substitutes %O %A %B
%L %P textually into the configured string and runs the whole thing through a
shell, and $(...) executes inside POSIX double quotes -- so quoting the
placeholder does not neutralise it. %O/%A/%B are git-generated temp names and
%L is an integer, but %P is the file's own path, chosen freely by any
contributor. A branch renaming a covered fixture to
evil$(touch PWNED_SENTINEL).json executed that command on the machine of every
maintainer who merged it, and the merge still reported success.

Fix removes the input rather than filtering it: %P is no longer registered, so
the driver receives no attacker-controlled argument at all. The marker records
a count instead of path names. A metacharacter filter would have been a guess
about shell grammar; passing nothing is a property. Re-ran the identical
exploit against the fixed command: nothing executed, conflict still resolved.

Two regressions guard it -- a platform-independent assertion that the
registered command carries no %P, and a real merge driven by the actual
planInstall output with a $(...) filename.

Also from review: CLI dispatch had no coverage at all (CONTRIBUTING's
"CLI and command routing" matrix), which is why runInstall/runStatus now take
{repoRoot} -- hardcoding REPO_ROOT was what made them untestable. Renamed
planResolution to resolveAndRecord since the plan* prefix promised purity it
did not have. Reconciled the eleven-vs-twelve generator count across
CONTEXT.md, CONTRIBUTING.md and the changeset.

Refs #2721

* test(#2721): scope safe.directory for the check-attr helper

The 66f4d85a run failed 11 assertions, all in the .gitattributes scoping block,
with "fatal: detected dubious ownership in repository at '/work'". The test
container checks the repo out at a path its user does not own, so git refuses
check-attr outright. Everything else passed (27,185).

`check-attr` is a pure read of .gitattributes -- no hooks, no filters -- so the
exemption is scoped to that one invocation. It is deliberately NOT applied to
the driver's own production `git config` calls, which run in the user's own
clone and should keep the protection.

Refs #2721

* test(#2721): delete the stale assertion that the driver command carries %P

The plex2 run on bdfd0856 left exactly two failures, both this test: it still
asserted the pre-fix command string, i.e. the vulnerable behaviour. Deleted
rather than relaxed, per RULESET.TESTS.delete-bad-tests -- its useful half is
already covered, in both directions, by
registeredDriverCommandNeverPassesThePlaceholderForTheFilePath.

Refs #2721

* test(#2721): drive the end-to-end merges from the real planInstall output

The e2e helper hand-rolled its own driver registration, and still carried %P.
That meant the five real-git tests were not exercising the production command
string at all -- planInstall could drift and they would keep passing. They now
register exactly what a contributor gets from npm run setup:merge-driver.

Refs #2721

* chore(#2721): backfill changeset pr number to 2730
2026-07-27 19:55:37 -04:00
Tom Boucher
e48eb44003 fix(#1856): give the executor-worktree refusal a handoff instead of a dead end (#2727)
* test(#1856): failing-first contract for the orchestrator cwd-drift guard handoff

The #48 guard correctly refuses to execute waves from an agent worktree, but the
refusal is a dead end: that worktree can hold committed fixes AND uncommitted
work, and "re-run from the orchestrator's own worktree" silently means
abandoning them. The reporter was left choosing between continuing from a
blocked worktree and losing the work.

The guard is shell embedded in execute-phase.md, so these tests extract the
block by a stable marker and EXECUTE it against real git fixtures — the shipped
text is the runtime contract. Covers the stranded-commit and dirty-tree report,
the integration commands, both agent- namespaces, commit-count boundaries 0/1/2,
and the constraints the guard's own comment records: it must NOT fire on an
ordinary branch, on 'agentic-refactor', or on a legitimate feature worktree under
.claude/worktrees/, and must degrade cleanly with no resolvable base or a
detached HEAD.

RED expected: the marker does not exist, so extraction fails and every case errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): give the executor-worktree refusal a handoff instead of a dead end

The #48 cwd-drift guard correctly refuses to execute waves from an agent
worktree — its comment records why ("this is how a wrong-base merge nearly
shipped ~1000 files"), and that refusal is untouched here. The defect is that it
was a dead end.

At the moment it fires, the worktree can hold committed product work, uncommitted
product and planning changes, and the live gap-planning context. Telling the user
to "re-run from the orchestrator's own worktree" silently means abandoning all of
it, because the orchestrator worktree cannot see commits that live only on the
agent branch. The reporter was left choosing between continuing from a blocked
worktree and losing five commits plus uncommitted work.

The refusal now reports what is actually stranded — the commit count and log
against the resolved base, and the uncommitted files — followed by the concrete
integration sequence (commit here, switch to an orchestrator-safe checkout,
merge or cherry-pick, re-run) and a verify command. Nothing is claimed that is
not there: a clean worktree with no commits ahead prints the plain refusal with
no empty sections.

Every added command is diagnostic and `|| true`-guarded, so a failure degrades to
the original refusal rather than crashing before the message prints. Verified: an
unresolvable base still refuses cleanly.

Deliberately NOT done: auto-merging or auto-cherry-picking the agent branch. That
is precisely the operation #48 exists to stop the orchestrator performing from a
drifted cwd, at the moment it has least confidence about which tree is which.
Reporting beats acting here.

The guard block carries a `gsd:guard=orchestrator-cwd-drift` marker so the new
contract test can extract and EXECUTE the shipped shell against real git
fixtures rather than asserting on its characters.

Fixes #1856

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): report true counts, add the changeset, and document the seam

Review findings, all from the isolated adversarial pass:

- The dirty-file list was capped at 20 with no indication, so a worktree with 27
  uncommitted files reported 20 — under-informing the user about exactly what is
  stranded, which is the entire point of this change. Both lists now count BEFORE
  truncating and print "… and N more". Verified with 25 commits / 27 dirty files.
- The has-commits condition was written out twice and could drift on a future
  edit. Collapsed to a single _WT_HAS_COMMITS flag.
- The changeset fragment existed but was untracked, so it was in neither commit
  on this branch and the PR gate would have failed against real history.
- CONTEXT.md:122 documents this exact seam ("the orchestrator runs a cwd-drift
  guard at execute_waves entry…") and was not extended. Now records the handoff
  report, that the refusal condition and exit code are unchanged, and that every
  added command is diagnostic and || true-guarded.

Verified NOT a defect, correcting the review's premise: the guard block does break
when its line endings are CRLF, but .gitattributes:2 is `* text=auto eol=lf`,
which OVERRIDES core.autocrlf and forces LF on checkout on every platform
including Windows — so the shipped file is LF there too, and the installer copies
it through Node without translating endings. The reproduction (mine and the
reviewer's) required injecting CRLF by hand. It is also not fixable from inside
the script: a \r breaks the shell parse at the block's first line, before any
#1856 code runs. Neither introduced nor amplified by this change.

Also noted and left as-is by design: the review flagged that #1856's "offer an
explicit recovery option" could be read as requiring an interactive/automated
integration rather than printed instructions. Deliberate — see the commit that
added the block: auto-merging is the exact operation #48 exists to prevent the
orchestrator performing from a drifted cwd. Called out in the PR body for a
maintainer decision rather than silently chosen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#1856): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 18:32:00 -04:00
Tom Boucher
9a76ca6783 fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter

extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.

Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.

The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.

Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (3eb1cede2), making it
the only non-text file under src. file(1) reported it as data and text tools
silently skipped it, defeating the audit rule that says to search the authored
source; tsc passed because the bytes sat inside a comment, so no gate caught it.
It is live on next.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): pin unterminated-frontmatter detection and its negative space

Covers the discriminator on both sides. The positive rows are the issue's own
repro (LF and CRLF) plus the key-count boundary 0/1/2 around the ">= 1 parsed
key" threshold. The negative rows are the documents that reach the same branch
and must stay silent -- above all a Markdown thematic break at byte 0, which is
how this class of check has previously shipped a false positive on valid
Markdown.

Deduplication is tested on both halves of the composite key: a repeat of the
same (path, cause) is suppressed, a genuine second failure in a different file
is not, and a Windows and POSIX spelling of one path resolve to a single key.
The reset seam is asserted to actually clear -- #2674 is the precedent where a
reset that silently failed to clear made every later dedup assertion a vacuous
pass, and the cases only passed because each happened to pick an unused key, so
every case here uses a path unique to itself.

Assertions are on typed surfaces throughout -- the frozen reason enum and the
dedup-set size -- never on diagnostic prose. The one CLI-level case asserts a
differential between two runs (whether stderr is empty) rather than matching a
message, and is the wired user-reachable surface for this fix. Stream failure is
injected by overriding process.stderr.write and restoring it, never chmod 0o000,
which root bypasses.

Two properties guard the ~50 call sites of the changed function: the new
optional path argument is inert with respect to the parsed value, and LF/CRLF
spellings of a document still parse identically.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): raise the truncation threshold and repair the dedup key

Isolated adversarial review found the one-key discriminator false-positives on
ordinary Markdown: a thematic break above a single labelled line -- `Note:`,
`Author:`, `TODO:`, `See:` -- parses as exactly one key and was reported as
corruption, which is the precise failure the design claimed to prevent and the
changeset promised was fixed. The threshold is now two keys. A file truncated
after exactly one key becomes a false negative; that is the same
precision-over-recall direction already taken at zero keys, and every GSD
artefact this guards carries two or more frontmatter keys.

Three dedup-key defects, each of which could silently swallow a real diagnostic:

- Backslash normalization is removed. A backslash is a legal filename character
  on Linux and macOS, so folding it to a forward slash made two genuinely
  different files share one key. Two spellings of one Windows path may now
  report twice; two distinct files can never silence each other. Lost signal is
  the worse failure.
- The key namespaces are tagged so a file literally named like the unnamed
  digest fallback can no longer collide with a path-less caller whose content
  hashes to that digest -- computable for any predictable content, no brute
  force needed.
- The source identity is computed once rather than hashed twice per emission.

Corrects the previous commit. The two NUL bytes in src/config-loader.cts were
NOT in a JSDoc comment as that message claimed; they were deliberate separators
in the live dedup key, and stripping them degraded it to bare concatenation.
They are restored as escape sequences -- byte-identical runtime string, and the
file is text again so grep can see it. The diagnostic script that misled me
indexed a character-offset string with a byte offset.

Also threads sourcePath through the STATE.md and PLAN.md readers so the two
artefacts epic #1879 is actually about name their file rather than reporting
under a content digest.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): correct fixtures and assertions left behind by the review fixes

The previous commit changed two behaviours deliberately and the suite still
encoded the old ones, so gsd-test came back red with six failures across both
lanes -- all of them mine.

Fixtures carrying a single frontmatter key no longer clear the two-key
truncation threshold, so the CLI differential and the two path-less dedup cases
were asserting a diagnostic that is now correctly withheld. They now carry two
keys, which is what a real interrupted write of a GSD artefact looks like.

The Windows/POSIX case asserted that two spellings of one path collapse to a
single key -- the exact folding that was removed because it also collapsed
genuinely distinct POSIX files whose names contain a backslash. Inverted to
assert they now report separately, with the reasoning recorded inline so the
trade is not silently reversed later: mild duplicate noise on one Windows path
is acceptable, a swallowed diagnostic is not.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): name the file at every read site, and report each file once

The diagnostic reached only the four frontmatter CLI verbs, so ~47 of 53 call
sites reported a truncated file under an anonymous content digest instead of
naming it. Since naming the file is the whole point -- it is what an operator
can act on -- that was a gap in the deliverable, not a scoping choice. 43 of 53
sites now pass the resolved path.

Closing it surfaced a defect the original design missed. A single truncated
STATE.md is parsed twice in a normal run: once by the read wrapper, which holds
the path, and again by a pure core downstream, which is handed only the string
and cannot know it. Those two parses keyed separately, so one file produced two
diagnostics -- and wiring more sites made the collision more likely, not less.
Every emission now registers both identities the input could be known by and
checks both before writing, so whichever caller arrives first speaks and the
other is suppressed. Distinct files with distinct content still report
separately, which is the property ADR-1411 actually requires; two files whose
truncated content is byte-identical collapse to one report, which stays the
documented limit.

Ten call sites deliberately keep no path. Two are frontmatter's own round-trip
checks during set and merge, where passing a path would report on every write.
The other eight are the state-transition pure cores, which ADR-1769 defines as
(content, intent, deps) -> newContent with injected I/O; threading a path
through them would contradict that recorded decision, so it is surfaced rather
than taken unilaterally. With the widened key they no longer double-report, and
in the normal flow the named parse runs first, so the file is still named.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): inject the STATE.md path into the transition cores

The six state-transition cores parsed STATE.md frontmatter without knowing
which file it came from, so a truncated STATE.md reached the operator as an
anonymous content digest on exactly the artefact epic #1879 is named for.

ADR-1769 section 3 shapes these as (content, intent, deps) -> newContent with
injected deps, and deps is the seam for precisely this: something the core
cannot derive without doing I/O. It already carries roadmapProvider and a
phase-inventory provider on that basis, each documented as injected rather than
imported so the core stays pure and testable without disk access. A resolved
path is data, not I/O, so an optional sourcePath member extends the established
pattern rather than contradicting it, and every existing stub keeps compiling
because the member is optional.

updateCore and reconcileCurrentPosition take no deps and are left alone. With
the widened dedup key they cannot double-report, and in the normal flow the read
wrapper has already named the file by the time they run.

Also regenerates gsd-core/bin/lib/state-transition.cjs. That artifact is tracked
rather than gitignored, unlike most of its siblings, so leaving it stale would
have shipped a runtime without this change to anyone reading the repo without
building. tsc had skipped the re-emit because its incremental build info still
recorded an emit that had since been reverted, so the stale output survived a
clean build; clearing tsconfig.build.tsbuildinfo forced it. The
compiled-artifact-sync gate is what surfaced the drift and now reports all nine
tracked artifacts matching their source.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): stop the widened dedup key from hiding a second file

The previous commit widened the dedup guard so one file parsed twice -- once by
a read wrapper holding the path, once by a pure core holding only the string --
reported once instead of twice. It did that by checking BOTH keys before
emitting, which silently traded one defect for a worse one: two DIFFERENT files
whose truncated content happened to be byte-identical now collided on the shared
content digest, and the second file's diagnostic was swallowed. That is the
over-coarse keying ADR-1411 explicitly forbids, reintroduced while fixing
something else.

The guard now checks only the key matching what the caller actually knows -- a
named read checks its path key, a path-less read checks its digest key -- while
still recording every key the input could later be identified by. The redundant
path-less re-parse of an already-named file stays silent, and two distinct files
always both report.

Verified across all six orderings: same file named-then-anonymous reports once;
two different files with identical content report twice; two different files
with different content report twice; the same path twice reports once; two
path-less parses of identical content report once; two path-less parses of
different content report twice.

The suite caught this -- twenty failures, all in the unusable-input tests that
reuse one truncated fixture across different paths. The local probe written
alongside the broken change did not, because it compared two files with
different content and could therefore only confirm the expected behaviour.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): count diagnostics emitted, not identities interned

The suite measured the size of the dedup set as a stand-in for "how many
diagnostics were emitted". That held only while one emission recorded exactly
one key. Once an emission began recording every identity the input could later
be matched by -- a path key and a content key for the same file -- the set grew
by two per write and twenty assertions read 2 where they expected 1.

The production behaviour was correct throughout; the proxy was not. Set size
counts identities, which is an implementation detail of the guard. The
behavioural claim these tests exist to make is how many diagnostics an operator
actually saw, so the module now exposes that directly as an emission counter and
the suite asserts on it. The set-size accessor stays for assertions genuinely
about key shape.

The local probe written alongside the change did not catch this because it
counted process.stderr.write calls -- the right thing -- while the suite counted
set growth. Verification now asserts both and requires them to agree, so a
future divergence between the counter and real writes fails immediately rather
than being discovered a bench run later.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): retire two assertions that outlived the behaviour they described

Both tests encoded assumptions the dedup fix invalidated, and both were caught
by the suite rather than by the probe written alongside the change.

The forged-path case asserted that a file named like the anonymous digest
fallback must not suppress a later path-less report. That premise is gone: an
emission now records every identity its input could be matched by, so ANY named
report of some content silences the anonymous re-parse of that same content --
which is the same-file guard working as intended, and has nothing to do with the
crafted name. The property still worth defending is that a crafted filename can
never silence a real file reported under its own path, so that is what the test
now asserts, with the deliberate suppression documented beside it.

The reset-seam case ended by reading the size of the dedup set and expecting 1.
Set size counts interned identities, not diagnostics written, and one emission
now interns two. It asserts the emission counter for the event and keeps a
weaker set-size check for the interning.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): close the review findings on the discriminator, dry-run and counter

Three orthogonal review passes ran against the final diff. Their findings:

A labelled preamble under a leading rule was still misreported. Raising the key
threshold to two only moved the boundary, because two colon-labelled lines are
as common in ordinary prose as one -- a document opening with a rule over an
Author and a Reviewed-by line, then prose, was called corrupt. Key count alone
cannot separate the two. What does is what follows: a write interrupted part way
through a frontmatter block ends mid-block, so every line of the region is still
frontmatter-shaped, whereas a document merely opening with a rule goes on to
prose. Both conditions are now required, and each closes a false-positive class
the other leaves open. Nested list values and indented continuations stay
frontmatter-shaped, so legitimate truncations are unaffected.

`state rebuild --dry-run` reported a truncated STATE.md anonymously. The write
path is named only because readModifyWriteStateMd parses with the path first;
the dry-run branch reads the file directly and never did. Dry-run is the
read-only mode an operator reaches for first when they suspect corruption, so it
is the one that most needed to name the file. reconcileCurrentPosition takes the
path as an optional argument now and rebuildCore passes it down. That function
was previously left alone on the grounds that a read wrapper always names the
file first -- this is the flow that disproves it.

The emission counter counted write attempts rather than writes, so on a broken
stderr it claimed a diagnostic had reached the operator when nothing had. It is
incremented only after a write that completed, and the broken-stderr test now
asserts the count as well as the return value.

Two documentation defects. The module described a guarantee it does not keep:
one file yields one diagnostic only when the named read comes first. The reverse
ordering emits twice, and that is deliberate -- a path-less caller cannot
identify its file, so suppressing the later named report would also suppress a
genuine second failure in a different file whenever two files share identical
truncated bytes, which ADR-1411 ranks the worse failure. The comment now states
the asymmetric guarantee and a test pins it. Separately, the CONTEXT.md glossary
entry still described backslash normalization that a later commit removed, and
asserted the opposite of what the tests pin; no lint checks prose against code,
so nothing caught it.

Also converts three body-level try/finally blocks to t.after(), per
CONTRIBUTING.md's rule that try/finally belongs only in helpers with no test
context -- the file's own emissionsDuring helper already did this correctly.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#1882): tell the operator what the truncated-frontmatter warning means

A user who has just seen the new warning is acting, not studying, so this lands
in the How-To quadrant beside the other "if you see X" branches in
debug-a-failed-execution, not in reference or explanation. It gives them what
the warning means for this run, three steps to restore the file, and the fact
that the warning changes no return value or exit code.

It also states the case that matters more than the warning itself: silence does
not prove the file is intact. GSD says nothing when the partial block carries
fewer than two fields or reads as prose, because a Markdown document opening
with a horizontal rule is indistinguishable from one of those. A reader chasing
missing metadata needs to know not to treat quiet as clean. Why that threshold
exists is explanation and deliberately stays out of a how-to.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1882): backfill changeset pr number to 2712

* test(#1882): constrain each branch of the frontmatter-shape check

CI's mutation gate came in at 61.56 against a threshold of 62, and the surviving
mutants were concentrated in isFrontmatterShaped -- the function added last, in
response to review, and the only one never given tests of its own. It was
exercised solely through extractFrontmatter, which covers the composite decision
but leaves each branch of the predicate unconstrained: drop the blank-line
filter, or any one of the three shape alternatives, and every existing assertion
still passed.

Four cases now pin the halves independently. A blank line inside an interrupted
block must not disqualify it, which constrains the filter and its comparison. An
unindented list item and an indented folded-scalar continuation each exercise one
shape alternative that no other case reaches on its own -- the folded line is
neither a key nor a list item, so it is the only input that distinguishes the
indented branch. And two keys followed by prose must stay silent, which is the
negative half: it fails if the predicate is ever mutated to accept everything,
and it is the case that proves key count alone was never sufficient.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): register the unusable-input suite with the frontmatter mutation shard

The mutation gate reported an identical 61.56 across two runs whose only
difference was four added tests. That is the tell: the tests were never
executed. The frontmatter shard runs a fixed file list in stryker.config.mjs and
scripts/mutation-matrix.cjs, and tests/unusable-input.test.cjs was in neither, so
the entire suite covering the new unterminated-fence branch was invisible to the
gate while passing perfectly well in the normal run.

So the score was not measuring weak tests, it was measuring absent ones: #1882
added mutants to frontmatter.cjs and no test in the shard covered them. Both
lists gain the file; the config already notes they must stay in sync.

This is a registration ripple a new test file carries when it covers a
mutation-tracked module, alongside the .gitignore, eslint, inventory, glossary
and size-baseline ripples a new module carries. Nothing warned about it, which
is why two runs were spent before the identical score gave it away.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:50:12 -04:00
Tom Boucher
09477f925e fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading

The Workflow backend emitted every agent() call with no model, so
model_overrides / model_policy / model_profile were silently inert on that path
while the inline path honored them (ADR-1411). Neither existing suite contained
the string 'model' at all.

The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor')
rather than hardcoding either, so it asserts backend parity rather than a fixed
string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when
nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids
reaching the code generator, the #2285 composed seam, CLI config-defaulting, and
a fast-check round-trip property.

RED expected: no model key is emitted anywhere, and --executor-model does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2686): thread the resolved executor model into the Workflow backend

The Workflow backend emitted every agent() call with no model at all, so
model_overrides / model_policy / model_profile_overrides / model_profile were
silently inert on that path while the inline path honored all of them. The model
was not dropped at the last step — it was absent from the whole seam:
agentOptions() took no model, EmitInput had no field to carry one, and
ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could
not forward one. The generated script asserted the parity it broke.

VERIFY-FIRST, which #2686 flags as the question that decides the fix: the
Workflow tool's agent() DOES accept a per-call model. Its documented signature is
  agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? })
so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled
out. ADR-1143:24's option enumeration omitting `model` is an incomplete
enumeration, not a decision to exclude it.

- agentOptions(p, executorModel) emits `model` only when it is a non-empty string
  that is not "inherit" (#2517: an empty model 404s on runtimes without native
  tier aliases). A non-string is a malformed config: omit, never throw.
- executorModel threaded through EmitInput and ResolveWaveDispatchInput.
- The CLI resolves gsd-executor from project config by DEFAULT rather than
  requiring a flag, reading the same source the inline path reads. An
  orchestrator that never learns about a new flag would otherwise silently keep
  the old bug. --executor-model exists only to pin/override.
- ADR-1411 provenance: the generated header now states which model was applied,
  or that none resolved and why. A fallback must be a visible value.

Compatibility: when nothing resolves, the emitted options object is byte-identical
to before, so every existing caller and assertion is unaffected.

Behavior change (Hyrum's Law): opted-in users move from session inheritance to the
catalog-resolved executor model. Adding a `model` key also changes agent() opts,
which invalidates the cached prefix of any in-flight resumeFromRunId run — a
one-time re-execution. Both disclosed in the changeset.

Fixes #2686

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2686): reject script-breaking model ids and share the emit predicate

The isolated adversarial review found a BLOCKER in my own provenance comment,
proven by execution (the emitted script exited 42 from an injected statement).

U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment
in EVERY engine — the ES2019 change legalized them inside string LITERALS only.
So quoteString (JSON.stringify) is sufficient for the `model: "..."` object
literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in
a model id closed the comment and made the rest of the line live top-level code.
The value is reachable from `.planning/config.json` (model_overrides /
model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any
non-claude runtime — attacker-influenceable in a cloned repo.

`emitWorkflowScript` now rejects a string executorModel carrying any character in
UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to
phaseDir/runId, which is proof the codebase knew this hazard. Rejection is
ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an
emit failure to the inline backend WITH that reason, so the degradation is
visible. A non-string stays on the existing defensive path (omit, never throw) —
that is malformed config, not an injection attempt.

Also from the reviews:

- The predicate deciding "is this model emittable" was duplicated between the
  emission and the comment asserting it. Extracted to emittableModel() so a
  generated comment can never claim something the generator did not do — the
  exact failure class #2686 was filed for.
- That predicate now trims and lower-cases before comparing, closing a real
  #2517-class gap: " " and "INHERIT" were previously emitted verbatim.
- The adversarial test was pass-always against this very vulnerability — it
  asserted only that JSON.stringify appeared. Replaced with the real contract
  (rejection) plus an execution-level check that no LineTerminator survives into
  the comment. A raw U+2028 had also been committed into that test's fixture
  array where a tab was intended; both are now explicit \u escapes.
- optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated
  model contained — silently not testing what it claimed. Now brace- and
  string-aware.

Stale-test corrections in tests/fix-2285-*: three assertions froze the exact
options literal `{ agentType: "gsd-executor" }`. The object legitimately gained
an optional additive `model` key, so they now assert the invariant they exist to
protect (agentType present, isolation absent) rather than a frozen literal. The
CLI-vs-pure equality test pins --executor-model on both sides; otherwise it
compared a config-resolved CLI run against a pure call given no model.

CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the
new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* test(#2686): fix the options extractor and model the rejection path

Two defects in my own test helper, caught by the full matrix:

- optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The
  emitted shape is agent("brief", { ... }), so that never matched and the helper
  returned an empty array, making every assertion over it vacuously true. It now
  anchors on agent( and takes the first balanced, string-aware {...} after it.

- The fast-check property predated the security fix and asserted ok:true for any
  generated string. Strings carrying an unscriptable character are now rejected,
  so the property models the real three-way contract: unscriptable -> ok:false;
  trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the
  trimmed value.

Verified locally against the built module: omit values clean, both plans carry
the model on the parity path, property passes 500 runs at seed 42. Test file
re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are
explicit \u escapes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* test(#2686): scope no-control-regex on the mirrored unscriptable-char class

The class is the point of the assertion — those bytes are exactly what must be
rejected — so the rule is disabled at that line rather than the class weakened.
UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence
the mirror.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#2686): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 16:24:21 -04:00
Tom Boucher
1008aabd31 fix(#2615): document the effortSurface axis in the host-integration matrix (#2698)
* fix(#2615): document the effortSurface axis in the host-integration matrix

#2481 added `effortSurface` as the ninth negotiated `hostIntegration` axis and
wrote documentation-sourced values into 18 descriptors, but never touched
`docs/reference/host-integration-capability-matrix.md`. The matrix that ADR-1239
designates the cited source of truth had zero occurrences of the axis: no entry in
the axes legend, and no row in any of the per-runtime tables. `src/host-integration.cts`
states "every value is documented or explicitly 'undocumented'" — for this axis
that was false for every runtime.

Adds the legend entry (the `argv` / `none` / `undocumented` vocabulary, plus why
there is deliberately no config-file member) and an `effortSurface` row to all 19
per-runtime tables. Every citation is carried over from #2481's own commit message,
where the values were sourced:

- claude   argv -- `claude --help` documents `--effort <level>`
- opencode argv -- `opencode run --help` documents `--variant`
- codex    argv -- `model_reasoning_effort` is a config.toml key, not a dedicated
                   flag, so the generic `-c key=value` override is the only argv
                   route (still argv)
- 15 hosts undocumented -- their docs state no reasoning setting

kimi-code is the nineteenth section (added by #2603 after #2481) and is the one
runtime with no declared value. Its row and a Documentation-gaps entry record why
rather than inventing one: Kimi Code documents `/effort` (alias `/thinking`), but
only as an INTERACTIVE slash command — `-m, --model` is the only model-adjacent
argv. Neither vocabulary member is accurate (`none` would deny a mechanism the host
has, `argv` would claim one it does not expose), so closing that gap needs a
vocabulary decision, which is a negotiation change and not a documentation one. The
absent value already degrades closed exactly as the sentinel does.

The regression test derives its runtime list from the registry rather than
hardcoding it, so a runtime added later fails until its matrix row exists — the
ratchet whose absence let #2481 add an axis with nothing catching the missing docs.

Closes #2615

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* docs(#2615): honest citations for the undocumented rows; fix the four stale 8-axis lists

Two findings from the orthogonal review of the first commit.

1. The 15 `undocumented` rows shared byte-identical text — "searched the runtime's
   official docs (see Sources consulted above)" — which is weaker than this file's
   own convention ("no authoritative doc — searched: <url>") and, worse, implies a
   per-host targeted search that did not happen: each section's Sources-consulted
   list was gathered for OTHER axes and contains no CLI-reference or
   reasoning-effort source. The rows now say plainly what the finding is — an
   ABSENCE established by #2481's cross-host survey — and cite that survey rather
   than implying a URL was checked per host.

2. Four normative docs still described "the eight negotiated axes" and omitted
   effortSurface entirely. The worst of them is
   docs/how-to/add-or-update-a-host-integration.md — the maintainer's own guide for
   onboarding a host, whose Step 2 axis table would have a maintainer reproduce
   exactly the gap #2615 exists to close. Also fixed:
   docs/reference/host-integration-interface.md (which calls itself the normative
   reference and had no effortSurface row at all),
   docs/how-to/author-a-host-plugin.md, docs/registries/README.md ("**exactly** the
   eight … axes keys"), and CONTEXT.md's matching EoS-registry sentence.

Deliberately NOT changed, because they are historical records rather than current
contract: docs/whats-new-1.7.0.md and docs/FEATURES.md's 1.7.0 entry (effortSurface
shipped in 1.8.0 via #2481 — rewriting them would falsify the release history),
ADR-1239's pre-amendment body (already superseded by its own
"Amendment (2026-07-21): effortSurface axis (#2481)"), and ADR-1016's "original
eight axes", which refers to a different axis set entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2615): backfill changeset PR number (#2698)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:02:20 -04:00
Tom Boucher
a5633bb32f enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures

* feat(#2671): brand raw vs calibrated token types

* refactor(#2671): hoist type-compile into a before() hook

Two review responses:

- The fixture compile ran in the describe() body, so it executed at
  collection time even when the block was filtered out, and a failed
  precondition collapsed eight independent assertions into one opaque
  describe-level failure. A before() hook is this repo's documented
  idiom and preserves per-test granularity.

- parseTokensFlag now records WHY it returns an unbranded number: it
  validates the magnitude of --tokens, but the basis is decided by
  --calibrated, so branding here would be wrong for half its callers.
  The assertion belongs to cmdEstimateCheck, its only caller.

* test(#2671): pin each brand diagnostic to its OFFENDING marker

Adversarial review demonstrated that asserting only exactly-one-diagnostic-
at-code-N is not airtight. Repairing a fixture's brand violation while
injecting an unrelated error of the same code (a string passed as the
budget argument) still yielded exactly one TS2345, so the fixture would
have reported green while no longer testing its regression at all.

Each bad-* fixture now routes its violating value through a const named
OFFENDING, and the test asserts the diagnostic's start offset falls inside
that node — located through the AST, so it survives reformatting and never
pattern-matches source text. Replaying the proof-of-concept against the new
assertion rejects it: the diagnostic lands on the budget literal, not the
marker.

Also corrects a doc comment that claimed the program type-checks all of
src/; it covers phase-estimation.cts and its transitive dependencies.

* chore(#2671): backfill changeset PR number (#2676)
2026-07-26 21:42:50 -04:00
Tom Boucher
c3958018dd docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern

ADR-1411 reasons only about a resolution miss. It is silent on input that
is present but not usable, which is how five engine read paths (#1879) could
fold an unusable input into the value meaning 'genuinely absent' without
contradicting an Accepted ADR.

Read together, ADR-1411 and ADR-227 converge and do not license throwing as
the cluster's answer: ADR-227 requires malformed input to be coerced rather
than propagated and carves out only genuinely-fatal fields, while ADR-1411
already permits a fallback provided it is 'a visible value, not a silent
substitution'. The defect in these five sites is therefore not that they fall
back but that they fall back invisibly.

Records the pattern that follows: every current return value is preserved, and
the cause is made visible in-band where the result already carries a
provenance envelope, or out-of-band via a deduplicated stderr diagnostic where
it returns a bare value it cannot extend. Throwing stays confined to ADR-227's
genuinely-fatal carve-out, decided per call. Also names the per-applier caller
audit and the lint-resolution-provenance registry gap.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2674): prove the warning-state reset misses the unknown-key dedup set

The two existing cases in this suite only pass because each picks a key
name no other case reuses, so neither can observe whether the reset the
beforeEach calls actually runs.

Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after
_resetRuntimeWarningCacheForTests(). It is not - the helper clears only
_warnedConfigKeys despite documenting itself as resetting per-process
warning state.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2674): reset the unknown-key dedup set with the runtime warning cache

_resetRuntimeWarningCacheForTests documents itself as resetting per-process
warning state but cleared only _warnedConfigKeys, leaving
_warnedUnknownConfigKeys populated across cases. The suite that exists to
test that set - 'loadConfig - unknown-key warning dedup' - calls the helper
in beforeEach expecting exactly this, so the reset was a silent no-op for
it; both cases passed only because each picked a key name the other never
reused. Any later case reusing a key would have had its warning suppressed
by leaked state.

Found while amending ADR-1411, which names this dedup guard as the pattern
five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the
fix would have propagated the footgun to each of them. Folded in here per
CLAUDE.md's no-defer rule rather than filed.

RED verified on 3c4895841 (test only, no fix): linux-node22 reported
'FAIL tests/config-loader.test.cjs - the documented per-process
warning-state reset must clear the unknown-key dedup set too'.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): document src/ in the changeset-lint trigger list

CONTRIBUTING.md presented the Changeset Required trigger list as bin/,
gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which
scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the
TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is
the most-edited user-facing path in the repo and the omission sends any
contributor who touches it into a CI failure the doc says cannot happen.

Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so
running it bare locally reports success without evaluating the branch. This
PR hit exactly that: a local run said ok_fragment_present and CI failed
fail_missing_fragment on the same diff.

Found while opening this PR; folded in per the no-defer rule.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): restore the round-2 review corrections to the amendment

These edits were made in response to the second isolated review pass but
never staged: later commits used targeted `git add <file>` for the test and
the source fix, so the two markdown files stayed dirty and shipped nothing.
The branch carried the round-1 text, including the ADR-227 misquote the
reviewer raised as a blocker.

Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in
was never implemented, so citing it as the precedent was wrong), the dedup
key, #1882 folded into the out-of-band mechanism instead of a fourth
mechanism-less category, the narrowed caller-audit rationale, and the
test-methodology clause.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 20:25:39 -04:00
Tom Boucher
bd570618d4 feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop

* fix(#2632): calibrate against the raw projection so the loop converges

* test(#2632): add closed-loop convergence guard and codify the feedback-loop rule

* fix(#2632): pair calibration samples per plan; atomic write; amend adr

* chore(#2632): backfill changeset pr to 2672

* fix(#2632): retry renameSync on transient windows errnos and clean up the temp
2026-07-26 16:28:56 -04:00
Tom Boucher
46ba02acde feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs

* fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions

* fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests

* chore(#2630): backfill changeset pr to 2661
2026-07-26 01:42:47 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
4a66d62d10 feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver

Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.

worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.

resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.

Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: rebuild tracked state-transition.cjs to match #2400 source

The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit 2bcfaa2e2) added the progress.total_plans frontmatter sync to source but the tracked bin/lib/state-transition.cjs was never rebuilt, so the fix was not shipping to consumers of the compiled artifact. The mandatory build:lib step for Phase 2 surfaced the drift; recompiling makes the already-merged, already-changelogged #2400 fix effective. Artifact-only resync (no source/test change); drift class tracked by #2591.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 21:23:06 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
BeeHiggs
bf9fe4630d feat(#2249): bracket phase-id core grammar — parse/render/toDir round-trip pair (epic #612 PR-1) (#2258)
* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards

PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md).
Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single
canonical owner — as a pure, additive extension. The 17 locked exports and
PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical,
so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs)
stays green.

New pure round-trippable model (ADR Decision 4):
- PhaseId { project, milestone, phase, subphase?, plan? }.
- parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token
  `GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens
  (`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new
  parser — normalizePhaseName and every legacy reader keep accepting those
  tokens unchanged (conservative default; no existing path gains a throw).
- renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug`
  with a slug guard that sanitizes path-traversal input.
- getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone
  from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning
  the `vN.0` form (parity with READING-A). The optional parameter keeps the helper
  pure (no config read) and byte-compatible — every existing single-arg caller
  resolves to the unchanged READING-A body (ADR Decision 6).
- extractPhaseToken(dirName, convention?): bracket dir branch GATED on
  convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is
  string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal
  family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no
  string-only discriminator is complete — an ungated auto-detect silently
  reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit
  convention signal keeps every existing convention-less call site byte-identical
  (pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs).
- comparator: no new code — comparePhaseNum already orders the dot-decimal
  `PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a
  PR-2 resolution concern (bracketQualifiedKey), not core grammar.
- SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999}
  non-milestone guard; the bracket-prefix reading is gated the same way (an
  ungated read called `P0.0-foundation` a sentinel), legacy leading-int form
  unchanged.
- BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately
  more permissive than parsePhaseId — a read-tolerance source for PR-2, not the
  emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt
  owner so PR-2 builds every bracket read regex from the canonical source and
  check:phase-id-drift stays green stack-wide.

The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*`
grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the
config permits parses. parsePhaseId has no live callers in PR-1, so this grammar
choice is forward-facing for PR-2 with zero PR-1 behavior impact.

Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips,
full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases),
extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens,
sentinel + slug guards, bare-token rejection, exported-source behavioral
assertions, and two generative fast-check properties: render∘parse identity over
well-formed displays, and the toDir/disk↔display bijection. Plus a #2043
numeric-tail characterization (single- AND multi-digit rows) in
tests/phase-id.test.cjs pinning the convention-less reading byte-identical.

The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish)
and rebuilt by CI, so it is intentionally not committed.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3)

PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4).

B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers,
over-padded numbers, multi-space separators, stray whitespace), so
render(parse(x)) === x did not hold for every well-formed x as ADR-612
Decision 4 requires. Both branches now enforce canonicality by construction:
parse permissively, rebuild the canonical string via the same emit path
(renderPhaseId for display, a hand-rebuilt token for dir/token), and throw
"parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's
entry is removed — the match anchors now reject leading/trailing whitespace
outright, folding into the existing "not a bracket phase id" rejection.

M1 (major): toDir only ever guarded the slug; project/milestone/phase/
subphase were interpolated unsanitized, so a hand-built PhaseId (a
structural, not nominal, type) could smuggle a path-traversal segment onto
disk. Every field is now validated against the exact shape parsePhaseId
itself would produce before use.

M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling
trailing hyphen in the emitted dir name. toDir now throws in that case.

M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable
from the dir-branch's plan tail, so it silently broke the disk<->identity
bijection on read-back. toDir now rejects all-digit slugs.

Nits: toDir now rejects a non-string slug instead of coercing it to the
literal token 'undefined'/'null'; sentinel boundary tests added for
milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999},
not an inclusive range — these were already correct, now locked by test).

Test-first: every new assertion (concrete examples + fast-check mutation
property for B1; concrete cases for M1-M3 and the nits) was written and
confirmed red before the implementation changes, per repo TDD convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): reformat changeset body to house convention (review Mi2)

The fragment added in ab26190a was a plain paragraph — no bold headline,
no trailing issue reference. Reformat to the repo's
`**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g.
.changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md).

Uses (#2249), the issue every commit on this branch references, not the
PR number already carried in frontmatter (`pr: 2258`) — the changelog
serializer appends `(#{pr})` unconditionally, so a body also ending in
`(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered
bullet directly via parseFragment + serializeChangelog: it now reads
`... (#2249) (#2258)`, matching the dominant convention across the other
fragments (frontmatter pr = merged PR, body reference = originating issue).

Also moved the docs-exempt marker back before the paragraph -> after it
(matching the file's original order): the marker sits on its own line and
is stripped before the body is used, but placing it first left a leading
blank line in front of the bold headline once reformatted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
two property-generator coverage gaps the reviewer flagged; no source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged).

Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no
property fed a 3+-digit milestone/phase/subphase/plan through parse/render/
toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen
numArb to 1–999 so the round-trip and disk↔display bijection properties both
span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no
leading zero, so canonicality still holds). Add a concrete regression pinning
the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and
toDirs to 'GSD.100-05-feature' without truncation.

Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered
milestone/phase pad + whitespace mutations but never a subphase pad. Add
unpad-subphase / overpad-subphase to the mutation set and a generated
`includeSub` boolean that decides whether the canonical carries a `.SS`
(forced in for the subphase mutations so there is always a `.SS` to mutate);
non-subphase mutations keep their original no-subphase coverage.

Non-vacuity verified against the compiled lib by temporarily probing each
widened/new property and confirming it fails: round-trip counterexample
["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit
tokens are genuinely generated and reach the body; a no-op unpad-subphase
mutation trips the mutated===canonical guard (counterexample
["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached
with a subphase present. Probes reverted; numRuns unchanged.

Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail;
`npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major)

BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that
re-derived the grammar as an unbounded `\d+` literal instead of consuming
PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket
path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug
whose first word is a year) over-collected the token as `01-14-2026` instead of
`01-14`.

Interpolating the cap verbatim at every position was rejected on evidence: the
bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent.
The exactly-2 cap at the others would under-collect ids toDir itself emits —
`PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit
sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and
`[GSD.100] 05` is a pinned regression. Those positions are delimiter-
disambiguated (a required field separator; a dot a slug can never contain),
not heuristically recognized, so they have no year collision to defend against.
Upstream draws the same line for the same reason: core-utils/phase cap the
paired PLAN component while the leading phase component stays unbounded.

So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and
each position takes the width its delimiter affords: leading unbounded, dash-1
and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner
seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of
the token grammar.

Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE
instead of re-spelling it as a literal, so the emit-side gate and the read-side
token source are one rule — the same single-owner discipline this fix is about.
Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant).

Refs #2249

* test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major)

The comment block alone cannot hold the divergence: src/phase-id.cts is exempt
from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not
catch the bracket token source drifting from the seam. Per the Generative Fix
Divergence rule, the divergence is pinned behaviorally instead.

Surface 6 joins the existing #2232 parity gate rather than starting a rival one:
the review named the bracket token source "a sixth continuation-recognition
site", and continuation-grammar-parity.test.cjs is already the invariant-named
home where the five #2043 sites agree with the owner on a shared width corpus.
Surface 6 asserts the same contract at the bracket run's slug-adjacent position
(`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so
the bracket path now fails the same gate the other five do.

A second block pins the DELIBERATE half — the wider canonical width at the
delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out
of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would
look like a cleanup rather than a regression.

The generative property ties the READ side to the EMIT side metamorphically: for
every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that
id's numeric run — no more, no less. It needed a new arbitrary: the existing
slugArb generates one [a-z0-9] word and so can never produce the number-leading
slug the collision requires.

Probe-falsified, both directions (probes reverted):
- reverting the source to the old unbounded `\d+` fails 8: the parity gate
  reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the
  review's scenario verbatim — and the property shrinks to
  ["A",1,1,undefined,"100-a"].
- interpolating the seam at EVERY position (the rejected verbatim option) leaves
  the repro and parity green but fails the divergence gate `'02' !== '02-105'`
  and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the
  evidence that a verbatim cap under-collects ids toDir emits.
Width 2 stays green under both probes — the corpus agrees with the owner exactly
where the old and new rules coincide, so the gate discriminates rather than
merely mirroring the regex.

Refs #2249

* docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major)

* test(#2249): pin deterministic grammar boundary cases (re-review m1)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
the m1 proof gap — the grammar's bounds were exercised only incidentally
through the fast-check domain (1-999, [a-z0-9] slugs). No source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged).

Adds a deterministic boundary block (7 describe groups, +22 tests) pinning
the compiled lib's CURRENT behavior — a proof gap, not a behavior gap:

- m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse
  (display + dir) -> render/toDir round-trip byte-equality. The plan
  position is identity-symmetric (parse/render accept 99/100/101) but toDir
  drops it (filename-surface dimension only).
- m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs
  99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps
  the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of
  the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry.
- m1.3 leading-zero 007 -> not-canonical rejection at every position/form.
- m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing
  slug (never stored, never mis-read as a plan) and rejects a line
  terminator; toDir's allow-list sanitizer collapses each to a safe
  [a-z0-9-] token or rejects sanitize-to-empty.
- m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test);
  an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE;
  an abs-path string is not a bracket id; an abs-path dir slug is dropped to
  a clean tuple.
- m1.6 whitespace-only -> not-a-bracket-phase-id.
- m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral):
  garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes.

No accept-not-reject case is a src bug: parse never STORES an abusive slug
(dropped from the identity tuple) and toDir independently re-sanitizes on
emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by
parse is the documented positional design (toDir drops the plan; the
read-token caps it) — divergence pinned, not papered over.

Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both
its parse-side and emit-side), ran -> 8 distinct named failures, reverted ->
66/66 green. Confirms every new group executes and can fail.

Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar +
continuation-grammar-parity + collision-characterization + phase-id family
175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green
except one pre-existing, unrelated env failure (npm-integrity-gate: a live
npm-audit advisory in the production dep tree — reproduces with this change
stashed; no package.json/lock change here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 12:49:53 -04:00
Tom Boucher
bf8f320083 feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI)

PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting
GSD's kimi support into two distinct products per the user's directive:
- kimi       (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python)
- kimi-code  (new):      Moonshot's Node Kimi Code CLI (~/.kimi-code,
                         runtime: node, KIMI_CODE_HOME env)

Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json
descriptors, not hardcoded branches in install.js. The new descriptor uses
the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml
hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs).

Critical Kimi Code constraint reflected in the descriptor:
  hostIntegration.dispatch.namedDispatch: false
  hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan']
  hostBehaviors.namedSubagentsSupported: false
Kimi Code's official docs confirm only 3 built-in subagents with NO custom-
subagent registration (the [subagent] table only has timeout_ms). The
kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in
kimi-code's artifactLayout.

Schema adjustments:
- subagentToolkit set to 'undocumented' (the existing escape hatch); the
  schema enum (full/read-only) lacks a 'limited'/'built-in-only' value.
  A follow-up PR can extend the schema enum to add 'built-in-only' as a
  first-class axis value reflecting Kimi Code's documented model.

Registration:
- capabilities/kimi-code/capability.json (new descriptor, modeled on codex)
- bin/install.js: allRuntimes array + --all list + --kimi-code flag
- gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases
  (kimi-code, kimicode, kimi_code)
- src/runtime-name-policy.cts: FALLBACK_ALIASES map
- gsd-core/bin/lib/capability-registry.cjs: regenerated via
  scripts/gen-capability-registry.cjs --write

Tests:
- tests/multi-runtime-select.test.cjs updated for the new runtime count (18)
  + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19.

Out of scope for PR 1 (follow-up PRs in the sequence):
- Install-time decision logic (kimi vs kimi-code detection / prompt)
- agent-install-check semantics for kimi-code (verify Agent Skills presence)
- cmdAgentSkills fallback returning subagent prompt content
- Workflow template mapping (named agents → built-in coder/explore/plan)
- Migration guidance for users currently on 'kimi' who are actually on Kimi Code
- Schema enum extension for subagentToolkit: 'built-in-only'

Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS).

* fix(#2454): complete drift-guard registrations for kimi-code runtime

The drift guards caught every surface that pins runtime enumeration. Each
update is mechanical, driven by the guard's named failure mode:

- src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code
- src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the
  isKimiCode predicate generator
- bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream
  entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19
- gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry
  (null/null/null — same as kimi, no model tier defaults until configured)
- docs/reference/capability-matrix.md: regenerated via
  scripts/gen-capability-matrix.cjs --write (kimi-code row added)
- tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP:
  kimi-code → '.kimi-code'
- tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden
  (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated)

The capability-registry is already regenerated from the prior commit.

* test(#2454): update drift-guard tests for kimi-code runtime registration

Multiple drift guards pin runtime enumeration counts and option numbering.
Each update is mechanical, driven by the guard's named failure mode:

- tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17);
  'all 16 flags' → 'all 17 flags' in test names + messages.
- tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering
  cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15,
  trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice
  test for kimi-code (option 11). Prompt test updated for new numbering.
- tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains
  kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs);
  EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per
  docs, same as Python kimi/opencode).
- tests/global-config-home-fragment.test.cjs: table-count test renamed
  13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier).

* fix(#2454): empty artifactLayout for kimi-code (PR 1 scope)

The skills kind requires a converter (existing converters are per-runtime
like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only
registers the descriptor; the actual Agent Skills converter (and a new
'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside
the install-time decision logic. Empty artifactLayout.global is valid and
means 'nothing to install yet via the layout seam'.

Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs
(localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as
option 11 in install.js's buildRuntimePromptText (renumbered downstream
options 11..17 → 12..18, All 18 → 19).

* fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode)

The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved)
for the new kimi-code runtime id. Property names with hyphens are awkward for
consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new
runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing
the conventional PascalCase flag name. The 16 prior single-word runtime ids
are unaffected (the regex finds no hyphens).

* fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures

- tests/runtime-flags.test.cjs drift guard: use proper kebab-case
  conversion (isKimiCode → kimi-code, not 'kimicode') so the registry
  comparison doesn't false-positive on hyphenated runtime ids.
- tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae
  single-choice tests for the renumbered options (kilo 11→12, opencode
  12→13, pi 13→14, qwen 14→15, trae 15→16).
- tests/install.test.cjs: Kilo integration option 11→12, prompt test
  regex updated.
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
  The kimi-code install produces the standard GSD install layout (skills,
  contexts, references, etc.) — 436 paths, same shape as other runtimes
  that have no custom converter yet.

* fix(#2454): add kimi-code install contract + global config home fragment

- src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code
  → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path
  instead of falling through to the default '.claude'.
- tests/installer-migration-install.integration.test.cjs
  RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for
  PR 1; PR 2 will specialize once the Agent Skills converter lands).
- tests/multi-runtime-select.test.cjs: fix space-separated-choices test
  for the renumbered kilo option (11 → 12).
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: regenerated after rebasing onto current next (new
  planner-reversibility.md from #2471 etc. now included).

* test(#2454): skip kimi-code install contract until PR 2 ships install layout

The end-to-end install test (tests/installer-migration-install.integration
.test.cjs) asserts every allRuntimes entry installs a runtime-specific
artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the
capability descriptor + flags + labels, but the install LAYOUT (Agent
Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands
in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and
self-removing — PR 2 removes the entry alongside adding the install
surface, restoring the contract loop to full coverage.

* fix(#2454): restore compact model-catalog.json format (M1 review)

Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations'
commit used python json.dump(indent=2) which inflated the file from 165→607
lines (every nested entry got expanded) and lost the trailing newline. The
semantic change was just a 3-line kimi-code entry. Restored the original
hybrid format (top-level indent=2 + inner entries' one-line style) and
added kimi-code in matching form.

Regenerated golden install parity + install tree fixtures since the
model-catalog.json hash changed.

* fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code)

CI lint-tests job failed on the glossary drift guard
(scripts/check-glossary-refs.cjs --check):
  ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but
    bin/install.js's allRuntimes array has 18.
  ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js
    (missing from CONTEXT.md's list: kimi-code).

Missed in the prior commits because gsd-test does not run the glossary
check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims
to 18 values + kimi-code in the member list.

* chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511)

* docs(changeset): Phase 1 kimi-code runtime Added (#2511)

* test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands

* docs(changeset): backfill PR #2519 for Phase 1 (#2511)
2026-07-22 00:27:35 -04:00