Commit Graph

855 Commits

Author SHA1 Message Date
0xdhx
ecea537194 docs(#2665): the guard watches config.toml but GSD also writes <root>/hooks/ there
Found pre-push by this round's third adversarial review pass. Not a rebase
regression — round 3 shipped it and #2755 doubled it.

resolveExtraWatchTargets watches one config.toml per non-registry descriptor,
and its comment asserted "GSD writes ONE named file into these third-party
roots". That is false: bin/install.js also calls installSharedHooksBundle on the
same root, populating <root>/hooks/ with GSD's hook scripts and a CommonJS
marker. So a suite-produced leak of a hook bundle into a developer's real
~/.kimi or ~/.kimi-code passes this guard silently — #2665's own hazard, in
#2665's own safety net.

Behaviour is deliberately unchanged and the gap is disclosed instead. Closing it
is a layout decision rather than one more path, for the same reason
getGlobalSkillsBase is already a deliberate non-target: the snapshot applies the
config-root layout beneath every root it is given, and these roots are not ours.
Happy to fix it here or take it as a separate issue — the maintainer's call.

The enumeration-relative test could not have caught this: it asserts one target
PER DESCRIPTOR and nothing about whether one per descriptor is enough, because
its expectation is derived from the same array it checks. That is exactly the
scope boundary round-2 Nit 7 asked to be marked, biting one layer up from where
it was marked; the test now says so.

479979c4's message says "there are three" — that is three WATCHED targets, not a
count of write surfaces. The hooks bundle is a fourth, and unwatched.

lint:ci rc=0; tests/live-config-guard.test.cjs 24/24. Comments and catalog only.
2026-08-08 05:50:49 -05:00
0xdhx
664bab3b48 docs(#2665): the scrub-set and guard-target seams understated their own mechanism
Both predicates were written in round 3 and not updated when round 4 widened
what they describe, so the catalog that exists to stop a future omission had
two of its own.

CONFIG.LOCATION.SEAM.scrub-set said "four sources" and listed four. There are
five rungs: the two descriptor rungs each additionally walk skillsHome.env (the
maintainer's round-4 Minor 3), and the fifth is WRITE_ESCAPE_PERMISSION_ENV_KEYS,
which is a permission rather than a location and so is reachable by no other rung.

LIVE-CONFIG.GUARD.SEAM.non-root-targets said "the two live write surfaces" and
named a single config.toml via resolveKimiHooksTomlDir. Since #2755 landed
KIMI_CODE_HOOKS_TOML_DESCRIPTOR there are three, and the guard derives them by
iterating NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than calling a named
resolver. Both bounds are stated rather than left open: skills bases are a
deliberate non-target (the config-root layout misfires beneath them), and a
further descriptor is only free if it owns the same NON_REGISTRY_OWNED_FILE —
the residual the guard already names at its own definition.

LIVE-CONFIG.GUARD.SEAM.scope's "gsd--prefixed" read as a two-hyphen prefix; the
selector is startsWith(GSD_ARTIFACT_PREFIX) where that constant is 'gsd-'.

CONFIG.LOCATION.SEAM.kimi-two-homes was checked and is NOT stale: #2755 made
Kimi Code a separate runtime, so Kimi CLI still declares exactly two.

Both CONTEXT-INDEX mirrors regenerated; predicate ids unchanged (425), only
their descriptions move.
2026-08-08 05:50:49 -05:00
0xdhx
123ba26fc5 chore(#2665): regenerate docs/CONTEXT-INDEX.json for the round-4 CONTEXT.md edits
The round-4 seam-entry rewrites (packaging fact on SEAM.module, strict-mode
state on SEAM.severity/ci-blind) left the derived index stale, which failed
lint:generated-sync and the #2944 sync test across ten CI contexts. Derived
artifact regen only; no content change beyond what CONTEXT.md already says.
2026-08-08 05:50:49 -05:00
0xdhx
434d71b03b docs(#2665): catalog the config-location and live-config-guard seams in CONTEXT.md
Round 2, Minor. "Workspace seams" carried WORKTREE.SEAM.* and
CONFIG.SEAM.loadConfig-context but nothing for this PR's mechanism or its env
vars, so the one machine-readable place a future author would look said nothing
about the class that has now recurred three times.

Ten predicates across two groups:

  CONFIG.LOCATION.SEAM.*  — the scrub set's four derivation sources and the rule
                            that a new var is made ENUMERABLE rather than
                            appended; the two-families distinction (runtime
                            configHomes vs GSD's own GSD_HOME/GSD_AGENTS_DIR)
                            that round 2 turned on; kimi's two config-location
                            vars; and the in-process scrub requirement, since
                            HOME sandboxing alone is the trap that produced
                            Blocker 1 twice.
  LIVE-CONFIG.GUARD.SEAM.* — module + exports + why it is scripts/ and not
                            scripts/lib/; ownership-based scope; the two
                            non-root targets and their asymmetric treatment;
                            the truncation contract; the report-not-fatal
                            severity ratchet; and that CI is structurally blind
                            here, so green CI is not evidence.

Both generated indexes regenerated. docs/CONTEXT-INDEX.json is checked by
lint:generated-sync (`gen-context-index.cjs --check`) LINE-NUMBER-SENSITIVELY,
and the example's own committed index is separately checked by
lint-example-parser-parity.cjs, which the first regen did not satisfy — editing
CONTEXT.md requires both, and only one of them says so in its error text.

Verified: parity lint rc=0, gen-context-index --check rc=0, full
lint:generated-sync rc=0. Regen diff audited — 10 predicates added, 0 removed,
0 values changed; the example index's remaining churn is line-number re-baking,
which is exactly why the parity lint excludes line numbers.
2026-08-08 05:50:49 -05:00
Tom Boucher
343835facc refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans

scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one
independent re-derivations across seven modules now route through it, and
scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather
than an allowlist (ADR-3180 Decision 4a).

The epic scoped this at three copies. A whole-repo guard found twenty-six sites
across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3
narrows to window plus sentinel enumeration.

Two sites are exempt with a documented reason rather than a bare allowlist:
audit.cts scans one quick task's own directory for a single completion record,
and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time
import. Neither is a phase directory.

scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one
owner answers both questions: verify.cts's numbering-gap check wants every plan
on disk, its pairing check wants the live set. Both fields are additive.

Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave
scheduling, was scheduling status:superseded plans into waves and reporting zero
plans for the post-#3139 nested layout.

filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned
them and only their own tests still called them.

New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator,
with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and
the CONTEXT.md glossary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice

The contract held; the phase boundary did not. The whole-repo drift guard found
26 re-derivations across 9 files against the epic's estimate of 3, and
cmdProgressRender re-derives both enumeration and plan counting on adjacent
lines, so DW4 was unsatisfiable within Phase 1's original file scope.

Records the amended scope, scanPhasePlans's new allPlanFiles field,
findOrphanSummaries, the two documented exemptions, the re-derived Tier-2
table, and the describeNonCanonicalPlans trap for later phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): complete the canonical pairing rule and gate the naming diagnostic

The remote runner went red with 13 deterministic failures on both lanes,
and they were right: replacing verify.cts's canonicalPlanStem pairing with
summaryCandidates dropped a case the bespoke rule covered. A plan carrying a
descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its
canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no
such candidate, so the plan read unsummarized.

The fix is to complete the one rule rather than restore a second:
summaryCandidates gains a canonical-id candidate, narrowed to fire only when an
id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans
and findOrphanSummaries all inherit it. The two-plans-one-summary collision
behaviour of the original rule is preserved deliberately and documented in
place.

Second defect, independently root-caused while verifying: routing the #2893
naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i
fallback, which is correct for counting and wrong for a naming check — a
non-canonically-named file was accepted as a valid plan and the diagnostic
went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect
with a strict isCanonicalPlanFile predicate before reporting names.

Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180:
a question about file naming wants the physical, strictly-matched set; only a
question about outstanding work wants the live set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): register planning-scope.cjs in the eslint migration list

tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor
ignored per its ADR-457 migration state. The new planning-scope module closed
five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md
and the CONTEXT.md glossary - but not eslint, because that one is enforced by a
test rather than by lint:ci, so the local pipeline stayed green while it was
missing.

Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is
linted, matching every other migrated module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): replace the plan-count drift detector with a literal tokenizer

CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the
backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an
escaped \.md". Five review rounds found it had two defects, not one:

  - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.`
    pair be consumed either as one escape or as two class characters, which is
    exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`.
    Excluding `\` from the class killed that but left a cubic path — 23ms at
    N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with
    no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI
    runs on fork pull requests, so a crafted src/*.cts could stall the job.
  - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g.
    `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the
    literal at that `/`, so the scan never reached `\.md` and the guard missed
    it entirely. (Classes holding an ESCAPED `\/` were already matched; the
    tests cover those separately as parity, not as regressions.)

Both defects have one root cause: regex-literal grammar — `\x` escapes, and
`/` inside `[...]` not terminating — is not expressible in a backtracking
regex. So the detector is now a tokenizer, not a regex.

readRegexLiteralAt reads the literal at a given `/` in a single left-to-right
pass with no backtracking, treating escapes as two-character units and
suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch
restarts it at every `/` on the line, preserving the old "find anywhere"
behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the
trailing-flag scan — which keeps the whole-line cost linear.

Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential
shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now
caught. Differential against the old regex over 28,474 lines (those matching
FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/
gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the
old regex hangs): 6 differences, all the tokenizer returning the fuller or
newly-correct literal, 0 old-only misses. The `\.md` token stays
case-insensitive, matching the `/i` the old regex carried.

Also closes three holes in the same new file:

  - walk() tested entry.isFile(), false for a symlink, so a symlinked
    src/*.cts was silently unscanned — an evasion of a guard whose stated
    principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist.
    It now resolves symlinks, but confined: file links must resolve inside the
    repo root, directory links inside the scanned dir itself. Every sibling
    drift guard in scripts/ uses the Dirent classification and never follows
    links, so following them unconfined would have made this the only linter
    able to read outside the tree — on fork PRs an arbitrary out-of-repo read
    whose matched fragments reach a public CI log. The narrower directory rule
    additionally stops `src/up -> ..` from sweeping the whole repo, and the
    skip list is now checked against resolved paths so `src/g -> ../.git`
    cannot reach .git/** or node_modules/**. Real paths are de-duplicated and
    files reported canonically, so a symlink alias cannot shift which
    FUNCTION_SCOPED_EXEMPTIONS key applies.
  - Both the reported fragment and the reported FILE PATH are attacker-
    controlled source text written straight to a CI log, and git permits
    control bytes in a filename. Both are now escaped — C0/C1/DEL plus the
    bidi and zero-width controls — so a crafted literal or filename cannot
    recolour the log, overwrite a line with CR, or fabricate a line that looks
    like this guard's own success output.

Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process
probe over both pathological shapes (catastrophic backtracking is synchronous
and would freeze the suite rather than fail one test), the bare-`/` class
shapes verified to fail against the parent-commit blob, root-confinement tests
covering the outside-file, outside-directory, cycle, broken-link and duplicate
cases, direct isInsideRoot coverage including the sibling-prefix case that a
bare startsWith would let through, sanitizeForReport coverage, and
limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the
exported constant. The earlier structural assertion was dropped — it checked
for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats
while staying exponential.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): backfill changeset PR number

Restores b77931869, which a force-push during the ReDoS remediation dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 01:20:55 -04:00
Tom Boucher
664d49e513 docs(#3182): ADR-3180 — planning semantic model single owner (#3196)
* docs(#3182): ADR-3180 — planning semantic model single owner

Phase 0 design lock for epic #3180. Names one canonical owner per
semantic derivation, specifies the frozen-enum scope contract that
distinguishes a genuinely-empty computation from a truncated or
unscoped one, and locks the drift-guard contract.

Ships no production code. Phases 1-5 execute against this ADR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* test(#3182): prune real issue 3182 from phantom-ref guard, fix empty-list regex

The guard's own header documents that its list rots: entries are phantom
only until the repo's shared issue/PR counter reaches them, and once the
counter passes an entry it must be deleted. Creating the Phase-0 sub-issue
advanced the counter past 3182, so the guard began rejecting a legitimate
citation of a real issue - the failure its header already records happening
twice, with PRs 2551 and 2361.

3182 was the last entry, and removing it exposed a latent bug: the regex
builder interpolated the list unconditionally, so an empty list yields
(?:#(?:)\b)|(?:issues/(?:)\b), whose empty alternation matches every issue
reference in the repo. Following the file's own maintenance instruction
would have turned a green guard into one failing on nearly every file.

buildRefRe() now returns null for an empty list and is exported, with
boundary coverage at 0/1/2 entries plus word-boundary and bare-digit
negative cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 20:25:12 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
sim
c643320cef docs(#3155): ADR-3128 — shipped default is off, not adaptive
The maintainer decided the default after the ADR first merged, by
consistency with the shipped workflow config: verification gates default
on (research, plan_check, verifier, nyquist_validation,
security_enforcement), agent autonomy defaults off (auto_advance,
research_before_questions, plan_bounce, cross_ai_execution). Installing
probes into tracked source without a second confirmation is autonomy,
not a gate.

Adds Decision 8, flips the legacy/absent-section default and the
precedence tail to off, and marks Open question 2 resolved.

Decision 1's justification is corrected rather than deleted. It rested
on 'adaptive carries no flag', which the new default makes false. The
conclusion is unchanged and the real reason is stronger: precedence
includes the saved session policy, so a resumed session that persisted
adaptive passes no flag either, and a flag-keyed atom would exclude the
section from exactly the sessions already running the protocol.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:24:10 -04:00
sim
55136a19e9 docs(#3155): ADR-3128 adaptive runtime evidence — Phase 0 design lock
Records the design decisions #3128's maintainer approval made a
condition: schema v1, the probe/artifact ownership model, and the
cleanup state machine that gates terminal transitions.

Also amends ADR-1671 with a RESERVED atom rather than a widening. The
vocabulary stays at 29 until #3128's implementation lands; the
reservation exists so the widening is a coordinated decision rather than
an organic edit found in review.

The load-bearing decision is the atom's shape. #3128's probe policy is
tri-state (adaptive|force|off), so gating on flag:--runtime-probes would
exclude the protocol section from every default invocation -- adaptive
carries no flag -- and the feature's primary mode could never activate.
That is admission gate (2)'s silent-exclusion failure arriving through a
different door: not a fact nobody computes, but a fact computed for only
one of three policies. The atom is therefore a resolved boolean folded
in cmdInitDebug.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:09:53 -04:00
sim
796bd2cb24 docs(#3149): use the hyphen slash form in reference docs
docs/ is never passed through the install-time slash-form converters, so
the colon form names a command no runtime registers. Caught by
lint-docs-command-form.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:10:23 -04:00
sim
2bead6ca1d feat(#3149): add dedicated init.debug entry point for /gsd:debug
/gsd:debug was one of the last workflows with no cmdInit* of its own: its
Step 0 made three separate round-trips (state.load, resolve-model
gsd-debugger, config-get workflow.tdd_mode) to assemble one context. Because
no debug-scoped fact was computed at any entry point, ADR-1671 admission gate
(2) could never be satisfied for debug — an applicability atom naming such a
fact would evaluate FALSE forever and silently exclude its section.

Adds cmdInitDebug (init.debug), registers it in the init router and the
command-alias table, and collapses debug.md Step 0 to one call. Every field
resolves through the same primitive the call it replaces used: loadConfig for
commit_docs, withProjectRoot for response_language (#2402), planningPaths for
debug_dir, resolveModelInternal for debugger_model, and the existing
Boolean(workflow.tdd_mode) idiom for tdd_mode.

PlanningPaths gains a debug field so state.load and init.debug share ONE
debug-directory expression rather than two kept in sync by hand. state.load
keeps emitting debug_dir: it is a shipped query surface with its own test
anchor, so narrowing it would break unseen consumers for no gain.

No WHEN_VOCABULARY atom and no gsd:section marker: gate (1), a consuming
section of at least 400 bytes, belongs to the change that adds the section.

Closes #3149

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:30:27 -04:00
Tom Boucher
0e6fa2e2cf enhance(#3118): close the dead injectables and the shell projection follow-on — Wave 4 (#3124)
* test(#3118): failing-first coverage for the dead injectables and the shell projection

Adds the counter-tests Wave 4 closes against, before any fix:

- antigravityWatermark had zero test references. The four existing tests
  that look like watermark coverage hand the fallback a literal mark and
  never call the producer, so nothing pinned whether a real run's mark is
  correct. Covers all six branches plus the non-object cache classes.
- Pins the fail-open: a transcript read that throws reports lines:0,
  indistinguishable from a genuinely empty transcript, and the consumer
  then replays a previous run's review as this run's.
- Pins the export-line escaping across the repair, persist and win32 bash
  lanes, including the parity assertion that they must not diverge.
- sliceCurrentPositionSection: empty-vs-absent, fenced heading, second
  occurrence, H3, CRLF.
- Proves deps.progressProvider is inert by supplying a throwing stub to
  all ten transition intents.

Verification through the remote runner only.

Refs #3118

* fix(#3118): distinguish an unreadable transcript from an empty one

antigravityWatermark's final read can throw on a transcript that
indisputably exists. It returned lines:0, which is the same value a
genuinely empty transcript produces, so the caller could not tell the
two apart.

antigravityTranscriptFallback derives its skip from that count. A mark
of {convId:'c1', lines:0} for a conversation that pre-dates the run
makes it skip nothing and return the last PLANNER_RESPONSE in a
transcript written before this run started — a previous review
presented as this one's, which is exactly what the function's own
'never stale' docstring promises cannot happen.

The unreadable case now sets unreadable:true and the fallback declines
for a same-conv-id unreadable mark. An absent or empty transcript is
untouched: those genuinely have zero prior lines.

* fix(#3118): escape the export line for the file it lands in, not the echo

Three lanes emit export PATH="<dir>:$PATH". repair escaped it with
escapePosixDoubleQuoted; persist and the win32 Git Bash lane escaped it
with escapeSingleQuotedShellLiteral instead.

The single-quoting is correct for the echo, so nothing runs when the
user pastes the command. But the bytes appended to ~/.bashrc are the
export line itself, and inside double quotes in an rc file a $(...) or
a backtick in the directory name is command substitution that runs on
every new shell. Those characters are legal in a path on both POSIX and
Windows, so the path was reachable.

projectPathExportLine is now the single source of that line and escapes
for its final rc-file context; each lane still applies its own transport
escaping on top. fish keeps the single-quote escaper — its value really
does stay single-quoted.

The cmd.exe lane interpolated into a cmd double-quoted string with no
cmd-level escaping, so a quote closed the region and &cmd& ran. A quote
is reserved on Windows and cannot appear in a real path, so there is no
correct command to suggest: the win32 lanes now fail closed for one.

Metacharacter-free paths render byte-identically on every lane.

* fix(#3118): drop a stray carriage return and a deps field nobody reads

locateCurrentPosition subtracted a fixed one byte to exclude the newline
before the next heading, which assumes LF. On a CRLF document the slice
kept an unpaired trailing carriage return. It now walks back over the
newline and over a preceding carriage return if there is one.

StateTransitionDeps also required a progressProvider that 33 sites
supplied and no site ever called. A required field nothing reads widens
the module's interface without changing its implementation, which is the
shape epic #3051 cites as its reason for refusing blanket injection.
Removed along with the ProgressRecord alias that existed only as its
return type; state-document.cts's unrelated interface of the same name
is untouched.

* fix(#3118): stop an empty span duplicating bytes, and name the empty results

Three findings from the isolated review pass.

locateCurrentPosition could return end < start when the section was
empty and the next heading followed with no blank line between. Every
mutator splices with slice(0,start) + body + slice(end), so an inverted
span duplicated the region between them — a blank line silently
inserted into STATE.md on every transition, two bytes on CRLF. The span
is now clamped, and an empty section is a zero-length span, which is
what it always meant.

The win32 fail-closed path left the installer printing 'Add it with one
of:' with nothing under it. An empty shellActions folded two different
facts together, so projectPathActionProjection now carries a frozen
PATH_ACTION_REASON and the installer branches on it. Two empty results
with different causes staying distinguishable is the subject of the
epic this belongs to.

fish_add_path parses a leading dash as an option, so a directory named
-v printed 'No paths to add' instead of being added. Verified against
fish 4.8.1: the end-of-options separator fixes it.

Replaces the console-prose test the second fix first arrived with — a
regex over captured stdout is what CONTRIBUTING prohibits, and the
typed reason is the surface it asks for instead.

* fix(#3118): escape TOML control characters, and stop a test name overstating

Five findings from the two review axes.

escapeTomlDoubleQuotedString escaped only backslash and quote. TOML
basic strings also require U+0000-U+0008, U+000A-U+001F and U+007F to be
escaped, so a value carrying a raw newline or NUL wrote a config.toml no
parser accepts — rejecting the whole file, not just that value. Four of
its call sites write real config. Tab stays raw; the grammar exempts it.

The byte-identity test claimed every lane was unchanged for an ordinary
path, which is false: fish now takes the end-of-options separator on
every path, not only hostile ones. Renamed, and the one intended delta
now has its own named test instead of hiding inside a claim that read
as broader than it was.

Also: exact-equality assertions in place of substring checks that could
pass on a subtly wrong escape, newline and null-byte cases for all five
quoting primitives, and a temp dir registered with t.after so it is
removed when an assertion fails.

* docs(#3118): add the changeset fragments

* fix(#3118): degrade instead of throwing on a null conversation cache

A cache file whose whole content is the literal null — what a truncated
or zeroed write leaves behind — made both antigravityWatermark and
antigravityTranscriptFallback throw. JSON.parse('null') succeeds, so the
try/catch wrapping the parse never fired, and resolveConvId then called
hasOwnProperty on null.

Both functions advertise the opposite; the existing test next to them is
named 'a missing cache or transcript degrades to empty, never throws'.
Parsing successfully is not the same fact as the payload being usable,
and a guard that only wraps the parse cannot tell them apart.

resolveConvId is now total for any non-object input, so one guard covers
both callers. Caught by the null case in this wave's own cache matrix.

* test(#3118): correct a stale fish expectation and a parity comparison

The pre-existing 'POSIX persist mode escapes single quotes' test pinned
fish_add_path without the end-of-options separator this wave adds, so it
asserted behavior that is no longer correct. A repo-wide scan found one
such hardcoded expectation; every other site derives its expectation
from the projection.

The new parity test compared the token from a POSIX path against the
win32 lane, which posix-normalizes its input first — two different
inputs, so the tokens differed for a reason that had nothing to do with
the parity it claims to check. It now derives the win32 expectation from
the same input the lane receives.

* docs(#3118): reword a comment the injection scanner reads as an instruction

The scanner pattern act\s+as\s+(?:a|an|the)\s+ carries no word
boundary, so 'the same fact as the payload' matched on the tail of
'fact'. Reworded per the documented remedy for this collision.

The missing boundary is a scanner defect rather than a prose problem —
any contributor writing 'fact as the' trips it — but the pattern is
gate plumbing, which the sibling epic owns, so it is surfaced rather
than changed here.

* chore(#3118): backfill changeset pr number to 3124

* chore(#3118): backfill changeset pr number to 3124

* fix(#2784): make the negation scan single-pass and index it correctly

Three defects in the negation suppression added by #3127, all in one
block, none of which had a test.

The pair scan was verbs.some(nouns.some(...)) with a slice and a split
per pair, so it grew cubically with clause length: 1.1ms before that PR
and 8462ms after, on 800 verb+noun pairs in one clause. api-coverage's
property test generates documents large enough to reach the runner's
600s file cap, which is why it hangs as 'fail 0, cancelled 1' rather
than failing an assertion. Every (verb, noun) window is a subset of the
single widest one, so one scan of that window answers the same question
in a linear pass. Verified equivalent against the old predicate over
20,000 generated clauses.

Both checks also subtracted clause.start from offsets that collectTerm-
Matches already returns clause-local. The first clause on a line has
start 0 so it worked there and nowhere else: later clauses went
negative, and slice reads a negative index from the end, so suppression
silently examined unrelated text.

The comment claimed 'without any API integration' was suppressed. It is
not — the qualifier sits outside the two-word lookback and the noun
precedes the verb. Widening the window would trade a false positive
that costs one declaration line for a false negative that slips a real
integration past a blocking gate, so the behavior stands and the
comment now says so. Pinned by a test.

The qualifier sets were also rebuilt for every line of every document.
2026-08-06 23:57:05 -04:00
Tom Boucher
610ebdebe8 docs(#3043): add caution blocks for --dangerously-skip-permissions (#3121)
* docs(#3043): add caution blocks for --dangerously-skip-permissions

The flag was presented without a caveat in docs/USER-GUIDE.md,
docs/tutorials/onboarding-an-existing-codebase.md, and all four translated
locales. Only the English first-project tutorial carried a proper [!CAUTION]
block. All 10 uncaveated occurrences now carry the same caution block
(optional flag, throwaway/low-stakes use, how to keep confirmations, link
to security model).

* chore(#3043): backfill changeset PR number 3121

---------

Co-authored-by: sim <sim@local>
2026-08-06 10:56:53 -04:00
Tom Boucher
955655407c fix(#2979): document ExitError plain-text carve-out in json-errors.md (#3093)
* fix(#2979): document ExitError plain-text carve-out in json-errors.md

The JSON-errors doc claimed every error emits a structured JSON envelope,
but usage errors (ExitError) intentionally emit plain text with their own
exit code (src/cli-exit.cts:36-39 catches ExitError before the envelope
branch). Anyone following the doc's 'always parse stderr as JSON' guidance
against a usage error got a parse failure.

Amended the Wire format + Overview + Writing tests sections to scope the
structured envelope to non-ExitError failures, stated the carve-out with a
pointer to cli-exit.cts, and scoped the JSON-parse instruction to the
envelope branch. Added a characterization test pinning both paths together
(ExitError -> plain text + own code; non-ExitError -> JSON envelope) so the
code cannot drift toward the doc's prior overstated claim. No runtime
change — the test passes before and after the doc edit.

Re-scoped per maintainer triage: the smart-entry --json part is already
satisfied (shipped payload exposes the command token); only the doc
correction + characterization test remain.

* chore(#2979): backfill changeset PR number 3093

---------

Co-authored-by: sim <sim@local>
2026-08-05 18:57:10 -04:00
Tom Boucher
7203011400 feat(#3072): ship the deferred MCP served catalog (resources + prompts) (#3083)
* test(#3072): add failing-first coverage for the mcp served catalog

55 input-class rows from the phase test matrix, across four suites: the
catalog module over injected readFile/readDir seams, the protocol surface
through handleMessage, the install-vs-catalog parity gate, and fast-check
properties for uri round-trip, traversal refusal, and pagination partition.

src/mcp-catalog.cts lands as a skeleton whose functions throw, so the suites
fail on BEHAVIOR rather than on a missing module. The REASON enum is real so
tests assert typed codes instead of message prose.

Hostile coverage for the one client-controlled path surface (resources/read):
dot-dot and backslash traversal, percent- and double-encoded traversal,
absolute posix and windows paths, file:// scheme, null byte, symlink escape,
unindexed sibling, non-string and empty uri, wrong root segment.

IO faults are injected by monkeypatching the seam, never chmod 0o000 - root
bypasses mode bits, so a permission-based test silently passes with zero
coverage in root CI.

Refs #3072

* feat(#3072): serve the mcp catalog as resources and prompts

gsd-mcp-server now serves GSD's own content alongside its three tools: the
workflow, reference and command tree as MCP resources (resources/list, cursor
paginated, and resources/read over gsd://<segment>/<relpath> uris) and the 71
commands/gsd/*.md as MCP prompts keyed by bare command name. initialize
advertises resources and prompts, and deliberately does not advertise
subscribe or listChanged - the catalog is fixed for a server process lifetime,
so declaring a notification we never send would be a lie a host acts on.

Composition scope is SHARED, not re-declared. shouldCompose lives in
src/mcp-catalog.cts and bin/install.js now imports it instead of carrying its
own regex, so the served catalog and the installed file floor cannot drift on
what gets composed. Proven behavior-preserving across all 2871 tracked paths
plus windows-backslash, absolute and near-miss-prefix cases: zero mismatches.
tests/mcp-catalog-parity.test.cjs asserts served text equals the installer
composition-stage text over the real tree, with anti-vacuity guards requiring
both a marker-bearing workflow and a non-composed file in the comparison set.

Two measurements corrected the literal issue text. Composition is scoped to
gsd-core/workflows/ only, because a reference or command that documents marker
syntax with an unfenced example would otherwise be parsed as carrying a real
marker and have that line lossily dropped. And parity is asserted at the
composition stage rather than against an emitted runtime tree, since install
applies per-runtime path rewrites afterwards and the catalog is host-agnostic,
so byte equality with any one runtime would be false by construction.

resources/read is the one client-controlled path surface and is guarded in two
independent layers: the uri must be an exact key in the prebuilt index, which
defeats every traversal string by construction, and the mapped path is then
re-checked with validatePath so a symlink planted inside a root after indexing
is still refused.

Also fixes a real drift defect found while here: SERVER_VERSION was hardcoded
1.7.0 while the package is at 1.9.1. It now resolves lazily from VERSION or
package.json, reusing the precedent in runtime-artifact-conversion.

Closes #3072

* test(#3072): make the catalog parity gate drive the real installer

Review found the parity gate vacuous: it never imported or spawned
bin/install.js, and recomputed the installer side with the SAME shouldCompose
and composeWorkflow the catalog calls internally. It therefore proved only
that src/mcp-catalog.cts is self-consistent. The old row 52 compared
shouldCompose against a regex literal frozen in the test file rather than
against the installer at all. An inline divergent regex re-added to
bin/install.js - the exact regression ADR-1671 asks this gate to catch - would
have left the suite green.

The gate now spawns a real bin/install.js and compares the composition
DECISION, observed as gsd:section marker survival, against what the catalog
serves for the same files. Marker presence is the right observable because the
installer applies per-runtime path rewrites after composing while the catalog
applies none, so raw byte equality between the two surfaces is false by
construction and must not be asserted.

Sensitivity was proven, not assumed: overlaying the shouldCompose export that
bin/install.js imports so it always returns false makes a real spawned install
leave autonomous.md's markers in place while the catalog still strips them,
and the row 48 assertion diverges.

Anti-vacuity guards are kept and extended - the comparison set must be
non-empty, must contain a workflow that actually carries markers, must contain
a file the predicate declines to compose, and the install must have emitted a
non-zero file count. The marker-documenting reference case has no instance in
the real tree, so it uses an overlay fixture built with the same technique
workflow-fragments-emission.install.test.cjs already uses.

Renamed to .install.test.cjs so it lands in the install suite it now belongs to.

Refs #3072

* test(#3072): retarget the unknown-method assertion off a now-implemented method

tests/gsd-mcp-server.test.cjs used 'resources/read' as its example of an
UNKNOWN JSON-RPC method. The served catalog implements that method, so it now
returns -32602 (no uri supplied) rather than -32601. The remote runner caught
it deterministically on both linux lanes: -32602 !== -32601.

The test's intent is still correct and worth keeping, so it is corrected
rather than deleted or weakened. It now uses 'resources/subscribe', which the
server deliberately does not implement and deliberately does not advertise in
initialize's capabilities, because it never sends the corresponding
notification. That turns the assertion into a real contract - the advertised
capability surface and the implemented method surface agree - instead of an
arbitrary method name a future feature could invalidate the same way.

Swept the rest of the suite for other assertions pinning the newly implemented
methods; this was the only one.

Refs #3072

* chore(#3072): backfill changeset PR number 3083

* test(#3072): make the catalog fake fs separator-agnostic for windows

CI caught this on windows-latest (22 and 24): every catalog fixture indexed
ZERO entries, surfaced by the anti-vacuity guards as 'fixture catalog must
actually index resources for this property to mean anything'.

Mechanism: makeFakeFs keyed its dirMap/fileMap on POSIX-joined paths
(${root}/${rel}), while production buildCatalog looks paths up with
path.join, which is backslash-separated on Windows. Every lookup missed,
tryReadDir returned null, and the catalog came back empty.

Production is NOT at fault and is unchanged. The same CI run proves it: on
windows-latest the real-filesystem tests all passed, including 'installer
composition decision matches the served catalog for every file in the real
installed tree' and the row-51 non-vacuity proof against a real spawned
installer. A real Windows fs accepts both separators; the FAKE did not, so the
fake was the unfaithful one and is what changed.

Lookup keys are now normalized unconditionally with .replace(/\\/g,'/') in
readDir and readFile - never path.sep-conditional, never platform-gated. The
row-42/43 injected-fault wrappers got the same treatment, since they compared
raw production paths against POSIX-literal fixtures.

No assertion was weakened, and the anti-vacuity guards that caught this are
untouched - they are the reason this surfaced as a loud failure instead of a
suite that silently asserted nothing on Windows.

Refs #3072

---------

Co-authored-by: sim <sim@local>
2026-08-05 13:30:55 -04:00
Tom Boucher
481ac7c71b fix(#2946): run milestone complete unstarted-phase guard independent of STATE (#3081)
* test(#2946): milestone complete unstarted-phase guard fails open on STATE desync

Row 1 of the test matrix: the regression test that fails first. Adds seven
cases to tests/milestone.test.cjs covering the desync, absent, no-file,
--force-override, mismatch-WARNING, fresh-project-noop, and sentinel-skip
behaviors. RED on next: the guard's entire scan is nested inside
`if (stateVersion && stateVersion === version)`, so any STATE.md milestone:
value that does not exactly equal the version argument skips the scan with
no warning — functionally an implicit --force on a one-way-door operation.

* fix(#2946): run milestone complete unstarted-phase guard independent of STATE

The entire ROADMAP phase-directory scan was nested inside
`if (stateVersion && stateVersion === version)`, so any STATE.md milestone:
value that did not exactly string-equal the version argument — a desynced
value, or no milestone: field at all — skipped the scan with no warning,
functionally an implicit --force. The operation the guard fronts is a one-way
door: ROADMAP.md and REQUIREMENTS.md are archived and phase directories are
MOVED into .planning/milestones/<version>-phases/.

The scan was already driven by the version argument through
getMilestonePhaseFilter / extractCurrentMilestone; the STATE match was a
redundant second gate that shadowed and broke it. Decouple: the scan now runs
whenever --force is absent, and a present-but-mismatched STATE milestone:
field emits a WARNING naming both values so the suspicious condition is
visible rather than silent. A fresh project with no Phase headings in the
scoped slice still yields an empty scan (no false positives) — the intent the
STATE-match short-circuit was reaching for, now achieved by the scan itself.

* docs(#2946): document milestone complete --force, --dry-run, and the unstarted-phase guard

The CLI-TOOLS reference signature omitted --force and --dry-run entirely,
and neither the unstarted-phase guard nor its override was documented
anywhere user-facing. Add a flags table and a factual guard description
to the Reference page (CLI-TOOLS.md), and a practical guard note to the
/gsd-complete-milestone How-to (COMMANDS.md) covering what to do when
the guard fires and the new STATE-mismatch WARNING (#2946).

American English per CONTRIBUTING.md language policy.

* fix(#2946): emit STATE-mismatch WARNING as JSON in --json-errors mode

Follow-up to the guard decoupling: a structured caller using --json-errors
parses stderr line-by-line as JSON, so the plain-text WARNING would break
such a parser. Honor getJsonErrorMode() and emit a structured JSON object
({ ok, level, message }) in that mode, plain text otherwise — mirroring
io.cts error()'s JSON shape. Addresses the isolated-review observation
(~45% but credible, since --json-errors is a documented CLI flag).

* fix(#2946): address review — drop JSON-mode WARNING scope creep, tighten test assertions

Standards + spec review findings (code-review two-axis + isolated adversarial):

1. The --json-errors JSON WARNING branch (commit dee719404) was scope creep
   the issue never asked for, AND emitted ok:true for a suspicious-condition
   warning (a category error — a stderr JSON parser keying on ok would treat
   the suspicious state as success), AND its comment falsely claimed to mirror
   io.cts error()'s {ok:false,reason,message} shape. Dropped: the WARNING is
   now plain-text stderr only, matching the existing [gsd-tools] WARNING
   convention (state.cts). The issue asked for 'at minimum warn', not a
   structured JSON surface.

2. The WARNING test asserted on /WARNING/ regex (raw-text matching on stderr
   prose). Tightened to assert on the stable operator-facing tokens — the
   WARNING: marker and both version literals the operator must see — not the
   surrounding formatter prose. Added a paired negative test confirming no
   WARNING is emitted for an absent milestone: field (a missing declaration
   is a normal fresh-project state, not suspicious drift).

* fix(#2946): sanitize STATE milestone value before stderr WARNING interpolation

Security review (minor): stateVersion is read from a user-controlled file
(STATE.md) and is not validated like the CLI version arg. Sanitize before
interpolating into the WARNING — strip ANSI/control chars
(/[\x00-\x1f\x7f]/g -> '?') and truncate to 80 chars — so a corrupted or
hostile STATE.md cannot echo terminal escapes or secret-looking strings
verbatim into a CI log or terminal aggregator (CONTRIBUTING.md security:
secret-looking values in stderr). version is already constrained to
[A-Za-z0-9._-] by ARCHIVE_VERSION_LABEL_RE upstream, so it needs no
sanitization.

* test(#2946): correct stale fixtures that relied on the guard being silently disabled

Four pre-existing tests broke under the #2946 fix because their fixtures
only passed thanks to the bug — the unstarted-phase guard was skipping on
STATE mismatch, so fixtures with missing or non-matching phase directories
slipped through. The tests exercise version-forwarding / version-scoping,
not the guard, so give them legitimate directories:

- milestone.test.cjs #3043: dirs were '103.old'/'104.old'/'108.new' (dot),
  which phaseTokenMatches rejects — renamed to hyphen form. The v3.6 stats
  scoping still yields 1 phase (getMilestonePhaseFilter scopes correctly);
  the guard now sees all three phases as having directories.
- milestone-archive.test.cjs 'returns version in response data': ROADMAP
  listed Phase 1 but no directory was created. Added 01-foundation so the
  scan is satisfied.

Per CONTRIBUTING.md, test-fixture corrections land as their own test:
commit, not bundled into fix: (release hotfix cherry-pick routes by prefix).

* chore(#2946): backfill changeset PR number 3081

---------

Co-authored-by: sim <sim@local>
2026-08-05 10:42:42 -04:00
Tom Boucher
d2e727d3b3 docs(#3074): correct adr-1671 mcp citation and stale runtime counts (#3080)
ADR-1671 attributed its MCP deferral to "ADR-857 §7 / #956" at five sites.
Neither source supports it: docs/adr/857-capability-system.md contains zero
MCP references (its Decision 7 is third-party code-loading, Decision 8 is
Runtime-as-Capability), and #956 is the closed first-party MemPalace plugin
pre-proposal that ADR-1239 explicitly disclaims in its own header.

The deferral itself is sound on ADR-1671's own runtime-partial reasoning and
never needed the borrowed citation. Ground it there, cross-reference ADR-1239
as the ADR that owns GSD's MCP surface, and record that a companion MCP server
shipped 2026-06-28 with three tools - so "MCP is deferred" is not misread as
"GSD has no MCP server". The deferral narrows to the served resources and
prompts catalog (#3072).

Also record that "deferred-tools", named alongside resources and prompts, is
not a deferred surface but an unbuildable one: MCP defines three server
primitives and the tools surface is tools/list plus tools/call, so schema
deferral is host behavior, not a server capability (#3075).

Correct two stale counts: 15 runtimes -> 19 (of 44 capability descriptors).

The ADR-857 reference in Open questions is a genuine Phase-6 completion
property and is deliberately left untouched.

Closes #3074

Co-authored-by: sim <sim@local>
2026-08-05 09:10:55 -04:00
Tom Boucher
8f75e27554 fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag

Every isolation gate already resolved correctly. The resolved value then reached
the executor through a prose instruction telling the model to substitute it into
a call the model composes itself, and nothing verified the substitution. When it
was dropped, the executor edited and committed in the user's primary checkout
with no consent and no warning.

A prose backstop would be the same class of artifact as the defect, so this is a
shipped PreToolUse hook on the Agent tool. It fires at the instant of the call
rather than being read once at the top of a workflow, which is the only placement
the model cannot skip.

The guard is inert unless it can positively establish that this is a GSD project,
that the project resolves to harness isolation, and that the dispatch targets an
executor. A non-GSD repo has no invariant to enforce. Where it cannot read the
configuration at all, it denies rather than assuming, with its own reason -- a
guard that cannot verify must not answer safe. A malformed payload allows rather
than throwing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3045): extend the isolation guard to Cursor

Cursor is the second of only two runtimes that resolve harness isolation, so
shipping the guard for Claude alone left half the exposed surface unguarded while
the changeset implied it was covered.

The two runtimes fail differently. On Claude the harness flag is a per-dispatch
kwarg the model must copy into a call it composes, and the defect is that it can
be dropped. On Cursor the flag is --worktree, which applies to the whole session,
and the subagent-start payload carries no isolation field at all. There is no
flag to check, so the guard verifies the effective state instead: whether the
workspace is genuinely running outside the user's primary checkout. That is a
stronger check than the Claude one because it tests reality rather than intent,
and it is commented so nobody later rewrites it into a flag check.

Isolation is established two ways, either sufficient: the workspace resolves to a
linked git worktree, or it sits under the worktree root Cursor manages. The
second matters because a directory Cursor placed there is a legitimate isolated
session even before it becomes a distinct git worktree, where linkage alone would
report no repository.

Detecting linkage required a new primitive rather than the existing context
resolver. That resolver short-circuits on finding a local .planning directory
before it ever compares the git directory to the common one -- and an isolation
worktree normally has its own checked-out .planning. Reusing it would have read a
correctly isolated session as unisolated and denied it, which is the failure
direction that gets a guard switched off. The comparison is now its own
shortcut-free function that the resolver delegates to after its own shortcut, so
existing behavior is unchanged, and the case that would have broken is pinned.

The subagent type is checked before any configuration is read, so an unreadable
config cannot deny a dispatch this guard would never have enforced against.

The input-schema comment on the Cursor hook documented only the fields common to
every event and omitted the ones specific to this one. That omission cost a
halt during this work; it now documents both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): enforce the resolved dispatch decision, not the host capability

The guard keyed on the registry's dispatch.isolation, which says only that a
runtime is CAPABLE of harness worktrees. The decision that actually governs a
dispatch is the one the workflow resolves after gating, and that legitimately
comes out as sequential in three documented cases: a project setting
use_worktrees false, a per-plan submodule intersection, and the base-check
auto-degrade. The workflow tells the model to omit the flag in exactly those
cases, and the guard was denying every one of them.

The third case matters most. The preceding fix made the base-check degrade on
git timeouts and a missing git binary, where it had previously answered "safe".
That correction is right, and it means a transient hang now degrades to
sequential far more often than before -- so the two changes composed into a trap
where the workflow behaved exactly as designed and the guard blocked it.

The workflow already resolves isolation in shell, deterministically, which is
what makes it a trustworthy source in a way the model-authored call is not. It
now records that resolved value through a dedicated verb, and both guards read
it first. A fresh record is authoritative, so sequential dispatches pass
untouched. Absent or stale, the guards fall back to the capability check
combined with the project's use_worktrees setting, which still covers the case
that never reaches the workflow.

Also widened the matcher to accept Task alongside Agent, since a host that names
the tool Task would otherwise leave the guard silently inert while implying
coverage; stopped assuming Claude when no runtime is declared, which is the
shipped default and would have demanded a Claude-only argument elsewhere; and
made a non-git project inert rather than denied, since advising a worktree
session is not actionable without a repository.

The original diagnosis never modeled sequential mode as legitimate. That
omission is what let this through, and it is now recorded there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): record at resolution and bind the record to its dispatch

Two independent reviews converged on the same failure: the guard was fail-open in
a default install, so it did not catch the defect it exists to catch. A shipped
project carries no runtime key, which made "runtime not confidently known" the
common case rather than a corner one. A record asserting that isolation was
required but carrying no flag then fell through to a capability lookup that
answered "none", and the dispatch was allowed. The flag itself only arrived from
a second shell block -- the same block a model dropping the argument would also
skip. A test had pinned that behavior as intended.

The record is now written by the resolver, as an unavoidable consequence of
asking for the value, rather than by a step the model is told in prose to go and
run. A guard against a prose-carried value cannot itself depend on prose. Mode,
flag and identifiers are written together and atomically, so the flagless window
is gone, and a record asserting isolation with no resolvable flag now denies
instead of degrading. Runtime is also resolved from the installer's own recorded
default, which makes confident resolution the normal case.

The per-plan submodule gate degrades after the phase-level decision and never
re-recorded, so a plan that legitimately ran sequentially was denied against a
still-fresh phase record. It now records its own, scoped to the plan.

A record also authorized any dispatch for four hours. One phase degrading to
sequential could silently license an unisolated dispatch in the next. Records
now carry phase and plan, the guards require them to match, and the window is
minutes rather than hours -- the resolver rewrites it before every dispatch, so
a long window bought nothing and only widened the hole.

The flag validator rejected any value beginning with two dashes, which is exactly
the form Cursor and Windsurf declare, so their real value could never have been
stored. Writer and reader also derived the record path differently and diverged
inside a linked worktree without local planning state.

The predictable path remains a way to silence the control without leaving a trace
in the diff. It grants no access an agent with shell does not already have, so it
is documented as accepted rather than redesigned around.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): correct the staleness boundary and unmask a vacuous parity test

The remote runner returned twenty failures. One was a real production defect the
boundary case existed to catch: a record whose age exactly equalled the staleness
window was treated as fresh, so it stayed authoritative for one tick past its own
expiry. Freshness is now strictly inside the window.

The parity test meant to stop the two guards' executor lists from drifting could
never have failed. Its project fixture was a bare directory rather than a
repository, so the non-git inert branch answered before the executor list was
ever consulted. It asserted agreement it never actually measured. The fixture is
now a real repository, like every sibling in the file.

A test also asserted that Windsurf declares the worktree flag. It does not --
Windsurf resolves to no isolation by design, having no named concurrent dispatch
to isolate. The test claimed a registry fact that was never true, and a comment
in the resolver repeated it. Both corrected, and the test now proves what it
should have all along: that the parser accepts any bare flag value, rather than
one runtime's supposed value.

The new guard was missing from the bundled-hook whitelist, which is the surface
that decides what actually ships, and the per-plan gate had gained calls to the
launcher without the preamble those calls require. The changeset carried
parenthetical product descriptions the purity rule forbids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3045): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3045): make the guard tests hold on Windows

Two tests redirect HOME to control where the installer-persisted runtime default
is read from. Node resolves the home directory from USERPROFILE on Windows and
never consults HOME, so both silently read the real runner profile, found no
recorded runtime, and asserted against a project the hook had not recognised. The
production code was already correct in asking the platform rather than the
variable; only the tests were wrong to assume one variable answers everywhere.
The helpers now mirror the override onto both.

The symlink spoofing test also created a directory symlink unconditionally, which
needs elevated privileges on Windows. It survived on this runner, but it would
fail on any host without them, so the creation is now attempted and the test
skips explicitly when it cannot be done -- a bare return would have counted as a
pass and hidden the gap.

Skipping alone would have left the platform uncovered, so the behaviour it proves
is now also driven in-process through an injected realpath, following the seam
already used for the clock. That case no longer depends on privileges at all, and
the end-to-end test keeps its original assertions wherever symlinks work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:42:16 -04:00
Tom Boucher
c899f5ada3 chore(#3065): build the deterministic load-bearing-fragment contract gate (#3068)
* test(#3065): build the load-bearing contract gate ADR-1671 promised

Epic #1671 Phase 7. A post-merge audit of every promise in ADR-1671 against the
merged tree found one mitigation asserted-but-absent and two stale records.

ADR-1671 names exactly one correctness risk — trimming a load-bearing fragment,
with the recorded history of a paraphrased META.RULE causing agent violations —
and #2931 amended its mitigation to a deterministic contract gate that proves no
load-bearing fragment was omitted or shrunk, treats a floored fragment as a
success, and asserts the isolate prefix survives byte-identical, with an explicit
anti-vacuity rule.

That gate did not exist. What existed was tests/context-composer.test.cjs:
synthetic unit tests of the composeWithinBudget primitive over invented
fragments, asserting nothing about real declared strategies. The ADR asserted a
mitigation that was never built, which is the promised-but-not-built shape the
epic's own coverage discipline exists to catch.

The gate derives its load-bearing set from declared verbatim strategies rather
than a hand-maintained list, so it cannot go stale as upstream changes. It sweeps
budgets from 4x total down to a quarter of total and asserts at every step that
no load-bearing id appears in omitted or shrunk, that isolatePrefix is
byte-identical, and that hardFailed is surfaced rather than silently passed.

Both anti-vacuity guards are EXECUTABLE, not comments. One proves the empty
load-bearing set guard actually throws. The other proves a sweep that never
applies pressure is rejected — because a gate that only ever runs unpressured is
exactly how the original mitigation went missing without anyone noticing.
Measured: underPressure true at 6 of 7 budgets, false only at 4x total.

Three ADR records corrected in the same change, all doc-vs-reality drift:

  - Decision item 2 describes a composer that trims by priority to fit a measured
    per-runtime cap. composeWorkflow in fact passes MAX_SAFE_INTEGER with every
    fragment verbatim (both verified in source), so no trimming happens there;
    the emitted-byte cap is a separate measure-and-fail gate and Windsurf's limit
    a bespoke truncation. The wording described an option as shipped behavior.
  - flag:--converge never reached a terminal state. #2992 withheld six atoms;
    five were resolved explicitly. This one was resolved in code by reusing
    state:plan-strategy-converge but recorded nowhere — the same gap #2995 closed
    for flag:--verify-only, and I closed five of six.
  - The open-questions list enumerated three questions while two Resolved-by
    blocks resolved an unlisted Question 4. It is now listed.

Refs #3065

* fix(#3065): make the gate assert over production, not a copy of it

The isolated review found a blocker, and it was fatal to the gate's purpose: it
hand-copied applyBudget's fragment array into the test, so flipping a strategy in
src/prompt-budget.cts — say roadmap from verbatim to drop — would leave the gate
computing from its own untouched copy and still passing. A guard built as an
instance of the very divergence class it exists to prevent
(DEFECT.GENERATIVE-FIX) is worse than no guard, because it reports green.

Fixed by eliminating the duplicate rather than adding a parity assertion, the
same resolution used for the FAMILIES table in #2996. applyBudget's inline
construction is extracted to an exported buildBudgetFragments(), which both
applyBudget and the gate now call; the 1024 plan floor is exported as
PLAN_FLOOR_CHARS instead of being re-declared in the test. The extraction is pure
— verified behavior-preserving at budget=2000: hardFailed false, omitted
['context'], projectMd shrunk, plan truncation ~27.8%, all headers present. There
is no longer a second copy to diverge from.

Also fixed a vacuous assertion the same review caught: isolatePrefix was pinned
across the sweep, but no production fragment sets isolate:true, so the value is
always '' and the check could never fail. The pinning assertion stays, with an
honest comment that nothing in production sets it today, and a second test now
constructs an isolate:true fragment set and proves the prefix is non-empty and
byte-identical across a roomy and a severely tight budget — which is what makes
the first assertion capable of detecting a real change.

Refs #3065

* chore(#3065): backfill changeset pr number to 3068

---------

Co-authored-by: sim <sim@local>
2026-08-04 22:41:01 -04:00
Tom Boucher
da062c0e0d chore(#2996): inventory the workflow fragment tree as its own manifest families (#3061)
* feat(#2996): inventory the workflow fragment tree as its own families

Epic #1671 Phase 6.5, the epic's last deliverable.

47 step files across 15 workflows and 13 mode files were invisible to
docs/INVENTORY-MANIFEST.json. Not through a missed row — through construction:
buildManifest walks each family with a flat readdirSync + isFile() and never
recurses, so nothing under gsd-core/workflows/<wf>/ could ever appear. modes/
has been invisible that way since #717 without any gate firing, which is the
evidence that this is a generator gap rather than someone forgetting a row.

Two new families, workflow_steps and workflow_modes, keyed by
<workflow>/<subdir>/<file> rather than a bare basename. That is deliberate: two
workflows may each own a regression-gate.md, and a step file may share a name
with a top-level workflow. The manifest is compared by JSON equality, so a
basename collision would silently drop an entry and read as "up to date".
Recursion is bounded at exactly one named subdirectory, and a limit+1 test pins
that bound so it cannot quietly become a general walk.

tests/inventory-manifest-sync.test.cjs carried its OWN duplicate copy of the
FAMILIES table — the DEFECT.GENERATIVE-FIX divergence class. Adding a family to
the generator alone would have left that test verifying six of eight families
while still reporting green. The table now lives once in the generator and is
imported, so the two surfaces cannot drift; runMain is guarded behind
require.main so importing does not execute the CLI.

The per-file roster stays in the generated manifest rather than being copied
into INVENTORY.md: 60 hand-maintained rows in lockstep with a generated artifact
is precisely the drift this file exists to catch.

CONTEXT.md's RULESET.MANIFEST-CANONICAL-KEY and DEFECT.INVENTORY-DRIFT both said
"six families" and now say eight, with the two key shapes and the import rule
recorded. The non-shipping example index was regenerated for the same edits.

Note on scope: this issue also asked for a one-fragment-edit proof. That landed
independently as PR #3046 and is not rebuilt here.

Refs #2996

* fix(#2996): correct a fabricated roster and an inert coverage pragma

Isolated review returned one blocker and three lesser findings. All four were
real; all four are fixed.

BLOCKER — docs/INVENTORY.md claimed the workflow_modes roster was
"discuss-phase, sketch". There is no gsd-core/workflows/sketch/ and never has
been; the second member is `help` (4 mode files), exactly as the manifest
generated by this same diff already listed. A doc contradicting the manifest it
describes, in the PR whose whole purpose is closing doc/reality drift. The
adjacent hand-maintained "15 workflows" count is also removed: an unenforced
number in a table cell is the same staleness class this file exists to catch,
and no test guards table-cell counts.

MAJOR — the CLI entry guard carried `/* istanbul ignore next */`, which excludes
nothing here. This repo measures coverage with c8 (test:coverage:scripts-floor,
55% floor over scripts/**/*.cjs), and c8/v8-to-istanbul honors only
`/* c8 ignore next */`. The pragma looked like it was doing something and was
not — the same failure shape as a marker that looks like working gating.

MINOR — collectNested called statSync/readdirSync unguarded, so a dangling
symlink or an EACCES directory under any workflow's steps/ would throw uncaught
and red the manifest gate for the entire repo. An entry that cannot be statted
is, for inventory purposes, not a countable file — the same disposition as "not
a directory". Row 13c pins the behavior with a real dangling symlink.

Refs #2996

* chore(#2996): backfill changeset pr number to 3061

* test(#2996): guard the dangling-symlink row on Windows

fs.symlinkSync throws EPERM on Windows without elevation or Developer Mode, so
row 13c would red the Windows lane. Guarded with the repo's idiom — a
process.platform check plus a genuine t.skip() carrying its reason, never a bare
return, which node:test counts as a PASS and would hide the gap.

Worth recording why this was not caught here: CI classified this PR's diff as
inert (no bin/, gsd-core/, or src/ changes), so the full test matrix was SKIPPED
entirely — the 'full test (${{ matrix.os }}, ...)' job shows as skipping with
its matrix expression unexpanded. The Windows lane never ran. It would have
fired on the next PR that does touch core code, in someone else's change.

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:51:12 -04:00
Tom Boucher
ed360cd99f chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point

Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers
only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim
into every runtime — and agent text is loaded into a subagent's context on every
dispatch.

The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op
for agents: agents never traverse that function. Agent content is read for
emission at five independent points, and the obvious chokepoint
`stageAgentsForProfile` short-circuits on the DEFAULT `full` profile
(`skills === '*'` returns the real unstaged directory), so a hook placed there is
dead code on most installs.

Composition now happens at two call sites instead of five parallel surfaces:
`stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind`
routed through it via an identity converter) and the inline agent loop in
bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` ->
`.windsurf/` regex can never reach inside a marker attribute — the ordering
#2930 established for workflows.

`installCodexConfig` was the fifth read point: Codex embeds each agent's prompt
into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it;
the exhaustive per-runtime emission sweep found it. That is why the new guard is
behavioral rather than structural — a sixth read point fails the sweep without
anyone remembering to extend a list.

tests/agent-fragments-emission.install.test.cjs spawns a real installer for every
runtime at every agent-bearing scope, derived from RUNTIME_META and the
capability registry at run time so a new runtime cannot be silently
under-covered. It asserts markers are absent AND the `when="always"` body is
retained, so marker-absence cannot be satisfied by dropping content. An
identity-composer negative control proves the assertion can fail.

Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope
paths; red before the wiring on claude(global+local), zcode(global+local),
kimi, codex and opencode.

Refs #2995

* chore(#2995): give the tightest agents headroom and correct the design lock

Epic #1671 Phase 6.4, second half.

`agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE
cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now
extract reference material to `gsd-core/references/` behind an @-reference — the
documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy:

  gsd-verifier  49,140 -> 46,371 B   headroom    12 -> 2,781
  gsd-debugger  57,197 -> 48,851 B   headroom   147 -> 8,493

Byte accounting proves no content was lost: the combined agent+reference delta
is exactly the new files' headers plus the agents' slim replacement blocks. Each
agent keeps its routing table and a one-line summary per entry, so it degrades
gracefully on a runtime that does not inline @-references.

`agents/gsd-planner.md` is untouched and still passes both char guards
(49,130 < 49,152); it needed no change, so it took none.

The other nine LARGE/XL agents carry NO gsd:section markers, and that is
deliberate, not deferred. `when=` selection is read from
gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives
from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key,
no per-agent init entry point. An agent atom therefore fails admission gate (2)
("a fact the init seam demonstrably computes at a real entry point") and would
evaluate false forever while looking like working gating. Marking agents would
manufacture exactly the silent-inertness rot the frozen vocabulary exists to
prevent.

ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage
found against what actually merged:

  - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR
    amendment, which that bullet's own rule forbids. Recorded now.
  - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to
    "the LARGE/XL rollout phase". Five shipped; this one is permanently
    rejected, and that disposition lived only in a merged PR body.
  - Phase 6.4's own finding: emission extends to agents/, gating does not.

CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still
listed the original 4-atom vocabulary and described when= as "not yet acted on",
and Section Manifest Module still described InvocationFacts as
{waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract.

Inventory manifest regenerated AFTER build:lib per the documented ordering
landmine; 19 install-tree fixtures pick up the two new references.

Refs #2995

* chore(#2995): correct the compose-site count and mark the raw stager

Self-review found two comment defects in the prior commit. The agentsKind
comment claimed composition lands at TWO call sites; it is three, since
installCodexConfig's per-agent .toml writer was added after that comment was
written. And stageAgentsForProfile is now production-dead — both callers route
through the composing stager — while staying exported and unit-tested, which
makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged
source directory under the default profile, so a future caller would silently
reintroduce the marker-shipping path. Its JSDoc now says so.

* test(#2995): guard the marker-documenting-doc class for agents

Widening the composer's scope to agents/ makes reachable the exact class #2930
narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an
unfenced example is indistinguishable from a real marker, so the composer drops
that line from the emitted artifact.

Three rows. A fenced example must compose byte-identically. No shipped agent may
carry a marker outside a fence — asserted by parsing every real agent and
requiring zero explicit sections, which is what makes the fence protection
load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED
marker IS parsed as a real marker, so if that ever stops being true the second
row is guarding nothing.

Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it
had no production caller, which is false — bin/install.js's _stageAgents still
calls it, and its consumers compose before writing. Corrected to state the
invariant instead. And a let/const nit in the emission sweep.

* fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture

The first remote run came back red with three failures. Both root causes were
mine.

1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally
   contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status
   vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent
   no longer had it.

   Byte accounting said no content was lost, and byte-wise that was true — but a
   contract required those tokens to live IN THE AGENT. That is ADR-1671:66's
   flexReserve floor stated concretely: a load-bearing fragment must not be
   trimmed out of its host, and "the bytes still exist somewhere" is not the
   test. The two status tables are restored to the agent and deliberately
   mirrored in the reference with a note saying so, so the procedure there still
   reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083,
   rather than the 2,781 the first attempt claimed.

2. Row 12b of the new marker-documentation guard asserted that an unfenced
   marker example parses as a real marker, and threw instead:
   "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The
   fixture had put the OPEN marker inline mid-sentence, so it was correctly not
   recognised as an open while the close, on its own line, was.

   That is a real refinement of the hazard this guard exists for: only a marker
   on its OWN line is mis-parsed — which is exactly how a documentation example
   is normally written. Row 12b now uses a whole-line marker, and a new row 12c
   pins the inline case as explicitly NOT a marker.

No test was weakened to accommodate the change; the change was corrected to
satisfy the tests.

Refs #2995

* chore(#2995): backfill changeset pr number to 3058

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:10:31 -04:00
Tom Boucher
4eb8e3648c fix(#3050): consolidate the spawn-timeout predicate and propagate the unresolved-root reason (#3060)
* chore(#3050): changeset and review artifacts for the follow-up

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3050): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 17:21:24 -04:00
Tom Boucher
ffd5370464 fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs

Docs told readers to type the colon form, which no runtime registers -- 18 of
19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an
example got an unrecognized command. Swept 178 occurrences across 53 files,
locale mirrors included so they do not re-diverge from English.

The colon form is a source-authoring token, not a user-facing one: install-time
converters key on it to produce the hyphen form runtimes actually register. So
the sweep is scoped, and three things are deliberately left alone:

- ADRs, which are a historical record; editing their prose falsifies what was
  written at the time.
- The legacy release-notes archive, pending a maintainer decision on whether it
  follows the same historical carve-out. Excluding it keeps a later reversal
  additive rather than a revert.
- Source artifacts under commands, workflows and agents, where the colon form is
  load-bearing. Rewriting those would break the installed-skill guarantee across
  every runtime -- the single largest hazard here.

The plugin namespace form is a real, separate token and survives untouched.

Adds a lint enforcing exactly that boundary, since the correct form genuinely
differs by directory and nothing previously caught the drift.

Also fixes a hardcoded colon form in the capability-matrix generator. The sweep
alone would have left the generated matrix disagreeing with the template that
produces it, so the fix is at the source and the output regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): stop the sweep misquoting source frontmatter

Adversarial review caught three lines where the sweep rewrote a citation of the
literal YAML name: key from a source command file. That key genuinely is the
colon form -- this change's own carve-out logic says source-authoring tokens keep
it -- so the docs ended up misquoting the real files. One of the three is an
acceptance-checklist assertion, which the sweep turned into a false statement.

Restored the three citations to match their sources verbatim, surgically: where a
line carried both a name: citation and a real reader-facing slash command, only
the citation reverted and the command stayed corrected.

The guard needed the same distinction, or it would have flagged the restoration
and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a
source token and is now permitted. The exemption is deliberately narrow -- a bare
gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness.

Also makes the detection case-insensitive. Review found /GSD:next slipped through
silently; no such casing exists in the tree today, so this closes a latent gap
rather than fixing a live one.

Swept the whole tree for further corrupted citations: none beyond the three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): retire the stale-next invariant and sweep next like every other command

Maintainer decision on a genuine conflict between two contracts.

Invariant #3054 banned the literal /gsd-next from user-facing docs because it
named a retired workflow-advance command. But commands/gsd/next.md is a live
command -- the state-aware smart-entry launcher -- and this issue requires docs
to use the hyphen form every runtime actually registers. Both could not hold for
this one command, so docs had been sidestepping the ban by keeping the colon
form, which is exactly the defect this issue exists to remove.

FEATURES.md already recorded the reassignment: the hyphen form "is not the
retired workflow-advance command; it is reserved for the state-aware smart-entry
launcher. Workflow advancement remains under /gsd-progress --next." With that
reassignment the invariant's premise is obsolete and the guard now contradicts
the documented command form, so it is retired with a comment recording why
rather than deleted silently.

next is now swept like every other command, and the earlier exemption added to
the new guard is removed so nothing is special-cased.

Four citations of the literal name: frontmatter key stay in colon form, because
the source file really does carry name: gsd:next and a doc quoting it must
reproduce it verbatim. Two of those lines were reworded to say which side is the
frontmatter key and which is the slash command, since they previously conflated
the two.

Verified the retired scan would now genuinely fail against this tree -- the
conflict was real and resolved, not dodged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2903): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 13:23:44 -04:00
Tom Boucher
ef823ca9d9 fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking

Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct,
transitive (2 and 3 hop), and diamond dependents of a halted plan across
both independent "which plans are incomplete" readers (phase-plan-index's
cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative
case (an unrelated decoupled plan stays runnable), and a parity check that
the two readers agree. Uses only modules that already exist at this
commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs)
so the test file loads and runs cleanly on a fresh clone of this exact
commit. These fail against current behavior: neither reader has any
concept of a halted plan or a blocked_by/runnable view yet.

* fix(#2830): a halted plan no longer leaves its dependents on the runnable work list

A plan that reaches a designed stop still writes a SUMMARY, so both
"which plans are incomplete" readers saw it as an ordinary completion and
reported its dependents as ordinary runnable work — never checking
whether an upstream plan had halted rather than finished.

- New `status: halted` frontmatter value, documented in all four SUMMARY
  templates alongside the existing `status: complete`.
- New shared src/plan-dependency-graph.cts: a single computeHaltPropagation
  pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and
  phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50
  dependent symbols across 5 command routers) now call, so the
  two-implementation divergence that caused this bug cannot recur. It
  accepts an optional precomputedOrder so cmdPhasePlanIndex — which already
  runs Kahn's algorithm in computeDependencyLevels for wave assignment —
  passes that order straight through instead of a second traversal;
  searchPhaseInDir (no prior traversal) lets the module derive its own.
  The two small duplicated predicates each reader would otherwise carry
  (is this status "halted"?, which summary file matches which plan id?)
  are centralized in the same module as isHaltedStatus/buildSummaryFileIndex.
- Additive fields only: `halted`/`blocked_by`/`runnable` on
  cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/
  `runnable_plans` on searchPhaseInDir's result. The pre-existing
  `incomplete`/`incomplete_plans` fields are unchanged in meaning and
  membership.
- execute-phase.md's discover_and_group_plans step now also skips any
  plan whose `blocked_by` is non-empty, reporting it by name with its
  blocking chain, in addition to (not instead of) the existing
  has_summary skip rule.

Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the
prior commit) with direct unit coverage of computeHaltPropagation
(including the precomputedOrder call shape) and a fast-check property
test — both only possible once this commit's new module exists.

Closes #2830

* fix(#2830): surface the halt-aware view from init execute-phase

The adopted work made phase-locator compute halted_plans / blocked_by /
runnable_plans, but cmdInitExecutePhase builds its output by explicitly
enumerating fields, so all three were computed and then silently dropped at
the exact consumer the issue names as regressed.

Forwards them additively -- incomplete_plans and incomplete_count keep their
name, type and semantics byte-for-byte -- and adds the same three empty
defaults to the roadmap-only fallback so the shape is consistent in both
branches. Covered by a new test that drives the real CLI end to end rather
than the locator function, since the locator already worked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect

Three review findings, all fixed:

- BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in
  the Kahn pass, so they were excluded from the topological order, never visited
  by the forward pass, and vanished from blocked_by entirely -- i.e. reported as
  runnable. The wave-grouping reader hard-fails on a cycle so it never hit this,
  but the phase-location reader does not, so init execute-phase offered a plan
  depending directly on a halted plan. Reproduced, then fixed in the shared
  engine so every consumer is safe regardless of pre-checks: a node absent from
  the order is now blocked with a deterministic, non-empty named cause. A plan
  silently missing from both blocked_by and runnable is the exact disappearance
  this issue exists to prevent.

- MAJOR (isolated adversarial). All four summary templates showed the field as
  an inline comment on the value line. Frontmatter parsing does not strip
  trailing comments, so an executor copying the templates' own presentation
  wrote a halt that parsed as a non-halted string, silently reproducing the
  original bug. Guidance moved off the value line, and the halt predicate now
  tolerates an unquoted trailing comment.

- HARD standards violation. A test regex-matched child-process stderr prose for
  /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal
  plus a differential assertion (same fixture without the cycle edge must
  succeed), so it stays cycle-specific without matching prose.

Also folds the duplicated read-summary-and-check-halted wrapper out of both
readers into the shared module -- centralizing only the predicate left the exact
two-copies-that-drift pattern the module exists to prevent -- and commits the
artifact-types documentation for the new status value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2830): stop the property generator hanging the whole suite

The remote runner did not fail -- it hung. Two containers sat in this file for
31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner
passes --test-timeout=0, so nothing ever reaps it: this would have hung CI
indefinitely, not reported a failure.

Root cause: the DAG generator built edges by rejection --

  from: fc.integer({ min: 0, max: n - 1 })
  to:   fc.integer({ min: 0, max: n - 1 })
  .filter(({ from, to }) => from < to)

With n === 1 both integers are forced to 0, so the predicate is unsatisfiable
and fast-check retries value generation forever. n is drawn from 1..12 and
fast-check biases toward boundary values, so n === 1 is reached almost at once.

This also explains why the failing-first run completed normally while the fixed
run hung: before the fix the graph module did not exist, so the property test
threw on import and never reached generation. It only starts hanging once the
code under test works.

Generates the DAG by construction instead -- `to` is drawn strictly above
`from`, with the degenerate single-node case short-circuited to an empty edge
list -- so no rejection sampling is involved. Switches the import to the shared
fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds
a bounded regression guard that samples the arbitrary directly, so a future
reintroduction fails loudly instead of hanging.

Verified: the file now completes in 2 seconds, 29 tests started and 29 finished,
zero failures, against an indefinite hang before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): restore the depends_on display contract and acknowledge the workflow growth

Full-suite run surfaced two things the focused harnesses could not.

1. Regression of a pinned pre-existing contract (#3785). A refactor routed the
   EMITTED depends_on field through the new dependency resolver, which also
   consults the canonical-prefix map. The original consulted the plan map only,
   so a short canonical prefix passed through verbatim -- '24-01' stayed
   '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a
   DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with
   a comment recording why it must not use the resolver; full resolution is
   still used for the wave DAG and halt propagation, which is what needs it.

2. The workflow file grew 518 bytes without an acknowledgment, from the
   halt-aware skip rule and the widened parse contract. Acknowledged.

Note on where the acknowledgment landed: the guidance is to add a NEW fragment,
but execute-phase.md is already named by an existing fragment and the linter
hard-fails when two ack sources name the same path. Appending to the owning
fragment, following its own established multi-PR pattern, was the only
lint-clean option.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2830): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 06:47:46 -04:00
Tom Boucher
ff4a57b78c chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model

Extract the --forensic-gated forensic_audit step to
workflows/progress/steps/forensic-audit.md behind a section marker, and
repair progress.md's init line to forward --forensic so the atom is
actually true in production rather than only under direct CLI tests.

progress.md shrinks 32630 -> 27207 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize the four manifest-wired workflows

new-project, quick, new-milestone and progress each already had a
dedicated cmdInit* entry point but zero marked sections. Extract nine
gated bodies to workflows/<wf>/steps/ behind section markers and repair
each init line to forward its flags.

Fold --full into the discuss/research/validate facts inside cmdInitQuick
so the when= grammar never sees an OR, per the chunked-mode precedent.

Fixes found while working, per the no-defer rule:
- cmdInitProgress passed no phase info to buildSectionManifestField, so
  state:phase-mvp-mode was permanently false — an atom in the vocabulary
  whose fact could never be computed.
- the quick init router folded flag tokens into the free-text
  description, which the new forwarding would have corrupted.
- a #2508 dispatch note was nested inside quick.md's Agent(prompt=)
  fence, leaking orchestrator guidance into the subagent prompt.
- progress.md had a 3-vs-4 backtick outer-fence imbalance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active

Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated
entry point that never emitted a manifest — and mark two sections.

state:ui-phase-active folds (plan:pre hooks include an active ui step) OR
(the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the
grammar still sees a single operator-free atom. The inner Playwright-MCP
check stays as prose inside the fragment: it is live session state and no
init seam can precompute it.

The MVP false-branch note is a real fallback, not redundant prose, so it
sits outside the marker — gating it away would delete the text needed
precisely when MVP mode is off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): follow moved workflow content in drift guards

Retarget every guard that asserted on content this branch moved into
workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was
verified to still fail when its step file is blanked, so none was
weakened into vacuity.

Three assertions in verify-mvp-uat were genuinely red. Three more were
worse than red — passing for the wrong reason:
- quick-commit-boundary and worktree-cleanup anchored on indexOf('Step
  5.6'), which matched a later cross-reference and sliced 16069 chars
  that coincidentally held the asserted substrings. Replaced with an
  expandWorkflowSections helper that splices step content back in place.
- phase6-review-capabilities lost its end boundary and widened to EOF.
- playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall
  back' in a subagent-dispatch line after the real content moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize code-review and complete-milestone, admit three atoms

Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points
alongside the shared generic ones rather than modifying them — init.phase-op
and init.manager carry a CRITICAL blast radius (179 dependents, 24
processes) and stay byte-identical for their other callers.

Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with
a consuming section and a fact its own entry point computes.

Both sections had the resolver-in-body hazard: the fallow config-gate and
the git.create_tag check each sat inside the very block being gated, so
gating would have disabled the resolver that decides the gate. Both are
hoisted into init and the bodies now consume the resolved fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): retarget code-review and milestone drift guards, fix two red tests

Retarget guards that asserted on content moved into steps/, proving
non-vacuity by blanking each step file and confirming failure.

Also fixes two genuinely red tests found while working, per the no-defer
rule:
- workflow-fragments' frozen-vocabulary lock was missing
  state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build
  both passed over it, which is why neither is sufficient verification.
- code-review's quick.md capability-hook assertion carried a stale
  delimiter after the 18ff35d20 extraction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge

Five sections share one atom, the pattern plan-phase already uses for
flag:--research-phase. The atom folds --converge OR --cross-ai into a
single boolean in cmdInitAutonomous so the grammar stays operator-free.

cmdInitAutonomous is additive; init.milestone-op, init.manager and
init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash
resolver is deliberately retained — ungated local-planning bullets still
read it, so the init-side fact supplements it rather than replacing it.

converge-fail-fast required splitting one bash fence so the always-run
CONVERGENCE_ARGS construction stays outside the marker. All three
flag-absent fallbacks were left outside their markers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize review and discuss-phase-assumptions

Admit state:reviewer-instances-configured (two peripheral notes share it;
the core reviewer-lane dispatch stays unmarked — it is the workflow's
primary always-evaluated logic, not an optional branch) and
state:auto-advance-active, which folds --auto OR two config keys into one
boolean so the grammar stays operator-free.

discuss-phase-assumptions was the highest-risk edit in this PR. Its
auto_advance step is a full if/elif/else; gating it whole would have
deleted the flag-absent fallback needed exactly when --auto is off. Split
verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both
stay outside the marker; only 653-667 is gated.

Adds emitted-drift acks for the two files that grew — review.md (+55 B)
and autonomous.md (+737 B from 80799211c, which had none and would have
red-gated the push.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A

Completes the 13-workflow rollout. Three of these had no init call at all
and gained a dedicated entry point plus their first gsd_run query line.

Admits state:is-monorepo and adds state:next-channel, state:workstream-active
and state:flat-mode. Vocabulary 26 -> 30 atoms.

Part A of new-milestone applies when NO workstream is active — the negation
of state:workstream-active. Rather than teach the grammar negation, which is
the Greenspun drift the frozen list exists to prevent, it gets a separate
positively-phrased atom whose fact is the inverse. Part B, which always runs,
stays outside the marker.

flag:--verify-only is deliberately NOT admitted: docs-update has no
contiguous purely-additive region for it, and an atom without a consuming
section is dead vocabulary. Evidence recorded in the slice report.

update.md reuses its existing resolved $GSD_TOOLS rather than prepending the
canonical preamble, which would have clobbered it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary

Two defects the new tests caught.

The automated-ui-verification step re-ran gsd_run loop render-hooks and
recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact
is already true — the circular self-disabling pattern this design forbids,
introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and
the step consumes it. Its launcher preamble goes too: no gsd_run remains.
The Playwright-MCP check stays as prose — that is live session state.

Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map
were admitted with a gate-1 claim that never materialized. flag:--full is
removed, redundant once quick folds it into discuss/research/validate.
state:needs-codebase-map gets the real consumer it always lacked, gating
new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now
without a consuming section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): add the atom-admission, inversion and resolver-hoist gates

The two existing parity guards prove vocabulary/predicate symmetry but
never that a fact is computed — an atom no cmdInit* assembles evaluates
false forever. These close that hole:

- per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion
  so the loop cannot silently cover zero atoms
- dead-vocabulary check against the shipped manifest
- inversion guard: the flag-absent fallbacks in discuss-phase-assumptions
  and verify-work must stay outside their markers
- data-driven resolver-hoist guard over the shipped manifest, so a future
  extraction cannot reintroduce the circular class
- compound-fold coverage (--full, --cross-ai, --rc, config-only --auto)
- null-vs-[] degraded/computed distinction, and flag value shapes

Also repairs the frozen-vocabulary lock, which was stale and red for the
seven atoms earlier commits on this branch shipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): add changeset for the fragment-model rollout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): cite the issue on the two new allow-test-rule exemptions

ADR-456 requires an issue ref on the same line as the annotation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): correct the atom-count claims after retiring flag:--full

The vocabulary doc comments still said 30 entries; it is 29 since
flag:--full was removed as dead vocabulary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): dedupe the phase-fallback block and harden --ws parsing

Review findings.

MAJOR: the three new init entry points each pasted a verbatim copy of the
guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four
copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and
folded six of the seven; each call site keeps its own field-set via a
closure. Duplication removed rather than papered over with a parity test.
cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a
byte-identical copy, and it is CRITICAL-radius.

LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a
workstream name holding glob metacharacters would expand against the
filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept —
it must word-split into two args and vanish when empty.

Also restores the vocabulary ordering convention, and fixes a masked test
bug the mandated run surfaced: the flag-forwarding guard checked only the
first init line per workflow, but new-milestone has two, so a real failure
was reporting exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the stale new-milestone emitted-drift ack

new-milestone.md was acked for a +406 B growth measured against an
intermediate commit. Net against origin/next it SHRANK by 8 bytes, so
nothing needed the ack and it explained nothing — which the differential
attribution check reports as a stale acknowledgment, not a pass.

update.md's entry stays: it genuinely grew +703 B.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): resolve the 15 failures from the full matrix run

All 15 were real and identical on both lanes.

REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's
40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the
five section stubs pushed it over. Extracted the 3a.5 UI Design Contract
body to references/; now 39968 chars, and the file nets -795 B vs base, so
its growth ack is deleted rather than left stale.

REAL DEFECT: docs referenced /gsd-transition, which is not a live
registered command. Reworded.

STALE FIXTURE: the emission byte-identity test hardcoded two marked
workflows; this branch legitimately marks fifteen. Fixture corrected — the
source was right.

The rest were drift guards over the eight workflows the earlier sweep did
not cover, retargeted at where the content now lives with non-vacuity
proven by blanking each step file and confirming failure. The GSD_WS
forwarding guard was checked as a possible real break and is not one: the
charclass narrowing is intact and forwarding works end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the ack for a newly-added reference file

A new file's emitted ripple is attributable to the diff that adds it, so
the acknowledgment explained nothing and the differential check reports it
as stale. Removing the last entry removes the fragment — an empty one
signals nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): retarget the UI-contract guards and clear two transitive advisories

The §3a.5 extraction that brought autonomous.md under the #2196 char cap
moved its body to references/autonomous-ui-design-contract.md, so ten
guards in autonomous-ui-steps and check-ui-safety-gate were asserting it
against the host. Retargeted via a combined read, each proven non-vacuous
by blanking the reference file and confirming failure.

This class had already bitten twice on this branch because each sweep was
scoped to the workflows touched at that moment, so this one was
exhaustive: ~70 test files across all 13 workflows, zero further broken or
vacuous assertions found.

Also clears two high transitive advisories the matrix flagged on one lane
— fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary
issues. Both pre-date this branch: package-lock.json was untouched until
now, so the production tree was byte-identical to the base. Lockfile-only,
package.json unchanged, verified against a real npm ci install.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): backfill changeset pr number to 3030

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 19:59:58 -04:00
Tom Boucher
c6ce4d1d9a fix(#2755): resolve the kimi hooks-TOML root per runtime (#3032)
* test(#2755): failing-first coverage for per-runtime kimi hooks root

Install/uninstall filesystem-shape rows over a sandbox HOME (no permission
tricks) plus resolver unit rows. Covers both uninstall directions, which is
where a fix applied only to the install call site would drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2755): resolve the kimi hooks-TOML root per runtime

resolveKimiHooksTomlDir took no runtime argument and hardcoded ~/.kimi, but
both kimi and kimi-code route through the single hooksSurface=kimi-hooks-toml
branch. A --kimi-code install therefore wrote its [[hooks]] block, hook bundle
and CommonJS marker into Kimi CLI's config file, and a --kimi-code uninstall
stripped Kimi CLI's block.

Adds a runtime selector to the resolver -- kimi keeps ~/.kimi + KIMI_SHARE_DIR,
kimi-code gets ~/.kimi-code + KIMI_CODE_HOME, per Kimi Code's own upstream
data-locations and hooks docs -- and passes the runtime at both the install and
uninstall call sites. An omitted or unrecognized runtime still resolves ~/.kimi,
so the exported no-arg contract is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2755): use centralized helpers and add a divergence guard

Review findings, all fixed in-PR:

- The new test block reimplemented runMinimalInstall, createTempDir and
  toPosixPath. Extends runMinimalInstall with optional root/extraEnv instead
  (back-compat: every existing caller passes neither) and uses the centralized
  helpers, per CONTRIBUTING's Use Centralized Test Helpers rule.

- Adds a parity assertion between the capability registry and the resolver: a
  third runtime declaring hooksSurface kimi-hooks-toml would silently inherit
  ~/.kimi, re-creating this very defect. The guard fires the moment those two
  surfaces drift.

- Adds an installer-level test proving KIMI_SHARE_DIR and KIMI_CODE_HOME do not
  interfere when both are set, which only the resolver unit covered before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2755): track the kimi-code hooks root in the emitted-artifact gates

The remote runner caught a real ripple: moving kimi-code hooks to ~/.kimi-code
made 31 emitted paths unattributable and 58 emitted hashes unexplained, because
three parallel surfaces keyed on the literal .kimi path.

- HOOK_CONFIG_RELATIVE_PATHS excluded only .kimi/config.toml, so kimi-code's
  config.toml became manifest-visible; it embeds a platform-varying node-runner
  command and must stay out for both products.
- HOOKS_ROOTS, the package.json-marker branch and the synthesized-install-metadata
  pattern each named .kimi only.
- tests/fixtures/install-tree/kimi-code.json still recorded the old paths;
  regenerated via gen:install-tree.

Adds the per-PR drift acknowledgment for the 58 paths whose bytes are unchanged
but whose destination moved - a ripple no source diff can show, since no hook
script was edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2755): clear production-tree security advisories

The remote runner's npm-integrity gate reported 2 high advisories in the
production dependency tree. My diff touches neither package.json nor
package-lock.json, so these come from the base -- but a red gate is not
something to wave off as pre-existing, so it is fixed here rather than deferred.

Lockfile-only, semver-in-range, via npm audit fix:
  fast-uri   3.1.4  -> 3.1.5   (host confusion via backslash authority introducer)
  ip-address 10.2.0 -> 10.4.0  (three SSRF / trust-boundary bypasses)
  hono       4.12.31 -> 4.13.0 (moderate; reverting it traded a high for a
                                moderate, so the full remedy is taken)

npm audit now reports 0 vulnerabilities at every severity, npm ci installs
clean from the updated lockfile, and the build and the kimi behavior both
re-verified afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2755): backfill changeset pr numbers

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:56:42 -04:00
Dennis Kim
d3ddcaba1c fix(#2785): implement missing gate predicate evaluators (#2816)
* fix(#2785): implement missing gate predicate evaluators

* fix(#2785): gate predicate numerical coercion

* fix(#2785): address evaluator review findings

* fix(#2785): use safe frontmatter read seam

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:16:56 -04:00
𝚌𝚕𝚎𝚣𝚌𝚘𝚍𝚒𝚗𝚐
88f6d9bd1b fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu

* fix: preserve installer executable mode

* chore: add changeset for PR #2812

* test(#2644): acknowledge Cursor emission changes

* test(#2644): drop spent emitted drift acknowledgments

* fix(#2644): remove retired Cursor command converter

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:05:45 -04:00
Tom Boucher
067a4d1c6c fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection

Regression test for gsd_stall_should_recover / gsd_stall_watch and the
planner.stall_* config keys, none of which exist yet — proves RED before
the fix lands in the next commit.

* fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls

Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug
#3212) but with a dispatch change the executor's prose-only surveillance
lacks: the standard planner spawn, chunked-outline planner spawn,
chunked-per-plan planner spawn, plan-checker spawn, and revision-loop
planner respawn now dispatch with run_in_background=true and are followed
by a real, bounded bash poll (gsd_stall_watch) that returns control to the
orchestrator on its own schedule instead of waiting indefinitely on a
subagent that may never return. On stall, the existing accept-plans/retry/
stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual
interrupt.

New config keys planner.stall_detect_interval_minutes (default 5) /
planner.stall_threshold_minutes (default 10) mirror executor.stall_*.

The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a
new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-
helpers.md rather than inline, and per-site prose is kept minimal, because
plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate
(tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at
baseline; the net effect is plan-phase.md.md ships slightly SMALLER than
before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at
the five touched sites, superseded by the bounded watcher).

Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that
still described the per-file workflow-size-baseline.json guard removed by
#2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism —
discovered while verifying this fix's own byte budget.

Researcher and pattern-mapper spawns are untouched (out of scope per the
issue's Agent Brief).

* fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs

Two review findings addressed on top of the prior commit:

1. gsd_stall_watch previously looped internally for the full
   threshold+interval duration inside ONE Bash tool call (up to 15 min at
   defaults) — a single call blocking that long risks the host tool's own
   timeout killing it before it ever prints a result, silently defeating the
   fix. Redesigned to a single sleep-and-check cycle per call, taking an
   explicit dispatch_ts so the orchestrator prose can repeat the (short,
   default 5 min) call until it resolves; the outer threshold is now
   enforced by dispatch_ts accumulating across calls, not by one call's
   duration. Documented the resulting trade-off (up to one interval of
   added latency on the success path) in the changeset and reference doc.

2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled
   values that flow into bash arithmetic ($(( ))). A review flagged this as
   command injection; empirically verified against both macOS bash 3.2.57
   and Docker bash:5 that this is NOT actually exploitable (bash hard-errors
   on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an
   unvalidated malformed value WOULD abort the stall-watcher itself with
   that bash error, silently defeating the exact hang-recovery this issue
   ships. Added integer validation with safe-default fallback, both at the
   config-resolution point and defensively inside gsd_stall_should_recover.

Also adds the previously-missing integration coverage for gsd_stall_watch's
real execution (grep/find/date plumbing), not just the pure classifier.

* fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose

The AC2 regression test asserted the stall-detection-helpers.md step file
never contains the substring "teams-status" at all, but the file's own
prose explicitly documents its independence from that guard (containing
the word by design). Narrowed the assertion to what actually matters: no
second `query teams-status` call site and no gating on it, not a blanket
absence of the word.

* test(#2650): regenerate golden install-tree fixtures for the new step file

gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an
emitted file (installed for every runtime), so adding it changes the
install tree even though it is invisible to docs/INVENTORY.md and
docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive
gsd-core/workflows/*.md — verified against the execute-phase #2930 and
pre-existing plan-phase step-file precedent, which are equally absent from
both inventory artifacts). The golden install tree snapshots the sorted
list of emitted relative paths per runtime, so a file invisible to the
inventory is still visible here. Regenerated via `npm run gen:install-tree`
— one line added per runtime fixture (19 files), no other drift.

* fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble

Two more consequences of extracting helper bodies out of plan-phase.md,
both caught by verification (0017e1a78, 9 unique failures):

1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7
   "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one
   per agent spawn site. Moving the full explanatory blocks to
   plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels
   out with them (only the untouched researcher/pattern-mapper sites kept
   theirs). Restored a short label at each of the 5 stall-watch sites,
   trimmed a few more redundant words ("Per 7.99, " — already established
   by the adjacent step-7.99 pointer) to stay under the frozen
   PRE_PHASE6 cap (94497 bytes, 21 bytes headroom).

2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one
   canonical gsd_run preamble, byte-equal to
   gsd-core/workflows/_runtime-launcher.snippet.sh, before the first
   gsd_run call in any workflow .md that calls it (recursive scan under
   gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance
   checks). The new step file's config-get calls use gsd_run without one.
   Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1
   preamble occurrence, before the first call, including the .claude/ and
   .codex/ home fallback arms.

Also verified (no fix needed, evidence recorded): the generic
`gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs
(roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any
new gsd-core/workflows/** path to itself, so the new step file needs no
drift-ack entry — consistent with plan-phase.md's own net shrinkage
requiring none either.

* test(#2650): acknowledge plan-phase.md's +14 byte drift

Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped
plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth
against baseline (94483 -> 94497), which the differential attribution
size ratchet (tests/emitted-attribution.test.cjs) correctly flags as
unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json, keyed on the bare filename plan-phase.md per the
existing fragment schema (see tests/emitted-drift-acks/2649-diagnose-
execute-plan-base-check.json), explaining the growth as exactly the 5
restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519)
and satisfying #913's 7-label requirement.

* fix(#2650): bind {outputFile} from the real Agent() return — was dead code

Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were
read by every gsd_stall_watch call but never assigned anywhere in the
diff. With the variable permanently empty, `[ -f "$output_file" ]` was
always false, marker_found could never become true, and marker_received
was unreachable — the marker-based detection path was permanently dead.

Worse for the plan-checker spawn specifically: a checker that PASSES
touches no *-PLAN.md files, so it had no working completion signal at
all without the marker path. A healthy plan-checker finishing cleanly in
two minutes would be declared stalled once planner.stall_threshold_minutes
elapsed and the recovery menu would fire on an already-succeeded agent —
worse than the original unbounded hang.

Fixed by replacing the dead bash variable with the `{outputFile}`
orchestrator-substitution token, the same convention docs-update.md:471
already uses for a real run_in_background=true Agent() return ("Read
tool: file_path: `{outputFile from README agent result}`"). This is a
net BYTE SAVING at each site (`"{outputFile}"` is shorter than
`"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding
explanation — including why plan-checker's *-PLAN.md glob alone is not
a working completion signal — into the lazily-loaded reference file to
stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13
over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase-
stall-detection.json).

Added a regression test asserting plan-phase.md itself binds {outputFile}
at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE /
$CHECKER_OUTPUT_FILE reference — the previous test suite only exercised
gsd_stall_watch's behavior when handed a valid argument, which is why
the dead production wiring survived two rounds of review. Also fixed
tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw
try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention.

* chore(#2650): backfill changeset PR number to 3015

* fix: normalize CRLF at the read boundary in all .md-bash-extraction tests

Maintainer-authorized scope expansion, folded into this PR rather than
deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase-
stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE
(CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a
one-off. Ten test files parse a fenced ```bash block out of a workflow
.md file and execute it via spawnSync/execFileSync; a Windows checkout
can yield CRLF line endings despite .gitattributes eol=lf, and bash then
treats the trailing \r on every extracted line as part of the token —
"unexpected EOF while looking for matching `"'" or a bare syntax error,
partway through the script.

Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the
read boundary, before any fence-slicing or regex runs, so every
downstream operation is correct by construction. Migrated all ten call
sites to it:

Previously broken (fs.readFileSync with no normalization anywhere
between read and spawn):
- tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a
  misleading comment claiming the fence regex alone was "CRLF-safe"; it
  protected only the fence delimiters, never the captured body.
- tests/new-milestone-clear-phases.test.cjs (extractFenceBetween,
  extractFenceContaining)
- tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript)
- tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file
  comparison read in the same test)
- tests/graphify-visualization.test.cjs (extractStep3Block)
- tests/pause-work-improvements.test.cjs (extractCheckBlock)
- tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock
  and the inline post-config-gate resolution-block slices)

Already correct (split(/\r?\n/) then join('\n')), migrated to the shared
helper for consistency rather than a fourth/fifth/sixth copy of the same
fix:
- tests/git-base-branch.test.cjs (extractHandleBranchingBash)
- tests/quick-branching.test.cjs (extractStep25Bash)
- tests/runtime-launcher-parity.test.cjs (extractResolverSnippet)

Verified against a simulated Windows CRLF checkout (not assumed): for
both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs
extraction shapes, confirmed the pre-fix code produces a real bash syntax
error on CRLF input and the post-fix code does not.

One eslint follow-up: local/no-crlf-fragile-split statically flags any
bare `\n` inside a markdown-fence-shaped regex, regardless of whether the
receiver was already normalized — it cannot see the readFileNormalized()
data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant
but harmless on pre-normalized input) rather than fight the rule.

Scope note: this diff is broader than issue #2650's own change (plan-
phase.md stall detection) because the Windows lane surfaced a genuine
repo-wide defect class while verifying that fix, and the maintainer
authorized fixing it here rather than filing it separately and shipping
a known-broken pattern.

Runtime impact: none — this is a test-harness-only defect. The live
orchestrator (Claude Code or another runtime) does not do a byte-exact
extract-and-pipe of .md content into a shell the way these tests do; it
reads the instructions and generates its own bash invocation text, which
does not reproduce a raw CRLF pass-through the same way.

Not touched: tests/plan-review-convergence.test.cjs's separate, tracked
spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on
unmodified next) — unrelated load-sensitivity, not a CRLF symptom.

* fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged
plan-phase.md's own emitted-path hash move, but plan-phase.md is directly
edited in this diff. Per the emitted-attribution law (ADR-2719,
tests/emitted-attribution.test.cjs), a workflow's emitted key equals its
own source path (gsd-core-verbatim identity rule), so a direct edit to the
source is self-explaining and auto-attributed — no ack was ever needed.

Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run
through npm run lint:ci with a fully cleared eslint cache): it passes clean
with the fragment removed, confirming no contradiction between the lint and
the runtime attribution gate — this was simply an unnecessary fragment.

* fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next

tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in
the previous commit because, against an earlier verification base, it was
inert: it explained a moved emitted hash that a direct edit to plan-phase.md
already self-attributes. Against origin/next@f1af47766a the demand is
different: plan-phase.md is 13 bytes larger than the base copy, which trips
the emitted-attribution size ratchet — a job this same ack also performs.

Recreated in the documented shape, keyed on the bare filename plan-phase.md
(not the full path, and not restating the byte delta per review guidance),
describing the actual change: the {outputFile} binding fix for the dead
PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored
ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn
sites, with explanatory bodies living in the lazily-loaded
gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference.

Confirmed no other fragment (on this branch or on next) claims the bare key
"plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's
duplicate check is an exact string match, and the only other mention of
plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json)
uses the full path as its key, so there is no collision.

* fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF

The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong,
not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows
checkout never receives CRLF for stall-detection-helpers.md, and the
extracted fence's line 64 is byte-identical and correctly balanced on every
platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense
script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])`
PLUS four more positional args. Windows has no execve — Node serializes
that whole argv into a single CreateProcess command-line string, and Git
Bash's MSYS layer re-splits and unescapes it with its own rules. The
boundary between the script and the trailing args was not stable across
that round trip (live evidence: one failure's stderr was prefixed
`gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:`
— arg0 did not).

Fixed by writing the script to a temp file and running `bash <file> <args>`
instead — the four values are now normal, quote-free positional args, and
the script itself never enters argv transport at all. Mirrors
tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already
uses this exact shape and is green on Windows on `next`.
tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on
`bash -c` but never appends extra positional args beyond the script itself,
so it never hits the same boundary — checked both siblings per review, not
assumed.

Corrected the now-actively-misleading CRLF comment in
extractStallHelpersBash(), and corrected the changeset's claim that the
repo-wide CRLF-normalization fix (folded into this branch, maintainer-
authorized) explains this PR's own Windows failure — it doesn't, though it
remains defensible on its own merits as general test-portability hardening.

Separately, while auditing the shipped (non-test) gsd_stall_watch for
Windows portability per review request, found and fixed a second, real
user-facing defect: the artifact-freshness check used GNU find's
`-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on
macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live
against /usr/bin/find on both a stale and a genuinely fresh file). With the
adjacent `2>/dev/null`, that failed silently and permanently degraded
artifact_fresh to false on every macOS run — a plan-checker or planner
actively writing plan files could still be reported "stalled." Replaced
with `find $glob -mmin -N` ("modified less than N minutes ago"), which
needs no date-string parsing and is supported identically by GNU find and
BSD find; verified live that the old shape fails and the new shape passes
against the same real fresh file. Added a real-execution regression test
(gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't
actually wait, but the real `find ... -mmin` line still runs) proving the
fix, replacing the prior "not integration-tested" note for that path.

Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm
the Windows fix — only the actual windows-latest CI lane can.

* fix(#2650): route the third bash -c call site through the same temp-file seam

runWatch() and a `-mmin` regression test still passed their script via
`bash -c <script>` after the previous commit only converted
runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism:
failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and
`shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in
this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe
block, which runWatch() serves. runWatch() passes NO extra positional args
at all, so this also rules out the trailing-args theory from the prior
commit: the ~73-line, quote-dense script itself is what does not survive
Windows argv serialization when passed as a single `-c` element, regardless
of how many (if any) further argv elements follow it.

Extracted one shared runBashScript(script, args, opts) helper — write to a
fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` —
and routed all three bash-invoking call sites in this file through it
(runShouldRecover, runWatch, and the -mmin freshness test that builds its
own script inline for the `sleep` stub). One transport seam means a fourth
call site in this file cannot silently reintroduce the bug in isolation,
which is exactly what happened here with a second call site.

Corrected extractStallHelpersBash()'s doc comment a second time to state
the mechanism precisely (script content, not argv-element count) and cite
the live evidence (11->4 failures, shards 1 and 2 flipping green) so the
next reader does not have to rediscover it.

Audited every other bash-invoking call site in files this branch touches,
per review request:
- tests/code-review-pipeline-regression.test.cjs (runPostProcessing),
  tests/graphify-visualization.test.cjs (runBlock), and
  tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...])
  sites, one of them carrying the same giant runtime-launcher preamble
  text) — all pre-existing, UNCHANGED by this branch (only touched for the
  readFileNormalized() CRLF swap), and already exercised on `next`'s last
  six Windows CI runs per the reviewer's own citation. Left as-is: no
  evidence of failure, and converting untested pre-existing code outside
  #2650's scope on an unverifiable guess would be its own risk.
- tests/git-base-branch.test.cjs (runHandleBranchingStep) and
  tests/quick-branching.test.cjs (runStep) already use the same temp-file
  pattern. No action needed.
- tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but
  is explicitly `if (process.platform === 'win32') return '';` guarded off
  on Windows entirely, for an unrelated extension-less-PATH-stub reason —
  never reaches Windows argv transport at all. No action needed.
- tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as
  correct and verified; not touched, per instruction.

Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all
three confirmed correct in prior rounds and left untouched here.

Note: the remote gsd-test runner is Linux-only and cannot confirm this;
only the windows-latest lanes on #3015 can.

* fix(#2650): give runBashScript a default timeout

runShouldRecover() was the only one of the three call sites through
runBashScript() with no timeout — runWatch() and the -mmin test both pass
timeout: 10000 explicitly. Not a regression (this path never had a bound
before), but CONTEXT.md's unbounded-subprocess guidance applies directly,
and runShouldRecover() is driven repeatedly by a fast-check property test:
one pathological input that fails to terminate would hang CI indefinitely
instead of failing.

timeout: 10000 is now the helper's own default, with ...opts spread after
it so the two existing explicit timeout: 10000 call sites are unchanged
and any future caller inherits a bound automatically.

* fix(#2650): build the -mmin freshness test's glob with forward slashes

Windows CI on d6ddda6ea reported the last failure: the -mmin regression
test expected 'active' but got 'waiting' — find matched nothing, the same
silent-degradation shape as the macOS -newermt defect, but this time in the
test's own fixture rather than the shipped bash.

Traced what production actually passes: every gsd_stall_watch call site in
plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` —
PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole
thing runs under Git Bash regardless of host OS, so production's glob is
always forward-slash. The test instead built it with
`path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path
(C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash
escapes the next character, so that pattern can never match a real path —
find silently returns empty under the existing 2>/dev/null, same shape as
the macOS bug. Confirmed as a test artifact, not a production defect:
production never constructs the glob this way, so no Windows user is
affected.

Fixed by forward-slashing the tmp dir before appending the glob suffix,
matching production's own convention, with a comment recording why (so a
future "simplify this back to path.join" edit doesn't silently reintroduce
the failure). The shipped bash's unquoted $artifact_glob is untouched —
quoting it would break the multi-file glob expansion it exists for.

Note: the remote runner is Linux-only and already passed clean at
d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on
#3015 can confirm this fix.

* fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR)

The :353 fix (833c11da9) only converted the -mmin freshness test's glob.
Three sibling tests in the same describe block still built theirs with
path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows.

Two of those three were silently passing for the wrong reason: the
'-> stalled' and '-> waiting' tests both expect the glob to match nothing,
and on Windows a backslash path matches nothing regardless of whether the
directory is actually empty (bash eats each backslash as an escape before
the pattern is even evaluated). They would have passed identically with
glob expansion completely broken, which is a vacuous pass — not exercising
what they claim to. The third ('-> marker_received') is outcome-independent
of the glob, so it was merely inconsistent rather than wrong.

Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md`
construction already used at the -mmin test, so every glob in the file now
matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape,
and the two negative tests are meaningful on Windows instead of accidentally
correct. Reworded the trailing comment on the 'stalled' test's glob line:
it now describes the fixture (the tmp dir contains no *-PLAN.md files)
rather than the pattern, since "matches nothing" read as a property of the
glob syntax when it's a property of what's on disk.

No assertion, the sleep stub, runBashScript, or the shipped bash changed.
Smoke-tested all three updated tests manually before committing (not via
node --test): marker_received / stalled / waiting, all correct.

* fix(#2650): fix own regression tests for #2993's plan-phase.md relocation

531101843's merge with origin/next brought in #2993 (unrelated, epic #1671
Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode"
section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md,
leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard.
test.cjs (#913) was already updated to read the combined surface (host file
+ every steps/*.md) so its label count didn't go blind — my own #2650
regression tests were not, and searched plan-phase.md alone for the two
chunked spawn sites' headings, which no longer exist there. Two tests
failed outright (indexOf returning -1); a third ("standard planner spawn")
was silently weakened to an unbounded slice-to-EOF by the same relocation,
since its own end-boundary heading also moved — passing by accident rather
than by testing what it claimed.

Promoted the drift guard's local readPlanPhaseCombined() to a shared,
exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file +
sorted steps/*.md, CRLF-normalized at the read boundary) so a second,
divergent implementation is never written — the drift guard now delegates
to it via a same-named local wrapper, unchanged at every existing call site.

Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection.
test.cjs:
- "standard planner spawn (step 8)": end boundary changed from the now-gone
  "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return",
  which still exists in plan-phase.md.
- "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now
  read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md
  directly (not the generic multi-file combined blob, whose file-sort
  ordering would put unrelated step files between 8.5.2's slice and any
  downstream anchor) — the same heading-to-heading slicing as before still
  works because the file is small and self-contained.
- Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check
  to also scan chunked-planning-mode.md, since two of the five spawn sites
  now live there.
- Added a new count-based test asserting exactly 5 (not "at least one")
  `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined
  surface, mirroring #913's own label-count guard, so every one of the five
  spawns stays provably bounded and a future relocation can't silently drop
  one without a test noticing.

Also added a small positive test that plan-phase.md's <!-- gsd:section -->
pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but
its presence is now load-bearing for where 2 of the 5 spawn sites live).

Audited every other test file in the repo for a stale reference to content
#2993 relocated (searched for the moved headings/prose and for
"chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this
file and the drift guard needed changes.
tests/issue-2762-plan-reviews-chunked.test.cjs already reads
chunked-planning-mode.md directly (brought in correct by the same merge).
gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs
reference "chunked-planning-mode" only as a manifest/section-id fixture
value for #2993 itself, not as a stale pointer to relocated content.

Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true,
the glob constructions, runBashScript, the -mmin change, the timeout
default, or the drift-ack fragment (confirmed correct against the stale
local `next` ref two rounds ago and left alone).

---------

Co-authored-by: sim <sim@local>
2026-08-03 10:46:22 -04:00
Tom Boucher
ad3b9ec486 chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model

Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and
carried zero markers; it was deferred out of the Phase 3 pilot for two
reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the
blocker it looked like — fragmentizing is net-negative on host source, so
the trim is what creates the room. The --mvp interleaving was resolved by
measurement in #2992 and no sub-line mechanism is built.

- widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671
  amendment: flag:--ingest, flag:--prd, flag:--research-phase,
  flag:--reviews, state:chunked-mode
- state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and
  that disjunction is resolved in the FACT, never in the grammar, so a
  compound condition never becomes an operator
- parse the new flags on the plan-phase route; extract six gated bodies to
  gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs
- prd-express-path.md was already extracted but read unconditionally; its
  wrapper is now gated, so the existing extraction finally pays off

plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36
bytes to 6,944.

Also closes a surfaced docs gap: five real plan-phase flags (--chunked,
--skip-ui, --bounce, --skip-bounce, --granularity) were documented in
neither the argument-hint nor help. Making --chunked load-bearing without
fixing its siblings would leave the defect class half-open.

Refs #2993

* fix(#2993): forward flags to the init bundle so section gating actually fires

Blocker found by the correctness review, confirmed directly, and missed by
both the isolated reviewer and every test in this branch.

Neither workflow forwarded its flags to the init CLI:

  plan-phase.md:71    INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM)
  execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}")

So every flag: atom was permanently false in production and its section
permanently excluded. For plan-phase that made the PRD express path
UNREACHABLE — a regression, since it was an unconditional read before.
For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating
that has never once been true, so `--wave` silently dropped its own
wave-filtering guidance. Fixed here under the no-defer rule.

Why every test missed it: they drive the init CLI directly with flags,
which works. Production goes through the workflow's bash line, which did
not pass them — the exact "assert against the shape production uses" trap
this branch's own test matrix warns about.

- parse and forward --prd/--ingest/--research-phase/--reviews/--chunked
  (plan-phase) and --wave (execute-phase), using the anchored regex idiom
  the neighbouring GRAN_PARAM line already uses
- add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X
  section, the owning workflow's init line must forward --X. It fails
  against the pre-fix files and covers any future atom, rather than
  spot-checking today's six.

Verified through the workflow shape, not the CLI shape: `3 --prd spec.md`
now yields ["prd-express-gate"] (was []), `2 --wave 2` yields
["partial-wave"] (was []).

Refs #2993

* test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures

Remote matrix was red with 46 unique failures, identical on both lanes.
Both causes are mechanical consequences of changing shipped workflow
content, and neither is visible to any local gate.

- emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM
  forwarding fix and was unacknowledged, while the ack fragment named
  plan-phase.md, which SHRANK and therefore needed no ack at all — a stale
  entry is itself a failure. The reason now names the real ripple.
  The entry had to merge into the existing 2930 fragment: the ack linter
  does unconditional cross-fragment duplicate-key detection with no
  spent/live exception, so a second fragment declaring execute-phase.md
  collides even when the first is already merged and inert. Resolved per
  the linter's own guidance and that file's precedent of appending
  successive ripple reasons to one entry.
- golden-install-tree: tests/fixtures/install-tree/*.json are committed and
  deliberately excluded from the ADR-2719 attribution cutover, so they must
  be regenerated when shipped tree content changes. Regenerated after
  build:lib per the ordering landmine. 19 runtimes each gained exactly the
  six new plan-phase step files; zero paths removed, which is the absolute
  failure shape those fixtures exist to catch.

Refs #2993

* fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards

Second red run: 26 unique failures, identical on both lanes, in two classes.

RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md
calls gsd_run but carried no canonical launcher preamble, which is what
DEFINES gsd_run(). On any non-Claude runtime that step would fail outright.
The preamble is now copied verbatim from the canonical source of truth,
gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to
column 0 to match the prd-express-path.md sibling (a list-continuation
indent breaks the byte-equal preamble match). prd-express-path.md already
had a correct one. This is the same defect #2932 hit when it extracted
steps; the parity test caught a real bug, not a stale assertion.

DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked,
skill-frontmatter-contract) — these assert plan-phase.md contains content
this branch moved into step files. Retargeted at where the content now
lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT
test now reads host + every step file so the count is preserved across the
split rather than reduced. Each retargeted guard was verified to still fail
when its step file is stripped, so none was weakened into vacuity.

No emitted-drift ack was needed: currentSizes() enumerates
gsd-core/workflows/*.md non-recursively, so files under
plan-phase/steps/ are never in the size ratchet's scope.

Refs #2993

* chore(#2993): backfill changeset pr number to 3019

---------

Co-authored-by: sim <sim@local>
2026-08-03 09:38:58 -04:00
Tom Boucher
9640968f8e fix(#2847): require gap_closure value in plan-gap-closure schema and bind validate_plan to it (#3018)
* test(#2847): add failing-first regression tests for gap-closure frontmatter schema gap

--gaps did not load a machine-checked requirement for gap_closure: true.
The planner's only validation gate (frontmatter.validate --schema plan)
never required it, and plan-phase.md's downstream_consumer contract never
mentioned it either, so gap-closure plans could pass validation while
missing the field that /gsd:execute-phase --gaps-only filters on.

These tests are RED against current production code: no plan-gap-closure
schema exists yet, and neither agents/gsd-planner.md's validate_plan step
nor plan-phase.md's downstream_consumer block references gap_closure
conditionally.

* fix(#2847): enforce gap_closure via plan-gap-closure schema

--gaps did not load a machine-checked requirement for gap_closure: true.
The planner's only validation gate (frontmatter.validate --schema plan)
never required it, so a gap-closure plan could pass validation while
missing the field /gsd:execute-phase --gaps-only filters on, silently
spawning zero executors.

Add a plan-gap-closure schema (every plan-required field plus
gap_closure) and make the planner's validate_plan step select it when
gap_closure mode is active, plan otherwise. Standard/reviews-mode plans
are unaffected: plan's required fields are unchanged.

plan-phase.md's downstream_consumer block was investigated for a
symmetric mention but deliberately left untouched: it sits 36 bytes
under the frozen ADR-857 PRE_PHASE6 ceiling and the validate_plan step
in gsd-planner.md is the actual call site, needing no help from
plan-phase.md's prose.

* fix(#2847): compact validate_plan edit under gsd-planner.md size caps

Merging origin/next (7 commits, including #2775's gsd-planner.md
STRIDE-row edit) left only 22 chars of headroom under four separate
hard-coded 49152-char caps on gsd-planner.md (planner-decomposition,
precondition-element, reversibility-tagging, security.test.cjs). The
verbose validate_plan prose from the previous commit overran all four.

Compact the edit to a single line (net +17 chars vs origin/next) while
keeping the functional content: schema name, mode condition, and the
unchanged base required-fields list.

Also:
- Fix a real bug in the fix-2847 negative-assertion test: plan-phase.md
  mentions the literal string "<downstream_consumer>" twice in
  backtick-quoted prose before the actual opening tag, so a plain
  indexOf() grabbed the wrong start position and swallowed ~10KB of
  unrelated content (including a "gap_closure" hit in a Mode: enum
  line), producing a false failure. Anchor on the tag starting its own
  line instead.
- Merge the emitted-drift-ack fragment for gsd-planner.md with the
  #2775 fragment brought in by the merge (both named the same path;
  two ack sources may never name the same path) and correct its byte
  delta to the actual final number.

* fix(#2847): drop stale merge-inherited emitted-drift-ack fragments

Merging origin/next brought in three new emitted-drift-ack fragments
(1700, 2658, 2775) relative to this branch's fork point. #2775
collided with my own gsd-planner.md key and was already consolidated.
#1700 and #2658 don't collide, but none of their entries name a path
this branch's actual diff touches (git diff --name-only
origin/next...HEAD) — the ripples they explain are already baked into
the current next baseline, so they explain nothing here and the
emitted-attribution gate correctly reports them as stale (verified
live: spike-wrap-up.md from #1700).

Delete both fragment files. Neither is referenced by any test beyond
a stray comment pointing at an unrelated diagnosis artifact path, not
the ack fragment itself.

* fix(#2847): restore merge-inherited ack fragments deleted in error

1700-spike-manifest-idea-scoping.json and 2658-trae-instruction-file-path.json
exist on origin/next (landed via other, already-merged PRs) and arrived
on this branch unchanged via the origin/next merge. The previous commit
deleted them to satisfy a stale-acknowledgment finding, but the finding
was about the acks being MODIFIED in this diff, not about needing to
stop existing — deleting them would have silently reverted two other
PRs' already-merged, already-justified byte growth.

Restored byte-identical to origin/next (git diff origin/next -- <path>
empty for both). 2775-planner-package-legitimacy-gate.json stays
consolidated into 2847-gap-closure-validate-plan-step.json: that one
was a genuine hard key-collision (two fragments naming the same
gsd-planner.md path, which lint-emitted-drift-ack hard-blocks), not a
pass-through case.

* fix(#2847): bind --schema to gap_closure mode, not hardcode it

Prior revision left the validate_plan bash invocation unconditional
(--schema plan)) while only the prose sentence above it described the
gap_closure-mode branch. An agent executing the shown line literally
always validated with the plan schema, so a gap-closure plan missing
gap_closure: true still reported valid:true — #2847 reproducing
unchanged. Existing tests didn't catch it: they checked for substring
presence anywhere in the step, which the prose alone satisfied.

Change the bash line to --schema "$SCHEMA" — a real shell-variable
reference in the same placeholder convention this file already uses
for "$PLAN_PATH" (never literally assigned; the agent resolves it from
context, same as PLAN_PATH). A genuine if/then bash conditional
already exists elsewhere in this file (load_project_state's
INIT @file: check), confirming executed conditionals, not merely
descriptive prose, are the established pattern here.

Rewrite the regression test to assert on the bash block's literal
--schema argument: reject a hardcoded plan) or plan-gap-closure)
literal, require a variable reference, and require the step's prose to
bind that same variable name. Verified RED against the prior revision
and GREEN against this one before committing either state.

* fix(#2847): CRLF-safe tests, drop unexplained ack, require gap_closure=true

Four items from independent review, all landing together per request:

1. The #2847 regression test file had two CRLF-fragile regexes
   (local/no-crlf-fragile-split): a bare \n on readFileSync content
   means a real \r\n checkout returns invocationLine === null and all
   four executable-content assertions stop asserting anything while
   still reporting green. Both now use \r?\n. Prior lint report of
   exit 0 was a false green from a stale eslint cache.

2. The 2847 drift-ack fragment explained nothing: a direct edit to
   agents/gsd-planner.md is self-explaining, drift-acks exist for
   emitted-artifact ripple that cannot be traced to a changed source
   path. Deleted. Restored the 2775 fragment byte-identical to next
   (git diff --name-status next...HEAD -- tests/emitted-drift-acks/
   now prints nothing) — it only conflicted with the now-deleted 2847
   fragment, never needed touching itself.

3. plan-gap-closure validated gap_closure by PRESENCE only
   (unchanged since the original #2847 fix), so gap_closure: false
   satisfied it — --gaps-only filters strictly on gap_closure === true,
   so a false-valued plan still validates green and still spawns zero
   executors: #2847's exact reported symptom, one value away. Added an
   optional requiredValues map to FRONTMATTER_SCHEMAS; plan-gap-closure
   now requires gap_closure to equal the string "true" (extractFrontmatter
   parses every scalar as a string) in addition to being present. Every
   other schema/field keeps the original presence-only contract. The
   row that had documented the hole instead of closing it now asserts
   the fix; a matching unit test locks requiredValues on
   FRONTMATTER_SCHEMAS.

4. The "names the plain plan schema" assertion matched the bare
   substring "plan" anywhere in the step, which verify.plan-structure
   satisfies incidentally a few lines below — the assertion could not
   fail even if the plain-plan branch were deleted from the prose.
   Changed to match the standalone backtick-quoted plan token.

* fix(#2847): remove contradictory leftover assertion in Row 6 test

The gap_closure:false test asserted !present.includes('gap_closure')
(correct — matches the implementation's fold-wrong-value-into-missing
semantics) immediately followed by a stale, unedited leftover from an
earlier draft of the same test asserting the opposite:
present.includes('gap_closure'). The second could never pass once the
first did; both were in the same diff.

Verified before committing: searched every consumer of frontmatter.validate
output (agents/gsd-planner.md, docs/CLI-TOOLS.md, all other test files)
for any read of the present field — none exist. Nothing depends on
"present" meaning "physically exists regardless of value correctness",
so the implementation's fold (present/missing stay a full partition of
required) is the right call; the test needed to agree with it, not the
other way around. Manually replayed all six rows in the plan-gap-closure
describe block against the built CLI to confirm each now passes.

* fix(#2847): prototype-key guard, wrong-value diagnostic, doc fixes, vacuous tests

Six items from an independent SHIP_VERDICT:no review, landing together
per request:

1. Prototype-key crash (src/frontmatter.cts): FRONTMATTER_SCHEMAS[schemaName]
   was an unguarded lookup, so --schema __proto__ (also constructor,
   toString, hasOwnProperty, valueOf) resolved to an Object.prototype
   member instead of undefined, the `!schema` check never fired, and the
   command crashed with an uncaught TypeError and a stack trace instead of
   "Unknown schema". Now reachable from prompt state (--schema is an
   agent-bound $SCHEMA), not just an unreachable literal. Guarded with
   Object.prototype.hasOwnProperty.call before the lookup, checked and
   rejected before assignment so `schema`'s type stays non-optional. Added
   a test for all five prototype keys.

2. Wrong-value diagnostic (src/frontmatter.cts, agents/gsd-planner.md):
   the strict gap_closure === "true" check from the previous fix was
   correct (fail-closed) but silent about WHY — a plan with
   gap_closure: True got "missing", indistinguishable from genuinely
   absent, even though the field is plainly in the file. Added an
   `invalidValue` field to the validate JSON (present but wrong-valued,
   disjoint from missing/present) and updated validate_plan's prose to
   state the exact required literal and explain invalidValue, within the
   remaining byte budget (49130/49152).

3. docs/reference/plan-md.md: fixed three inaccuracies in the gap_closure
   row — "this field plus every field above" implied `requirements`
   (documented Required: Yes) is schema-enforced, it is not; "Type:
   boolean" implied YAML True/TRUE/yes/1 are accepted, they are rejected
   (exact string match on literal lowercase true); "must never carry it"
   stated an unenforced rule as fact. Also switched /gsd:plan-phase and
   /gsd:execute-phase to the house-style hyphen form for docs/.

4. Vacuous negative assertions (tests/fix-2847-gap-closure-frontmatter.test.cjs):
   RegExp#test coerces a null invocationLine to the string "null", so both
   hardcoded-literal checks passed vacuously even if the step or its bash
   block were deleted entirely. Added a truthy precondition check first.

5. Deleted vacuous/pass-always tests: four in tests/frontmatter.unit.test.cjs
   strictly subsumed by (or, for the "superset" test, tautologically
   guaranteed by the same spread as) the deepEqual exact-list test; two
   describe blocks in the #2847 regression file that were already GREEN at
   the RED commit (5e5897cd2f17ebf2fc55757bae651bbbeb236289) and pinned
   untouched files rather than covering anything this change altered — one
   of them additionally forbade any future legitimate gap_closure mention
   in plan-phase.md, a trap for whoever frees up that file's byte budget
   later.

6. .changeset/clever-newts-wake.md: switched /gsd:plan-phase and
   /gsd:execute-phase to /gsd-plan-phase and /gsd-execute-phase — changesets
   render verbatim into CHANGELOG.md with no converter in the path, so the
   colon form would have reached readers naming a command no runtime
   registers.

* chore(#2847): backfill changeset pr number (#3018)

---------

Co-authored-by: sim <sim@local>
2026-08-03 09:16:29 -04:00
Tom Boucher
f1af47766a chore(#1671): widen the when= grammar and key the section manifest per workflow — Phase 6.1 (#3013)
* chore(#2992): widen the when= grammar and key the section manifest per workflow

Epic #1671 Phase 6.1. Two blockers stopped the fragment model reaching any
file beyond execute-phase.md: the when= vocabulary was frozen at 4 atoms
(3 execute-phase-specific), and the section manifest was single-workflow by
construction with 'execute-phase' hardcoded into buildSectionManifestField.

- widen WHEN_VOCABULARY 4 -> 14 via a coordinated ADR-1671 amendment; the
  grammar stays CLOSED (one atom, no operators, negation or nesting) and
  WHEN_PREDICATES stays a hand-written literal map, never deriving a
  predicate from its atom string
- InvocationFacts gains flags: ReadonlySet<string> plus three computed state
  booleans; add the missing reverse vocabulary/predicate parity guard
- key the manifest artifact per workflow; a stale flat {sections:[...]}
  artifact now fails shape validation instead of being misattributed
- wire the field into six init entry points and parse the flags each needs

An atom ships only with both a real consuming section and a fact the init
seam actually computes. Six surveyed atoms are withheld because their
workflows have no dedicated init entry point; an atom without a computed
fact evaluates false forever and silently disables its own section.

Fixes a defect found while wiring: parseNamedArgs always materializes a
boolean flag key, so folding its false into the absent sentinel is required
or every flag reads as present and gating is silently always-on.

Also resolves ADR-1671:194 by measurement: --mvp stays unmarkable, because
its interleaved sites are always-run flag resolution and a ~340 byte block
that already delegates lazily.

Refs #2992

* fix(#2992): treat any falsy option value as an absent flag and reject unsafe manifest read paths

Findings from two orthogonal reviews (Claude /code-review + an isolated
adversarial pass); both independently reproduced the first one.

- MAJOR: the flags-builder treated only `undefined` as absent, but
  parseNamedArgs yields `null` for an absent value-flag and `false` for an
  absent boolean-flag, so `--granularity` read as present on every
  plan-phase invocation. Fixed at the root: a flag is present iff its
  option value is truthy. The six per-handler `|| undefined` folds are now
  redundant and removed, which also closes the duplicate-translation and
  missed-onboard-handler findings.
- MAJOR: state:needs-codebase-map had zero coverage. Added unit, property
  and real-CLI integration tests.
- MINOR: reject absolute, UNC/drive and `..`-traversing `read` paths in the
  manifest, degrading the whole load to null like every other shape
  violation. Verified: `/etc/passwd` previously reached section_manifest.read.
- MINOR: corrected a stale "4 to 20" doc comment; the vocabulary is 14.

Refs #2992

* test(#2992): update the generator suite for the per-workflow manifest shape

The remote matrix went red with 5 unique failures, identical on
linux-node22 and linux-node24, all in tests/gen-section-manifest.test.cjs.
Re-keying the artifact to {workflows:{...}} left this suite asserting the
old flat {sections:[...]} shape; nothing else in the tree still does.

- three tests read manifest.sections.length, now undefined; retargeted at
  workflows.<name> with their original intent preserved (a fenced or
  loop-host marker still asserts NO section is produced, not merely a
  changed count)
- the stale-manifest test wrote its fixture in the OLD shape, so it tripped
  shape validation and stopped exercising staleness at all. Its fixture is
  now valid-but-mismatched so FAIL_STALE is genuinely reached again.
- added the coverage that exposed: a pre-6.1 flat artifact must report
  FAIL_MANIFEST_MALFORMED_SHAPE. That is the real upgrade path for an
  installed tree and nothing covered it.

Refs #2992

* chore(#2992): backfill changeset pr number to 3013

---------

Co-authored-by: sim <sim@local>
2026-08-02 22:36:45 -04:00
Tom Boucher
fd07e1a357 fix(#2850): resolve the active workstream in the statusline GSD-state segment (#3012)
* test(#2850): add failing-first tests for workstream statusline state

readGsdState only ever reads the flat .planning/STATE.md via a directory
walk-up; it has no path for .planning/workstreams/<ws>/STATE.md and never
consults GSD_WORKSTREAM or the stored active-workstream pointer, so the
GSD-state segment silently disappears in workstream mode. These tests
prove the RED before the fix lands.

Uses shared saveSessionEnv/restoreSessionEnv/clearSessionEnv helpers now
added to tests/helpers.cjs (single source of truth for the session-env-var
save/clear/restore pattern also used by tests/active-workstream-store.unit.test.cjs,
which is updated here to consume the same shared helpers instead of its own
local, already-diverged copy).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2850): resolve active workstream in the statusline

readGsdState only ever walked up looking for a flat .planning/STATE.md; it
had no branch for .planning/workstreams/<ws>/STATE.md and never consulted
GSD_WORKSTREAM or the stored active-workstream pointer, so the GSD-state
segment silently vanished in workstream-mode projects with no root
STATE.md (exit 0, no diagnostic).

Reuses the existing CLI>env>store resolution seam (resolveActiveWorkstream,
active-workstream-store.cts) and the existing mode-detection/path-building
seam (listAvailableWorkstreams/planningPaths, planning-workspace.cts)
rather than re-implementing either inline. When workstream mode is
detected but nothing resolves, readGsdState now returns a
{noActiveWorkstream:true} sentinel that formatGsdState/formatGsdStateCompact
render as "no active workstream" -- observable, never silent emptiness.
Flat-mode behavior and the case where a resolved workstream has no
STATE.md yet are both unchanged.

Adds active-workstream-store.cts's peekActiveWorkstream: a read-only
sibling of getActiveWorkstream. resolveActiveWorkstream's default store
lookup self-heals a stale/invalid pointer by deleting it
(adapter.clear()) -- correct for a command, but not for a renderer
invoked once per prompt, which must never mutate persistent, possibly
cross-session state as a side effect of drawing a screen. The statusline
now injects peekActiveWorkstream via resolveActiveWorkstream's own
getStored override, keeping the env>store precedence itself fully reused
while removing only the store tier's write side effect. This satisfies
the issue's AC4 ("the fix is purely additive to what's displayed").

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2850): backfill changeset PR number to 3012

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 21:44:04 -04:00
Tom Boucher
de78f2eef2 docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate

security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md,
FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors)
described the pre-ADR-0656 design: slopcheck as the install-or-degrade
gate, with unavailability degrading every package to [ASSUMED].
ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/
crates.io) are the gate; slopcheck is an optional escalate-only adapter
that no shipped configuration wires. Verified every replacement claim
against src/package-legitimacy.cts (checkPackages, classifyPackage,
lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so
the corrected prose matches the live implementation rather than
restating the ADR from memory.

Restored docs/explanation/security-model.md:79-84 (and its ja-JP
mirror) to original wording after an orthogonal spec review caught
that an earlier draft had edited the "Why WebSearch packages are
always [ASSUMED]" paragraph — inside the range issue #2775 explicitly
named as correct and to leave alone.

The ja-JP mirror was missing the closing clause present in the
corrected English original ("its absence leaves registry-API verdicts
intact rather than downgrading everything to [ASSUMED]") — added for
parity. This completes the ja-JP mirror the issue's acceptance
criteria named explicitly.

zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale
design) get the mechanical portion of the same fix: command-string
swaps, table headers, ARCHITECTURE.md diagram labels, and technical-
term swaps that reuse a word already attested elsewhere in the same
file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose
untouched. The remainder in those three locales — full-paragraph
rewrites of the corrected degrade-path mechanism, deleted "External
dependency" bullets, and "manually install slopcheck" code blocks —
needs prose composed by a fluent speaker of each language and is filed
as open-gsd/gsd-core#3002 with an exact file:line inventory.

* test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix

agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE
supply-chain row correction (slopcheck -> package-legitimacy gate).
Emitted agent/workflow files are byte-tracked; this fragment
acknowledges the growth per tests/emitted-attribution.test.cjs's
"differential attribution over the real tree" check.

* docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading

docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration
"スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal
"slopcheck" grep, so it was missed when ja-JP parity was checked and
declared complete. Corrected to "正当性判定" (legitimacy verdict),
matching the term already established in ja-JP/explanation/
security-model.md and ja-JP/USER-GUIDE.md. This was the only
remaining ja-JP gap; a full sweep for the transliterated form across
docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely
at parity.

docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same
transliteration problem. Fixed inline to "적법성 판정:", reusing the
적법성/legitimacy word already attested two lines below in the same
table. A parallel sweep of zh-CN and pt-BR found no transliterated
forms of "slopcheck" in either locale. The remaining transliterated
occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the
pip-install code block) needs prose composition like the rest of that
block and is added to open-gsd/gsd-core#3002's inventory.

* chore(#2775): backfill changeset PR number to 3010

---------

Co-authored-by: sim <sim@local>
2026-08-02 20:24:42 -04:00
Tom Boucher
d770365753 docs(#2999): document the takeover process for a capability, reviewer lane, or EoS integration (#3000)
* docs(#2999): document the capability / reviewer-lane / EoS takeover process

The capability ecosystem documented a complete forward lifecycle — develop,
publish, version, import, update, remove, turn off — but nothing covering a
change of maintainer for an entry that already exists. Adds
docs/how-to/take-over-a-capability-or-eos.md defining four takeover modes
(consensual handoff, adoption fork, first-party absorption, retirement), the
PR shape each takes, a per-surface snapshot of the inherited user-visible
contract, and an install-continuity checklist.

Also corrects .github/PULL_REQUEST_TEMPLATE/registry-entry.md, which directed
contributors to a 'Registry' Discussions category that does not exist — the
category is named 'EoS Registry' per docs/registries/README.md, and because
'discussion' is a required field the thread must exist before the PR is
opened, so the wrong name stalled contributors at the first required step.

Closes #2999

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2999): backfill changeset pr number to 3000

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:02:58 -04:00
Tom Boucher
a987cf2731 chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init

Extends the init bundle with a typed per-invocation section manifest so an
invocation loads only the branch guidance it will actually take.

The three flag/state-gated branches in execute-phase.md move into their own
step files; the parent keeps its gsd:section markers wrapping a one-line
on-demand reference, so each section's prose lives in exactly one file and
the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator
derives the shipped section manifest from those markers, and a new pure
evaluator maps invocation facts to applicable section ids.

The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser
(Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary
and the predicate map stay exhaustively in sync.

Closes #2932

* fix(#2932): fail closed on prototype-chain when values

An isolated adversarial review found WHEN_PREDICATES[section.when] was a
bracket lookup on a plain-prototype object, so inherited Object.prototype
members resolved as predicates: "constructor"/"toString"/"valueOf"/
"hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and
"__proto__" threw an untyped TypeError carrying no .reason. Both violate
the module's documented fail-closed contract, and the manifest is read from
disk at run time so it cannot be assumed trustworthy.

Builds the predicate map on a null prototype and guards the lookup with an
explicit Object.hasOwn check. Adds table-driven coverage for nine
Object.prototype-shaped keys asserting the TYPED reason (asserting only
that it throws would still pass while broken) plus a fast-check property
injecting a hostile value at an arbitrary document position.

* test(#2932): retarget execute-phase step assertions at extracted step files

* fix(#2932): emit typed reasons for generator lib-load and write failures

* fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures

* chore(#2932): backfill changeset pr number to 2987

---------

Co-authored-by: sim <sim@local>
2026-08-02 12:34:41 -04:00
kyle-the-dev
ce38d44811 fix(#2777): remove stale codex local home metadata (#2831)
* fix(#2777): remove stale codex local home metadata

* chore(#2777): add changeset for codex local layout metadata

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-02 01:00:52 -04:00
Adnan
fc3fde05ee feat(#2646): surface unresolved deferred-items.md at milestone close (#2983)
* feat(#2646): surface unresolved deferred-items.md at milestone close

auditOpenArtifacts scanned eight categories; deferred-items.md was not
among them. #2287 made that file readable at the PHASE boundary
(audit-uat, /gsd-progress check 7), but one boundary up it stayed
invisible — and phase directories archive to milestones/vX.Y-phases/ by
default (#1871), so an out-of-scope discovery a phase agent correctly
recorded rather than fixed left the live tree at milestone close having
never reached the [R]/[A]/[C] prompt that exists to catch exactly this.

Adds deferred_items as a ninth scanner plus its count, its items entry
and its report section. The workflow needed no change: complete-milestone
branches on "any section with count > 0", so the new category flows
through the existing prompt.

The resolved/unresolved predicate is NOT reimplemented. uat.cjs already
exports parseDeferredItems, which owns the parsing rule (entries under a
`## Deferred Items` level-2 heading, else the whole file fail-safe;
RESOLVED only on an explicit case-insensitive `status: resolved` field).
The scanner requires it lazily, inside the scan, so audit-command-router's
property that a route never loads the module it does not need is
preserved. Two readers of one file sharing one predicate is the point —
duplicating the inequality is how they drift into disagreeing about what
"open" means.

Regression test proves fail-first: 9 of its 10 cases go red against the
pre-change tree. The tenth is the deliberate no-regression boundary (a
clean tree emits no section) and is green both ways.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

* docs(#2646): document the pre-close artifact audit and its nine categories

The /gsd-complete-milestone entry did not mention the audit at all, so
the gate that can stop a close was undocumented — and this change adds a
category to it. Tabulates all nine with their source artifact and what
makes each one "open", plus the [R]/[A]/[C] outcomes.

Also disambiguates the one genuinely confusing thing: the per-phase
deferred-items.md scanned here is NOT the `## Deferred Items` section the
[A] path writes into STATE.md. Same name, different artifact, opposite
ends of the flow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

* chore(#2646): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:27:33 -04:00
0xdhx
c61dd49d95 enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes

Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks
(decision: 'block', exit 2) a whole-file Write collapsing a curated
.planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md,
STATE.md) below 40% of its on-disk line count. Files under 40 lines
are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message)
bypasses for legitimate milestone resets.

Fix 3 of #973 — the only defense independent of per-agent tool config.
Registered on the Claude plugin surface (hooks.json), settings-json
runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi
spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and
INVENTORY regenerated; regression tests negative-controlled (16/16
RED with the hook absent, 16/16 GREEN with it present).

* chore(#2255): backfill changeset pr number to 2301

* enhance(#2255): address review — fail-closed reads, typed block output, registration, property test

Review fixes for trek-e's CHANGES_REQUESTED on PR #2301:

- Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES
  (no-shipping-drift test).
- Blocker 3: update the always-on hook enumerations in ADR-766 and
  CONTEXT.md from six to seven.
- Major 4: fail CLOSED on non-ENOENT read errors — only a missing file
  (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a
  typed readError field and the override still honored. Tested, with a
  negative control against the pre-fix hook.
- Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES
  budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the
  floor; sub-floor always exempt), boundary examples pinned.
- Major 6: block output now carries typed oldLines/newLines/
  overrideEnvVar fields; tests assert on those instead of regexing the
  free-form reason string.
- Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS
  bypass on macOS/Windows); limit+1 boundary tests added for both the
  floor and the ratio.

* enhance(#2255): engage the write guard on Kimi's native payload shape

The guard shipped with Claude-vocabulary checks (tool_name 'Write',
tool_input.file_path), which #2304 showed leaves a guard dormant on
Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli
forwards its native payload verbatim — tool_name 'WriteFile' (bare or
module-qualified) and tool_input.path per its tool schemas
(src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown
name, and exited 0.

Apply the same per-guard normalization PR #2326 gives the three
sibling guards (name + field mapping, inlined — hook scripts stage as
standalone files), and write the block reason to stderr as well as
stdout JSON: Kimi feeds stderr, not stdout, back to the model on
exit 2, so a stdout-only reason blocks without telling the model why
or naming the documented override.

Regression tests pipe Kimi-shaped payloads (engage, qualified-name,
stderr-reason) plus exemption pins (StrReplaceFile stays out of scope
by design; non-curated paths pass) — verified red against the pre-fix
guard, green after.

* enhance(#2255): rebase onto next; regenerate golden-parity fixtures

* enhance(#2255): wire the escape hatch into complete-milestone's reorganize step

Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP
reorganize — the tree's only legitimate milestone reset and the exact caller
GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the
rewrite through a shell write with the hatch set on the command (a hook
inherits the runtime env, so a bare Write cannot carry a per-step override),
and a binding test derives the env var name from the guard's typed output and
asserts (a) the workflow step sets it and (b) the guard passes the identical
catastrophic payload under it — so the next complete-milestone.md edit cannot
silently re-break the wiring.

* enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload

Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string
reconstruction were unreachable-by-effect — the guard exits 0 for any
tool_name !== 'Write', so nothing ever read the fields they set, leaving
guaranteed-surviving mutants against the Stryker bar. The map now carries
only WriteFile -> 'Write'; the StrReplaceFile exemption test message states
the fall-through it actually exercises.

* enhance(#2255): review minors — American spellings; writeSync before exit(2)

Minor 1: normalised/normalise -> American house style. Minor 2: the two
block paths wrote stdout+stderr via async pipe writes then exit(2) —
async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block
payload durable.

* enhance(#2255): assert stderr equals the typed reason, not raw prose

Minor 3: the last raw-text match in the suite pinned override-name prose on
stderr. The contract is "stderr carries the reason Kimi feeds back" — now
asserted as stderr non-empty and byte-equal to the parsed stdout.reason.

* enhance(#2255): bind the write-guard's Kimi normalization into the parity test

Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with
nothing binding it. This extends PR #2326's kimi-guard-normalization-parity
test (same path and helpers, authored as a superset so either merge order
resolves cleanly): sibling byte-parity is existence-gated zero-or-all —
trivially green until #2326 lands, full-strength after — and the write-guard
copy is bound semantically (map is the value-inverse of convertKimiToolName;
the Kimi name for Write must map, or the guard is dormant on Kimi; the
path -> file_path half must be present). Byte-parity is deliberately not
asserted for this copy: it legitimately omits the Edit-class mapping
(Major 1 — dead code in a Write-only guard).

* enhance(#2255): refresh golden-parity fixtures for revised guard + workflow

* chore(#2255): regenerate golden fixtures after rebase onto next

The committed fixture hashes were generated against a tree predating
next's latest 11 commits, which independently modified the same
install-parity surface. Rebased onto next and regenerated with
`npm run gen:golden`.

Verified: against upstream/next the regenerated fixtures differ by
exactly this PR's own entries -- hooks/gsd-write-guard.js (new),
hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and
gsd-core/workflows/complete-milestone.md. No unrelated drift.

* fix(#2255): regenerate workflow size baseline for complete-milestone

`complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2
review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step,
but tests/workflow-size-baseline.json was never regenerated. The
per-file workflow baseline test (issue #1074) failed on
ubuntu-latest/22 and both macOS shard 1/3 jobs.

The growth is justified: it is the escape-hatch binding requested in
review round 2 (the guard must not hard-block the tree's only
legitimate milestone reset), not incidental bloat.

Regenerated via `npm run size:baseline`; the diff is exactly the one
entry.

* chore(#2255): regenerate golden fixtures and size baseline after rebase onto next

* enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes

Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert
(no PreToolUse hook exists on Bash in this family; the write succeeded by
dodging the guard, not by the override firing) and the protection was prose.
The hatch is now a transport code consults: complete-milestone's reorganize
step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the
Write tool as the sanctioned path, and the guard — at the block point only —
verifies the sentinel is fresh (15 min) and names the pending target, then
CONSUMES it and allows that one write. Path-bound + single-use + freshness
keep it from becoming a standing unlock. The env var remains as the
interactive transport, where it can actually reach the hook.

Regression tests written first (negative control: 3 failed pre-fix): the
armed-sentinel Write passes and consumes; stale does not exempt; a token for
a different file neither exempts nor is consumed; the binding test now takes
the sentinel name from the guard's typed output (overrideSentinel), asserts
the step arms it, and asserts the step no longer routes the rewrite around
Write via a shell pipe.

Also in this commit, same file:
- m2: block emission is exception-safe — emitBlock() wraps both writeSync
  sites in their own try/catch that still exits 2, so an EPIPE can no longer
  convert fail-closed into the outer catch's fail-open.
- Header discloses the two reviewed design limits (cumulative sequential
  shrink; lexical match vs symlinked paths) per round-5 scoping.

* docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4)

USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the
runtime-hooks-surface registration comment, and the changeset now describe
both hatches — the single-use sentinel for workflow steps and the env var
for interactive use — instead of implying a per-step env can reach a hook.
The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)`
parenthetical the repo's fragments use (round-5 m4).

* chore(#2255): regenerate derived families on the rebased tree (full sweep)

Full generator sweep after rebasing onto next @ the body-parser-patched
lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every
regen delta verified to be either a PR-owned entry (gsd-write-guard.js,
complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's
committed value for entries our arbitrary-side conflict resolution had left
stale (all 18 runtime fixtures checked mechanically).

* test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync

The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY
retry budget; the sentinel disarm now rides it like every other teardown.

* chore(#2255): regenerate derived families after rebase onto next

Full sweep on the rebased tree (build -> gen-inventory-manifest ->
gen:golden -> size:baseline). Every delta is either a PR-owned entry
(hooks/gsd-write-guard.js, its registration surfaces
hooks/managed-hooks-registry.cjs and the two plugin buses,
gsd-core/workflows/complete-milestone.md) or exact convergence to
next's committed value across all 18 runtime fixtures.

* chore(#2255): regenerate derived families after rebase onto next @ a5180d96

Rebase onto current `next` (a5180d96) resolved 12 conflicting
golden-install-parity fixtures; all regenerated via the full generator
sweep (build, gen:golden, size:baseline) rather than a single generator.

`lint:generated-sync` reports every generated artifact in sync. All 45
differing fixture keys and the single workflow-size-baseline entry map
to files this PR actually touches; no foreign drift.

* fix(#2255): remove the stale unguarded reorganize_roadmap step (round-8 blocker)

complete-milestone.md carried a second ROADMAP-collapsing step,
`reorganize_roadmap`, distinct from the sentinel-armed
`reorganize_roadmap_and_delete_originals` this PR wired. It is a vestige
of the pre-archive-then-reorganize design: it sits BEFORE
archive_milestone, so executing it as written would collapse ROADMAP.md
before the archive snapshots the full phase detail — and its Write is
exactly the shape gsd-write-guard hard-blocks, with no hatch armed. The
file's own success criteria describe only one reorganize outcome
(Backlog-preserving, overwrite-in-place — the later step's properties),
and archive_milestone points forward to "the reorganize step".

Removed rather than wired, per the round-8 review's confirm-and-remove
option. A new binding test asserts the sentinel-armed step is the ONLY
reorganize step in the workflow, so an unguarded collapse step cannot be
silently reintroduced (negative-controlled: fails against the pre-fix
tree). Golden-parity fixtures and the size baseline regenerate for the
shrunk file; every changed fixture key is complete-milestone.md's own.

* test(#2255): document why the read-error injection is a path collision, not an fs monkeypatch

Round-8 nit: the non-ENOENT tests inject via a directory-at-target-path
collision instead of the repo's fs-method monkeypatch pattern. That is
deliberate, not drift — runHook exercises the hook as a spawnSync child
process, so an in-process fs.readFileSync patch (the pattern the cited
siblings use on require'd, in-process code) can never reach the code
under test. Record the reasoning at the injection site.

* chore(#2255): regenerate derived families after rebase onto next @ 0d08c320

Rebase onto current next (0d08c320) for the CONFLICTING/DIRTY state. All 32
conflicts were generated artifacts (19 golden-install-parity, 12 install-tree,
workflow-size-baseline); resolved arbitrarily and regenerated via a full
generator sweep (build, gen:golden, size:baseline, gen-inventory-manifest)
rather than hand-merged. No source conflicts.

Regen diff verified against the PR's changed-file set: 7 distinct differing
keys, all PR-owned (gsd-write-guard.js, managed-hooks-registry.cjs,
plugins/gsd-core.js, complete-milestone.md, and their .kimi mirrors).
lint:generated-sync clean.

* chore(#2255): regenerate derived families after rebase onto next @ 9138271b

Conflict set was 20 paths, every one a generated artifact, zero source
conflicts — resolved arbitrarily during the replay and regenerated here,
per the maintainer's round-9 recipe (never hand-merged).

Generator sweep (not just gen:golden): npm run build, gen:golden,
size:baseline, gen-inventory-manifest, gen:registry. INVENTORY-MANIFEST
came back byte-identical, so the merged value was already correct.

Regen diff verified == PR-touched entries: every differing leaf key
attributes to a file this PR changes (complete-milestone.md, the guard,
managed-hooks-registry.cjs, both plugin buses) — derived from the
changed-file set, not a feature-name filter.

* fix(#2255): port the #2595 Kimi-payload hardening the base landed for the sibling guards

Self-found via the round-10 base sweep: a8b40fa5 (#2547/PR #2595) hardened
gsd-prompt/read/worktree-path-guard's normalizeKimiPayload after this PR was
authored, and this guard carried all three pre-fix shapes:

- `path` now wins outright over a model-supplied `file_path` — kimi-cli's
  WriteFile schema carries no `file_path`, so a payload pairing a curated
  `path` with a spurious `file_path: ""` had the guard reading '' and
  exiting 0 while kimi-cli wrote to `path` (one-key bypass, no crash needed).
- null/primitive payloads pass through deliberately instead of throwing into
  the fail-open catch.
- the `file_path` read is typed (`typeof === 'string'`), closing the
  crash-to-allow on `[]`/`{}` — this is what the base-added
  kimi-guard-typed-payload-reads scan flags.

Negative-controlled: the shadowing payload exits 0 against the pre-fix guard
and 2 against this one.

* fix(#2255): round 9 Major 2 — stop printing the sentinel recipe in the agent-facing denial

The block message ended by telling the blocked agent exactly how to arm the
single-use sentinel — handing the #973 actor a mechanical self-authorization
in the moment it is blocked. The recipe is gone from the prose; the
corrective advice and the env override's name stay (the latter is a #2255
acceptance criterion, and a per-step env prefix cannot reach a hook anyway),
and the typed overrideSentinel field stays for the binding tests. The hatch
remains documented in USER-GUIDE.md and complete-milestone.md, where humans
and the workflow engine read.

* fix(#2255): round 9 Minors 1-2 — realpath-resolve the target before the curated match; disclose the /i Linux cost

Minor 1: a Write to a non-curated path that symlinks into a curated file was
not matched while writeFileSync followed the link — the target is now
realpath-resolved before the curated match (ENOENT keeps the lexical
resolution so new-file Writes still pass; any other realpath error falls
through to the read, which fails closed). Negative-controlled: the symlink
payload exits 0 against the pre-fix guard, 2 against this one. Test skips on
win32, where symlink creation needs privilege.

Minor 2: the header's design-limits block now names the unconditional /i
cost on case-sensitive Linux (a genuinely distinct .planning/roadmap.md is
also treated as curated) next to the stateless limit, and drops the closed
symlink limit.

* test(#2255): round 9 Minors 3-4 — CRLF counting pin + a passing Write leaves a fresh sentinel unburned

Minor 3: countLines' split('\n') is CRLF-safe for a count (the \r rides
along), confirmed by trace in the review — this pins it against this repo's
recurring CRLF regressions, on both sides of the compare and at the 40%
boundary.

Minor 4: consumeSentinelFor runs only after the ratio check would block, so
a within-tolerance Write never burns the workflow's token — true by
construction, previously un-asserted.

* fix(#2255): round 9 Major 3 — correct the stale env-var line in archive_milestone's summary

complete-milestone.md's "After archival" bullet still said the reorganize
happens "under GSD_ALLOW_PLANNING_SHRINK=1" — the wording from the round-2
design this PR's own history rejected in round 5 (a per-step env var cannot
reach a hook; setting it in a Bash step silently does nothing). It now points
at the sentinel mechanics the reorganize step actually documents, matching
that step and USER-GUIDE.md.

* docs(#2255): round 9 Major 1 — user-facing docs state the stateless per-Write limit

The changeset and USER-GUIDE described the guard as covering "catastrophically
shrinks" with no caveat, while the stateless design was disclosed only in the
hook header — an operator reading the shipped docs would conclude iterative
erosion is covered. Both surfaces now state the per-Write comparison and the
erosion non-goal explicitly, in line with what the guard does.

* fix(#2255): round 10 Blocker 1 — arm the hatch in the roadmapper's curated writes

gsd-roadmapper Step 7 Writes both .planning/ROADMAP.md and .planning/STATE.md
wholesale, and /gsd:new-milestone spawns it against the OUTGOING milestone's
files. new-milestone's `phases.clear` archives phase DIRECTORIES, never
ROADMAP.md, so nothing compacts it first, and no ordering rule forces
/gsd:complete-milestone to run before /gsd:new-milestone.

Measured against the shipped hook at the #973 file size (292 lines): a new
4-phase roadmap lands at 18.2% and an 8-phase one at 31.8% — both blocked; only
a 12-phase replacement (45.5%) clears. So the collapse is reachable on a
first-party flow, which is the failure mode #2255 para 4 says the guard must not
have.

Arms the single-use sentinel immediately before each Write, one per target
(the token is path-bound and single-use, so one arming cannot cover both), and
gates each arming on `[ -f ]` so the /gsd:new-project path — where the guard
already exempts via ENOENT — strands no unconsumed token.

Binding test derives the sentinel name from the guard's own typed output, so a
rename on either side fails the suite instead of silently unwiring the hatch.
Negative-controlled: 2 assertions fail against the pre-fix roadmapper.

* docs(#2255): round 10 Major 2 — scope the claim to the guarantee that ships

The sentinel is a plain file, so an agent that would reason past an advisory
can arm one with a single Bash call it is already permitted to make. The
"only defense independent of per-agent tool config" framing outran that, and
the changeset was on its way into CHANGELOG.md.

Retitles the claim on all three surfaces (changeset, guard header, USER-GUIDE)
to what the guard actually delivers: it blocks accidental and single-shot
collapse and is not a defense against a determined agent; what it converts is
"ignore a sentence" into "take one deliberate, path-bound, single-use,
auditable action".

Pinned by test on the DURABLE surfaces only — the guard header and USER-GUIDE.
The changeset fragment is deliberately not pinned: it is consumed at release,
so a test reading it would start failing the moment the release lands. The
bound-statement assertion normalizes comment markers and whitespace first, so
it pins the claim rather than the paragraph's line wrapping.

Negative-controlled: both assertions fail against the pre-fix surfaces.

* test(#2255): acknowledge the roadmapper growth from the round 10 Blocker 1 wiring

The emitted-attribution gate (#2719/#2767) flags gsd-roadmapper.md growing 1130
bytes without an acknowledgment. The growth is the Blocker 1 sentinel wiring
plus the rationale a future editor needs to keep it, so it gets an ack fragment
rather than a silencing regen — the gate's own message is explicit that there is
nothing left to regenerate.

Fragment is PR-scoped (2301-…) per the gate's naming instruction, and uses the
plain-string reason form the shipped fragments use.

Verified against the TRUE upstream tip, not the fork's origin/next: a stale
origin made this same gate report unrelated phantom drift (1 emitted path + 6
grown files + 5 stale acks) that vanishes when GSD_EMITTED_BASE is pinned.

* test(#2255): renumber the roadmapper PROSE_ALLOWLIST pin after the Step 7 wiring

CI red on shard 2/3, all four platforms. The #2751 gate keys PROSE_ALLOWLIST on
{file, line}; the Blocker 1 wiring added 18 lines above the allowlisted
parenthetical in agents/gsd-roadmapper.md, moving it 624 -> 642. Both halves of
the gate then fired: the moved line reads as a new offender, and the stale
entry no longer matches anything.

Line content at 642 is byte-identical to what the entry describes — a
descriptive "e.g." naming SDK queries a user could run — so this is a
renumber, not a re-classification.

Swept the defect class rather than the instance: agents/gsd-roadmapper.md is
the only line-pinned reference to any file this round changed.

Negative-controlled: both assertions fail against the un-renumbered allowlist.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:19:49 -04:00
0xdhx
cc3ee301a7 fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root

installSharedHooksBundle wrote `{"type":"commonjs"}` over
<configRoot>/package.json unconditionally — no existence check, no merge,
no backup — on every install and every /gsd-update re-install. On the 11
affected runtimes that file is often user-owned; on OpenCode and Kilo it is
the documented place to declare local-plugin npm dependencies, so a user's
name/type/dependencies/scripts were destroyed on each run.

The uninstall path already read the file and unlinked it only on an exact
content match. That asymmetry was the defect: the discipline existed in the
codebase, it just was not applied on the write side.

Move the marker into the directories GSD creates and fills with its own .js
files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the
nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop
writing the config root entirely. New src/commonjs-marker.cts owns the marker
string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed
on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker,
so install and uninstall cannot drift apart again.

Nothing else depended on the config-root marker: package identity is baked at
build time (#378/#498) and version resolution prefers gsd-core/VERSION and
already tolerates a missing root package.json (#1383) — Codex has installed
without one all along. A package.json in plugins/ or extensions/ is inert to
plugin discovery, which globs *.{ts,js} only (see installer-migration 006).

Uninstall retires the pre-fix config-root marker, so upgrading users are
cleaned up on removal, and still never touches a file it did not write.

* fix(#2544): point the changeset fragment at the filed PR

The fragment's `pr:` field is only knowable after `gh pr create` returns.

* fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set

bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the
linted source), so it belongs in the ADR-457 ignore list like its siblings.
Clears the lint-tests no-var failure and the repo-invariants
"linted xor ignored" migration-state test.

* fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root

The UPGRADE 1 test still asserted the pre-#2544 marker location
(~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the
directory GSD itself creates — matching the updated golden-install-parity
and install-tree fixtures. Also asserts the root marker is NOT written.

* fix(#2544): make the CommonJS marker write path non-fatal

Review round 2, Major 3 + Minor 1 + the stagedHooks nit.

ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site
caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole
install with a raw stack trace. Every other marker interaction in the module is
best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker
swallows read failures — and this was the write path, i.e. the one most likely
to fail on a locked-down config dir. It now returns a new 'failed' outcome and
both call sites warn and continue.

Sibling found while sweeping for the same defect class: fs.mkdirSync sat
OUTSIDE the try block, so an unwritable parent threw past the guard entirely.
Creating the directory is the same environmental hazard as writing into it, so
it moved inside.

Also in this file:

- The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks
  alone. stagedHooks is computed from the SOURCE listing before the copy loop,
  so it stays true when the copies land but verifyInstalled() then fails —
  marking a hooks/ GSD did not successfully populate claims an ownership the
  install did not earn.
- The uninstall rmdir of the native plugin dir is gated on GSD having actually
  removed something from it. Hoisting it out of the adapter-exists guard (so
  the marker-only case could prune) had silently widened it into deleting a
  user-created but empty plugins/ or extensions/ dir — the same "don't touch
  territory GSD didn't fill" principle this issue is about, inverted.
- Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the
  same call site that writes its replacement. That path is outside kimi's
  configDir, so installer-migration 007 structurally cannot reach it.

* fix(#2544): retire the stale config-root marker via installer-migration 007

Review round 2, Major 1 — the PR's headline claim was false for existing
installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale
{"type":"commonjs"} at the config root, so their config root stayed pinned to
CommonJS and their dependency manifest stayed gone until they uninstalled.

The migration is unusual in one way, and it is the part worth reviewing: the
config-root marker was never recorded in gsd-file-manifest.json (writeManifest
records hooks/, agents/, commands/, scripts/ and the native plugin, never a root
package.json), so classifyArtifact answers 'unknown' for it and the planner's
own guard downgrades a remove-managed on an 'unknown' classification to
preserve-user. 007 therefore supplies the "purpose-built detector for an old
GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions —
exact content match, the same predicate removeCommonJsMarker has always used —
and declares the resulting classification on the action. A package.json with any
other content is left untouched, and there is deliberately no backup-and-remove
branch: a non-matching file here is not a patched GSD artifact, it is somebody
else's file.

Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`:
validateStringArray requires the field to be non-empty WHEN PRESENT, while the
runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan
time and the migration never runs. The metadata test pins this.

Kimi is a deliberate carve-out, named in the migration's own header: its marker
lived at ~/.kimi, outside kimi's configDir, and migration relPaths are
structurally confined to configDir. It is retired by the installer instead.

Registration: shipped-migrations table, .gitignore for the emitted .cjs, the
EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not
copied from migration 006 by rote — 006 needs no entry because it imports
nothing, while 007 imports node builtins, so tsc emits its __importDefault
helper and the `var` in it trips no-var. This is the same lint gate that made
round 1 red.

* test(#2544): fault-injection and multi-runtime marker coverage

Review round 2, Major 2 + Minors 4 and 5.

Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and
the suite had no fs monkeypatching at all. Every branch now covered is one whose
doc comment claims it as the module's safety posture:

- classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule),
  with an ENOENT control alongside it so the test discriminates rather than
  just asserting one side
- classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never
  downgrades to the permissive answer) — the fixture's bytes are exactly GSD's
  marker, so the test fails if the code ever answers on content it could not read
- a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already
  covered with a real symlink, the directory case needs no injection at all)
- the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx'
- the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the
  mkdirSync that used to sit outside the guard
- removeCommonJsMarker unlink throw -> false

These save and restore fs methods in `finally` rather than using chmod 0o000,
which does not fault under root and would pass vacuously in root Docker and CI.

Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both
kimi locations now have behavioral coverage, install and uninstall, each paired
with a user-authored-file case proving GSD leaves it alone.

Minor 5 — the stagedHooks gate had no assertion behind its stated reason.
A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime
that declares skipSharedHooksInstall and asserted to stay marker-free, with its
user content intact.

Also regression-tests the uninstall rmdir gate from the previous commit: an
empty plugin dir GSD removed nothing from must survive.

* docs(#2544): correct stale marker prose, register the module, document the trade-off

Review round 2, Minors 2, 3 and 6.

Minor 2 — six files asserted the installed ROOT ships the synthetic marker.
None was load-bearing (all three walk-up consumers are VERSION-first with
try/catch and the marker never carried a `version`), but ADR-457:52 is the
rationale for keeping a generated module, so a future reader would mis-derive
the constraint from it. Each site is corrected to what is now true: the
installed tree carries no package.json with a .name at all, because the only
ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own
directories.

Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js
and the platform-gate test both described `require('../package.json').name`
resolving to undefined; post-#2544 that require does not resolve at all, so the
history is kept accurate and the present-tense claim corrected rather than just
moved. And src/runtime-artifact-conversion.cts described the no-root-package.json
case as Codex-only — it is now every runtime, which strengthens that comment's
own argument for lazy resolution. The generated .cjs sibling needs no edit: it
is gitignored build output, not a tracked file.

Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer
module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added,
including the fail-closed posture and the never-throws contract.

Minor 6 — the plugins//extensions/ marker shadows the config root for all .js
siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is
exactly what #2544's Fix section prescribed and it is disclosed in the PR body,
but the PR body is not documentation. It now lives in the OpenCode section of
docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than
a pure improvement, with the .ts mitigation and a fallback for ESM plugins.

* test(#2544): attribute the CommonJS marker in the emitted-provenance rules

The differential emitted-attribution gate (#2723, landed on `next` after this
branch was cut) went red on the macOS shards once this PR rebased onto it. Two
distinct causes, both real gaps rather than noise:

1. `plugins/package.json` and `extensions/package.json` matched NO rule — the
   `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an
   unattributed emitted family.
2. `hooks/package.json` fell through to `hooks-built`, which attributes an
   emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no
   `hooks/package.json` in the repo, so it resolved to a nonexistent path.

Cause 2 is exactly the failure already documented three lines above it for
Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix
follows that precedent rather than inventing one: `package.json` is excluded
from `hooks-built` the same way, and a dedicated `commonjs-marker` rule
attributes the family across all four roots it can appear in (both hooks roots
plus `plugins`/`extensions`) to the sources that actually emit it.

Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for
a one-off ripple and goes stale by design — the gate fails a stale ack precisely
so it cannot pre-clear the next change on that path. These markers are a
permanent part of the emitted tree from #2544 onward, so they need standing
attribution.

Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3
provenance errors + 12 unattributed paths before, 35/35 green after.

* fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker

#2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned
in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via
dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's
module on the two properties that matter:

  - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a
    DANGLING one, so a dangling package.json symlink classified as absent and
    the write went straight through it. Demonstrated: against the pre-fix copy,
    ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json
    symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory.
  - create: a plain writeFileSync leaves the classify->write window open, where
    commonjs-marker creates with flag:'wx' (O_EXCL).

Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's
own docstring already claimed was the single place these rules are enforced.
Exported signatures are unchanged (still boolean), so bin/install.js and the
#2717 tests are unaffected.

The new subtest is the only coverage that fails if the duplicate is ever
reintroduced — the two implementations agree on every non-adversarial input, so
the existing suites pass against both.

* test(#2544): pin the stagedHooks gate on zcode, not windsurf

The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall
kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the
marker was correctly absent.

#2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via
dedicated paths and get the marker beside those scripts. Measured on this tree,
windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning
behaviour that is now wrong, not the gate it was written for.

ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin
surface to spawn hooks, so GSD stages no .js there by either route (measured: 0
staged, no marker). The property under test is unchanged — a user-created hooks/
directory GSD never fills stays marker-free.

* test(#2544): use the shared cleanup helper in the migration test

Addresses the review's Major 1. The suppression's stated reason — "no helpers
import available" — was not correct: tests/helpers.cjs exports cleanup, and the
other test file added in this same PR imports it (tests/commonjs-marker.test.cjs).

The local reimplementation dropped two protections that are live on this repo's
windows-latest lane: the CWD guard (Windows cannot remove a directory that is the
current working directory) and the 20 x 250ms retry budget that absorbs the
deferred-scan handle Windows Defender holds on newly-written files.

Local function and suppression both removed; local/no-raw-rmsync-in-tests now
passes without one.

* test(#2544): expect hooks/package.json for the #2717 runtimes

The fresh-install contract table predates #2717, which stages cursor/windsurf/
codex .js hooks via dedicated paths and writes the CommonJS marker beside them.
All three therefore now receive hooks/package.json legitimately.

Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with
the marker; cline/copilot/trae/zcode stage none and get none, so their contracts
are unchanged.

* fix(#2544): gate the #2717 marker writes on having staged something

The three dedicated marker writers #2717 added ran unconditionally. Each one
mkdirs hooks/ up front and stages its scripts conditionally on the source
existing, so with an absent or empty hook source they created a directory,
filled it with nothing, and marked it as GSD's anyway.

That is the same write-into-someone-else's-territory this issue is about, and
installSharedHooksBundle already guards the identical case with `stagedHooks`.
The dedicated paths now carry the matching gate:

  - cursor / windsurf: `installedScripts.size > 0`
  - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that
    hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY
    entry landed.

Covered for cursor and windsurf by driving each writer against a src tree whose
hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its
trigger state needs a package tree where hooks/dist exists but holds none of the
allowlist, which is not constructible from a real checkout.

* test(#2544): scope the commonjs-marker sources per root

The rule declared one flat source list for every marker root, so
`extensions/package.json` was attributed to runtime-hooks-surface.cts (which
never writes there) and `.kimi/hooks/package.json` to install-engine.cts.

That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source,
so a flat list containing bin/install.js let any change anywhere in that
13k-line file authorise marker drift for every root — the blanket escape hatch
this file's own agents-verbatim comment refuses for exactly the same reason.

Sources are now derived per root from ctx.rel. Note the rule ctx is
`{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent
every path down one branch silently.

* test(#2544): state precisely what the zcode assertion pins

The comment claimed the test pinned installSharedHooksBundle's `stagedHooks`
gate. It does not, and neither did the windsurf version it replaced: zcode
declares skipSharedHooksInstall, so the outer guard skips that helper entirely
and the gate is never evaluated. The test passes on the runtime exclusion.

What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays
marker-free — is still worth having, and is what the review asked for. The two
`staging zero hook scripts` tests are the ones that pin a real staged-nothing
gate. Comment corrected rather than left implying coverage that is not there.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:00:23 -04:00
Tom Boucher
33fd203ccd test(#2966): loop QA walk — drive real scenarios across all five loop steps (#2976)
* test(#2966): loop QA walk — drive real scenarios across all five loop steps

Adds a headless walk that carries accumulating project state across
discuss -> plan -> execute -> verify -> ship against one temp project,
layered over the existing tests/helpers.cjs runGsdTools substrate.

Findings carry severity. A violation breaks a stated contract and fails
the build; a smell is legal under today's implementation but structurally
questionable, is recorded, and never reddens CI. Without that split an
oracle set derived from current behavior can only ever confirm current
behavior -- the harness could not say "this works and is still wrong".

The end-to-end test asserts the walk produces at least one smell: a QA
harness that reports nothing on a first run against a real engine is far
more likely mis-specified than the engine is perfect. It deliberately does
not pin smell ids or counts, which would re-freeze current behavior.

First run against the real engine: 0 violations, 3 smell classes --
init returns agents_dir outside the project tree; smart-entry emits prose
unconditionally so routing cannot be asserted; state-snapshot reports a
missing STATE.md through a payload key with exit 0.

Also fixes tests/fixtures/index.cjs: createFixture with git:true and
planning:false staged nothing, so the commit failed with "nothing to
commit". That combination was unreachable until greenfield needed it.

Extends RULESET.TESTS.feedback-loop-convergence from estimation to the
loop itself. Design lock: docs/adr/2966-loop-qa-walk.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): wire fault injection, make perturbations discriminating

Independent review found tests/qa/mutations.cjs entirely unwired: 462
lines exercised only by their own unit tests, with no mutation hook in
the scenario DSL and no scenario applying one, while the module header
and the ADR described fault injection in the present tense. Dead code
documented as live.

Adds a `mutate` step field, three perturbation scenarios, and a wiring
detector: a self-test scenario whose expectations are known-false and
which MUST fail. The previous anti-vacuity check asserted only that the
walk produced a smell, which passes on well-known engine behavior
regardless of whether the harness wiring works.

First perturbation attempt produced zero signal -- progress does not
structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A
perturbation that cannot fail is the same defect in a new costume.
Probes now target roadmap get-phase, and each mutated step runs a clean
baseline first so `mutationObserved` records whether the corruption
changed anything at all.

Also clears four review findings: classify() returned PROSE for exit-0
with empty stdout; `warnings` was structurally unpopulatable on the
success path (execFileSync discards it) and is now documented as
error-path-only; read-only-idempotence passed vacuously when asked to
check idempotence without the data to check it; the ADR miscounted the
oracles.

Discrimination matrix across 8 mutations x 6 commands: bom,
duplicate-phase-id and escaped-pipes are absorbed silently by every
probed surface, and progress / smart-entry / roadmap validate never
reacted to any mutation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): add path-containment guard for scenario-supplied targets

Security review found scenario-supplied paths joined to the temp project
with no containment check. step.mutate.target and agent.write keys were
validated only as non-empty strings, so a target of ../../../../etc/hosts
reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the
project. The symlink mutation was worst: it read the traversed file, wrote
a sibling copy, deleted the original and symlinked it back.

Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md
and scenarios are repo-committed, not runtime input. Fixed anyway: it is a
live primitive any future scenario or copied helper can reach.

Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL
bytes and empty input, normalizes separators unconditionally, and requires
containment by path segment so a sibling like <base>-evil is not treated as
inside. Non-existent targets resolve via nearest existing ancestor rather
than falling back to a lexical compare. Scenario load now rejects traversing
or absolute targets up front.

oracles.cjs previously carried its own copy of the containment logic; both
now share paths.cjs, since a duplicated containment check is exactly the
divergence class this repo calls out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): complete trajectory corpus, report emission, boundary-aware oracle

Adds the remaining trajectories and drives all 11 mutations end-to-end.
20 scenarios, 72 steps, 0 violations, 25 smells.

Adds qa-report.json with per-step verdicts and a copy-pasteable repro
command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro
line for a tree that was not preserved is marked NOT RUNNABLE rather than
emitting a command pointing at a deleted directory.

monotonic-progress is now boundary-aware. Two scenarios had been trimmed
to stop the oracle complaining at a milestone rollover, which destroys the
signal the trajectory exists to produce. Evidence: counters legitimately
reset to zero at milestone complete, but the payload milestone_version
lags until a new ROADMAP.md is written. So the oracle now scopes by
milestone plus workstream, keeps a same-scope decrease as a violation, and
records a boundary crossing as a smell. Both scenarios walk the real
boundary again.

Standards review fixes: oracle findings now carry a structured subject so
tests assert on typed fields instead of substring-matching the free-form
detail string, resolveWithin throws a typed EPATHESCAPE error, and the
absolute-path predicate scenario.cjs had re-implemented now comes from
paths.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): fix silently-vacuous fixtures and guard the class

Every fixture carried its #2371 provenance comment BEFORE the frontmatter
block, and extractFrontmatter returns {} when anything precedes the opening
---. So every scenario reading status/phase/name was operating on an empty
object and reporting green. Nine fixtures repositioned; the comment stays,
it just moves below the closing ---.

Both UAT fixtures lacked a parser-recognized result block, so
evaluateUatPassed saw checks.length===0 and could never return passed:true.
The uat-fail-then-remediate scenario could not have proven a remediation.
Its expect block only inspected blockers, which is empty before AND after,
which is why the corpus never noticed. Both fixtures now carry real result
blocks and the scenario asserts passed and no_uat_artifacts on each side of
the flip.

The actual deliverable is the guard: a fixture-integrity block asserting
every fixture with a frontmatter shape parses to a non-empty object, that
every fixture carries its provenance marker, and that the two UAT fixtures
produce opposite verdicts through the real evaluateUatPassed. The first
guard written required --- at byte 0, which would never have fired on the
regression it exists to prevent; it was rewritten and proven by deliberately
re-breaking a fixture.

No engine defect here. no_uat_artifacts means no parsed check items, not no
UAT files, and it was reporting correctly on fixtures that had none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): make the walk report — smell ratchet, baseline, CI job

The harness computed smells into a gitignored qa-report.json that nothing
read. In CI it surfaced nothing at all: violations failed the build, but the
half of the tool that says "this works and is still wrong" was inert. A QA
tool nobody hears is decoration.

Adds a ratchet on the same idiom this repo already uses three times over
(the regression-test-name allowlist, the emitted-drift acks, the size
baseline): a committed smell-baseline.json, per-PR acknowledgment fragments
under tests/qa/smell-acks/, and a ratchet script wired into CI.

The design invariant is preserved exactly. A smell still never fails a build
on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence
of a decision -- leaving an author two honest exits: fix it, or record a
fragment with a real reason. An empty reason is rejected. The baseline is
shrink-only, so a fixed smell must prune its entry. Violations remain
unacknowledgeable.

Fingerprints are composed only from stable fields (oracle id, scenario,
argv, subject discriminator) -- never temp paths, timestamps or counts.
Verified byte-identical across two runs in separate temp dirs; an unstable
fingerprint would have false-positived every CI run.

CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the
report with `if: always()` (it matters most when it failed), and renders a
summary a reviewer reads without downloading anything.

Also fixes the report runner invoking main() unconditionally on require, so
importing it double-ran every scenario and clobbered its own output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): every smell terminates in a defect or a fixed detector

The baseline accepted a smell with a free-text reason. That is a mechanism
for designing smells in -- an allowlist nobody revisits. The harness is
brand new, so nothing it found is inherited legacy; every finding is a
FIRST finding. Each must now terminate in exactly one of two states:

  REAL           -> an assigned defect, entry carries the issue number
  FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined

There is no third "accepted with a good explanation" state, so the ratchet
now requires a positive-integer `issue` on every entry. A reason may remain
as a human note but can never substitute. `--update` refuses to invent
issue numbers: a new smell is written with `issue: null` and a TODO, and
the next plain run rejects it, forcing triage rather than accumulation.

Working the 21 existing entries through that rule found 16 were my own
detectors being wrong:

value-hygiene (10) flagged $.agents_dir, a field whose entire contract is
to point at the install tree outside any project. Fixed with a leaf-key
allowlist of contractually-external fields, verified as the only such key
in the init payload. Genuinely unexpected out-of-project paths still smell.

monotonic-progress (6) fired on legitimate boundary crossings -- milestone
v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no
scope fields at all, where a change cannot even be known. Scope changes now
reset silently and scope-less observations are skipped. The same-scope
decrease remains a violation; that is the real invariant and is regression-
guarded.

The five survivors are real and now tracked: soft-error-exit-zero (#2980),
untyped-success (#2979). Baseline 25 -> 5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): keep the ratchet out of the tarball, unpin the qa CI job

The remote matrix returned failed -- 3 unique failures, identical on
node22 and node24, both root causes in this branch's own diff.

The ratchet lives under scripts/, which ships in the npm tarball, and it
requires three modules under tests/, which does not. In a published
install it is MODULE_NOT_FOUND at load. This is exactly the class the
#2858 guard was added to catch, and it caught it. Fixed the way #2858
fixed the same shape for its own repo-only CI script: a targeted files[]
negation, so the ratchet stays in the repo for CI and out of the tarball.
Not solved by moving or inlining the required modules -- the ratchet must
keep using the same code the harness uses, or the two drift.

Verified both directions: the script is no longer in the pack list, and
build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are
all still shipped. Over-negating there would have broken installs, since
bin/install.js requires them.

The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a
neighbouring job without the paired GSD_EMITTED_BASE, which the #2854
invariant forbids by name: diverging them makes the differential compare a
tree against a baseline from a different commit. The job runs only the qa
suite and the ratchet and invokes no emitted-attribution test, so it needs
no rebase-pinned base at all -- the step was removed rather than paired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): stop monotonic-progress going blind on scope-less payloads

The full remote suite caught a false NEGATIVE I introduced while fixing a
false positive. Silencing the boundary-crossing noise had made the oracle
skip ANY observation lacking milestone fields -- so a minimal payload like
{total_summaries: n} produced no violation at all, and the oracle stopped
catching the exact defect it exists to catch. For a QA tool that is
strictly worse than the noise it replaced.

Scope is only indeterminate when the two observations DISAGREE about
having it:

  both scoped, same scope, decrease -> VIOLATION
  both scoped, different scope      -> reset silently
  NEITHER scoped, decrease          -> VIOLATION   (the regression)
  mixed                             -> skip the comparison

Implementing the mixed case surfaced a second blind spot: advancing the
reference point on a skipped pair lets a scope-less observation sitting
between two same-scope ones mask a real decrease. Mixed now leaves the
reference untouched. All four branches carry explicit coverage; only one
did before, which is why this shipped.

The self-test that failed was right and the code was wrong, so the code
moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0
violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:13:28 -04:00
Tom Boucher
628648d63a chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes

Route every runtime brand swap through applyClaudeCodeBrandSwap so
"Claude Code" survives verbatim inside <runtime_compatibility> regions
(#2284b). The fix existed only in bin/install.js's local copies; the
src/*.cts exports still used a naive replace, so binding install.js to
the single source -- as this phase does for the Windsurf family --
would have silently regressed those runtimes. A table-driven parity
guard now covers all nine brand-swapping converters.

De-duplicate the Windsurf converter family: delete the six local copies
in bin/install.js and bind the four exported ones by reference, guarded
by reference-identity assertions (the ADR-1508/#1675 pattern). The two
unexported helpers and an unused tool table go with them.

Replace the Windsurf 12,000-byte throw with description truncation,
matching the bound its sibling skill converter already applied. The
throw could only fire on an ~11.7 KB frontmatter description: the
largest emitted workflow is 311 bytes. Truncation makes the cap
unreachable by construction and leaves 12,000 in exactly one place,
eliminating the dual-surface duplication rather than testing for it.

Add the emitted-byte cap gate: buildEmittedSizes captures LF- and
<HOME>-normalized bytes from the walk buildParityManifest already
performs, and evaluateEmittedCaps asserts them against a per-runtime
cap table with dead-rule detection. buildParityManifest's return shape
is deliberately unchanged -- diffEmitted compares its values with
===, so making them objects would report all 8,529 emitted paths as
moved. A regression test pins the values as strings.

Add a deterministic trim-safety gate over composeWithinBudget's
omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity
rule, replacing the model-graded eval gate the issue described.

* docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract

* fix(#2931): bound the windsurf command name and single-source the brand swap

Review findings from the orthogonal passes, all fixed inline.

The claim that removing the 12,000-byte throw left total emission
"bounded by construction" was false. The #1615 regex constrains the
character class but not the length, and commandName is interpolated
three times into the emitted workflow: a 20,000-character name emitted
60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate,
clearly-labelled size control that THROWS -- commandName is the @-ref
path target, so truncating it would point the workflow at a file that
does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The
#1615 security regex is untouched and still runs first. 128 is generous:
the longest shipped name is gsd-plan-review-convergence at 27.

Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe
truncation helper. It still used a UTF-16 slice(0,177) -- the exact
surrogate-splitting bug the helper was written to avoid, in the very
sibling the helper's comment cites as its model. Bounds are unchanged,
so output is byte-identical for every shipped command (descriptions max
out at 99 chars).

Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting
the local copy. Adding it to the .cts left two unlinked implementations
of identical logic -- the drift class this change exists to remove.
Verified byte-identical across eight fixtures and five sequential calls
before merging, and guarded by a reference-identity assertion.

Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344),
add fast-check property coverage for the trim-safety contract, and use
fc.pre instead of a bare return in a property callback.

* test(#2931): fix three test-authoring bugs the remote matrix caught

The remote runner returned 8 unique failures on 6f15cdeb8. All three
causes were in the test files, not the modules under test -- local
harnesses exercise the modules directly, so nothing executed the test
bodies until the matrix did.

`{ __proto__: [...] }` in an object literal sets the prototype instead
of an own key, so the JSON round-trip erased it and the cap table never
saw a reserved runtime key. The production rejection was already
correct; the test could not reach it. Use a computed key.

Two cap fixtures tripped orthogonal error paths rather than the paths
they name: one declared windsurf in the cap table but omitted it from
sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern
matching nothing (a genuine dead rule). Both now include a compliant
artifact so the intended branch is what is asserted. The dead-rule and
unknown-runtime contracts are deliberate and unchanged.

`const { root } = makeSyntheticConfig({ ... `${root}` })` referenced
`root` from inside its own initializer -- a temporal dead zone error.
makeSyntheticConfig now optionally takes a (root) => files factory.

Also raise the npm pack --dry-run bound 60s -> 120s in the shipped-
scripts packaging test. That failure is NOT from this branch: the file
is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s
here, and the run recorded 60,637ms against a 60,000ms bound -- a
timeout under 28,948-test parallel contention, not a slowdown. Fixed
rather than deferred because a bound that tight is fragile regardless
of which branch trips it.

* chore(#2931): backfill changeset pr number to 2984

---------

Co-authored-by: sim <sim@local>
2026-08-01 16:00:14 -04:00
Tom Boucher
640eaee16e chore(#2930): fragmentize execute-phase.md and prove per-runtime composed emission (#2972)
* feat(#2930): fragmentize plan-phase.md workflow into per-runtime-composed sections

Adds src/workflow-fragments.cts (in-file <!-- gsd:section --> marker
parser/composer, ADR-1671 epic #1671 Phase 3), wires it into
bin/install.js's copyWithPathReplacement emission path, and pilots the
marker grammar on gsd-core/workflows/plan-phase.md.

Bookkeeping ripple for the new src/*.cts module: .gitignore,
eslint.config.mjs, docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json,
and a CONTEXT.md glossary entry. Amends ADR-1671 with open questions 1
and 2 resolutions and records the closed when= applicability grammar.
Adds docs/reference/workflow-fragments.md and an ARCHITECTURE.md
section documenting the marker authoring model.

* fix(#2930): put allow-test-rule issue ref on the same line as the marker

lint-allow-test-rule-refs.cjs requires the #NNN issue reference on the
same source line as `allow-test-rule:`; it was one line below and read
as an unreferenced novel exemption.

* docs(#2930): link the orphaned gate-predicates reference from the docs index

Found while adding the workflow-fragments reference doc: docs/reference/gate-predicates.md
shipped without an entry in docs/README.md, so it was unreachable from the docs index.
Fixed inline rather than deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): scope composition to workflows, add typed failure reasons

Review findings from two orthogonal passes:

- Scope composeWorkflow to gsd-core/workflows/ only. It previously ran on
  every .md the installer copied, so a future agent/command/reference doc
  documenting the marker syntax with an unfenced example would have been
  mis-parsed and silently stripped — a lossy drop the phase forbids.
- Add a frozen REASON enum; failures attach a typed .reason and tests assert
  on it instead of matching free-form message text (CONTRIBUTING.md:635-694).
- Derive the property generator's when= values from WHEN_VOCABULARY instead
  of duplicating them (DEFECT.GENERATIVE-FIX).
- Add adversarial parser fixtures: Unicode headings, NUL, U+FFFD, BOM,
  fence-within-fence, tilde and indented fences, lone-CR marker line.
- Document why --mvp is structurally unmarkable: its content is interleaved,
  not sectioned, so the whole-line grammar cannot reach it.

Also fixes two stale tests on this branch, each reproduced on the unmodified
tree before correction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): retarget the pilot from plan-phase to execute-phase

The full remote matrix went red on both Linux lanes. Root cause was ours:
tests/phase6-capstone-conformance.test.cjs holds a PRE_PHASE6 ceiling of
94519 bytes for plan-phase.md, asserting an ADR-857 Phase-6 completion
property. That is a third size gate beyond the tier caps and the
differential ratchet, and it left plan-phase.md just 36 bytes of headroom
rather than the 3821 computed from the XL cap. The 330 marker bytes
overran it by 294.

Raising the ceiling is not an option: it is a red line certifying another
ADR's completion. plan-phase.md is reverted to byte-identical origin/next
and the pilot moves to execute-phase.md, which has 728 bytes of headroom
under its own ceiling and lands at 93147 with 3 marker pairs.

The vocabulary narrows to the atoms actually used: always, flag:--wave,
state:gap-closure-phase, state:has-prior-phases.

Recorded in the ADR: every branch the epic names lives in plan-phase.md,
which cannot be fragmentized until caps move from source to emitted bytes.
That is direct evidence for the epic's premise and may reorder phases 3-4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2930): backfill changeset PR number (#2972)

* fix(#2930): make the emission install tests portable on Windows

The windows-latest lane went red on three tests in the new install suite;
Linux was green. Both causes were in the test harness, not the module.

Root normalization: the opencode converter always embeds the install root
forward-slashed, but the tests stripped it with the native-separator string
from mkdtemp. On Windows that never matched, so the root leaked through
unstripped — and because the real and stub install roots have different
prefix lengths, that length difference landed directly in the byte-delta
assertion (344 observed vs 275 expected). Normalize both text and root to
one separator form before stripping.

@-ref resolution: the helper stripped only the @~/ and @$HOME/ forms, so a
Windows absolute ref (@C:/Users/...) fell through and was joined onto the
root, producing ...\@C:\Users\... Strip the @ first, then detect
absoluteness from the token's own shape (POSIX, drive-letter, or UNC) with
no platform branching, so every OS takes the same path.

Neither assertion was weakened; the exact-equality byte check is the point
of the test and still holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2930): document every REASON member and guard the doc/enum parity

Code review found the reference doc's 'Fails closed' list covering 10 of the
11 frozen REASON members — MALFORMED_ATTRIBUTES (parseAttrs rejects malformed
key="value" syntax) had no bullet, and it is distinct from
UNRECOGNIZED_ATTRIBUTE, which is valid syntax with an unknown key.

Two parallel surfaces sharing one constant with nothing asserting they agree is
the DEFECT.GENERATIVE-FIX class, so the same commit adds the parity assertion:
the test derives the enum side from the built module and the doc side by parsing
the reference page, keyed on the reason IDENTIFIER rather than prose so a
reworded bullet does not break it, and reports set differences in both
directions by name.

Proven non-vacuous: removing the MALFORMED_ATTRIBUTES bullet turns the suite
red naming that exact member; restoring it returns 44/44.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 12:12:20 -04:00
Tom Boucher
9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00
JusticeWay
7b204ad2ac enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages)

response_language is a free-form config value, but CHECKPOINT_FRAMES only
covered 9 languages — any other configured language silently fell back to
the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi,
Arabic, Vietnamese, and Indonesian frames plus their aliases, with a
regression test asserting each resolves instead of falling back.

Follow-up to #2402 (PR #2457).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: add changeset for #2527

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2530): point changeset fragment at PR #2557

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2530): address Unicode language-pack review

* fix: address checkpoint language review

* fix: count spacing combining marks in checkpoint width

* test: verify checkpoint aliases structurally

* fix: isolate RTL checkpoint frames

* fix: isolate RTL checkpoint frames correctly

* test(#2530): assert checkpoint aliases neither collide nor go unreachable

Review Minor #1. A duplicate alias key was invisible to the existing
catalog tests: the runtime object is well-formed after JS collapses the
literal, the self-alias assertion still holds, and the losing language
just stops resolving. tsc catches the byte-equal case (TS1117), but not
the two that survive compilation — an alias whose NFC-lowercase form
already belongs to another language, and an alias not in lookup form at
all, which resolveCheckpointFrame() can never produce.

The check reads the source literal rather than the object, since the
object no longer records what was written. Both assertions are
independently load-bearing: an NFD twin of an existing alias trips the
collision check, an uppercase alias trips the unreachability check.

Review Minor #2: changeset retyped Changed -> Added. Nine wholly new
supported response_language values are an addition under Keep a
Changelog, not a modification of existing behavior.

* test(#2530): check alias collisions on the catalog, not its source

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:01:50 -04:00
Tom Boucher
5a0a9f0972 fix(#2944): remove the catastrophic-backtracking regex from the ADR-1671 example (#2950)
* fix(#2944): remove the catastrophic-backtracking regex from the example

The non-shipping Option-E reference example carried its own copy of the
predicate-id regex, which nested a dot-containing character class inside a
dot-prefixed repeat. A run of N consecutive dots therefore had exponentially
many partitions. Measured on next before this change: 30 dots 54ms, 35 66ms,
40 807ms — so roughly 55-60 dots hangs for hours.

Not exploitable where it sits: the example is outside tsconfig.build.json,
outside the npm package files list, outside the installer and outside tests,
so no build step or CI job parses anything with it. Fixed because the entire
point of a reference example is that people copy it forward, and ADR-1671
presents this one as the pattern for the platform.

Ports the linear per-segment validation that #2928 gave the production module,
so the two copies agree: both parse the real CONTEXT.md to 415 predicates
across 20 classes with 0 duplicates. Doubled-dot ids are now rejected here
too, matching production, and the grammar comment records it.

Also refreshes the example's committed index, which #2928 made stale when it
removed the duplicate predicate from CONTEXT.md.

Closes #2944

* test(#2944): guard predicate-index sync and example/production parity

Two regression tests for the two defects in this PR.

Index sync: asserts the committed docs/CONTEXT-INDEX.json equals a fresh parse
of CONTEXT.md, naming any diverging predicate ids. The merge race that reddened
next was invisible to both PRs involved and only surfaced on the next PR to run
lint:ci; this puts the same check inside the suite, which runs on every PR, and
a mutation test proves the assertion is not vacuous.

Example/production parity: asserts both copies of the parser report the same
count, classes and duplicates for the real CONTEXT.md, and agree verdict-for-
verdict over a table of id shapes. The divergence WAS the bug — production went
linear-time while the example kept the backtracking regex, with nothing
asserting they agreed. Also pins the example rejecting a 60-dot id, with the
clean rejection as the binding assertion and wall-clock only as a smoke check.

Notes a real tension rather than hiding it: ADR-1671 says the example sits
outside tests/, and this imports it. The ADR's intent is that the example is
not compiled, packaged or installed — not that it may silently rot. A parity
guard does not ship it. The file states this so a reviewer can object.

* fix(#2944): address both isolated review passes

Two independent reviewers (correctness and security axes, neither the author).
Security found nothing — it measured linearity to 100k chars across dots,
hyphens, underscores and mixed classes, and showed prototype pollution is
structurally unreachable because the first-segment pattern forbids
lowercase and underscore-leading ids. The correctness pass found three
blockers, all real.

Blocker: the parity test violated ADR-1671 verbatim. The ADR lists FOUR
exclusions for the reference example, the fourth being the CI test suite, and
the test imported it from tests/ while its own justification comment cited only
three -- constructing a rationale around the exclusion it broke. Moved to
scripts/lint-example-parser-parity.cjs wired into lint:ci; a lint script is not
the test suite, so the exclusion stands. The test file keeps only the
docs/CONTEXT-INDEX.json sync check.

Blocker: the mutation test leaked its temp dir. Its callback took no `t`, so a
failing assertion skipped the bare cleanup call. Now registered via t.after(),
matching the convention adr-index-gate.test.cjs documents.

Blocker: the example's own committed index carries the identical merge-race
staleness this PR fixes for the production one, and nothing guarded it.
Deliberately NOT fixed by wiring the example's --check into CI: that artifact
bakes line numbers, so it re-drifts on any unrelated CONTEXT.md line shift --
exactly ADR-1671 open question 4 -- and would make CI routinely red. The new
lint asserts the line-INDEPENDENT facts instead: count, class map, duplicate
set, and every (id, value) pair. Proven non-vacuous both ways: mutating a value
fails and names the id, mutating only a line number passes.

Major: a real divergence the parity claim would have missed. Production rejects
values containing an embedded CR, LF, U+2028 or U+2029; the example did not, so
a value with an embedded lone CR was rejected by one copy and accepted by the
other. Ported, and now covered by the parity table.

Also, found while verifying rather than reported: malformed diagnostics covered
only empty values. A doubled-dot id, a space in an id, and a lowercase-leading
id were all dropped silently. That contradicts the module's own intent -- a
typo should be diagnosable, and a space in an id is a likely one -- and
predicates are contractually cited, so a silently vanished predicate is the
failure mode that matters. Each rejection class now carries a named reason in
both copies, while ordinary inline code still yields none.

Trues up counts my own change staled: the example README and ADR-1671's
prototype figures said 416 and 393/18 against a real 415/20/0.

Closes #2944

* chore(#2944): backfill changeset PR number 2950

---------

Co-authored-by: sim <sim@local>
2026-07-31 15:44:19 -04:00
Tom Boucher
07603df8f2 fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp

* fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp

The gsd-code-fixer agent hand-rolled its worktree at a hardcoded
`/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash
that landed OUTSIDE the project tree — outside the agent session's permission
allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's
MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path.

Place the worktree repo-relative under `.claude/worktrees/` (the same dir the
harness-managed executor worktrees use: gitignored via `.claude/`, inside the
session's permission scope), with a $$-PID + epoch suffix for concurrency
uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way
the cleanup tail already resolves it.

Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules.
The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test
asserts it). Failing-first regression added to the #2990 suite in
tests/agent-frontmatter.test.cjs.

* test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp

The #2686 regression test encoded the worktree location as a hardcoded
`/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that
breaks Windows/Git Bash (worktree outside the project tree → permission prompts;
mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to
require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`.
The #2686 isolation + cleanup assertions are unchanged.

* fix(#2647): word-boundary wt= parse + ack the fixer growth vs next

Two follow-ups to the #2647 GREEN run:
- parseWtAssignments matched `prior_wt=` (no word boundary), polluting the
  set and tripping the repo-relative + concurrency-unique assertions. Anchor
  on (?:^|\s)wt= so only the real worktree-path assignment is captured.
- emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update
  the emitted-drift-ack entry to attribute the #2647 worktree-path change
  (supersedes the prior #2825 attribution, whose growth is already in next).

* fix(#2647): address review — validate padded_phase at the sink + tighten test

Code-review + security-review both APPROVED with one actionable minor:
padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but
was only validated by the orchestrator (code-review-fix.md), not at the agent
sink. The agent prompt is a literal bash contract any caller can spawn, so add
a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell
metachars (defense-in-depth; not a present vuln — the only caller validates).

Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s)
(either-alone was too lax per review). Update the emitted-drift-ack reason to
cover the added validation growth.

* changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp

* changeset(#2647): backfill PR number 2942

* chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next

#2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR
emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json
generated from the OLD prose. lint:generated-sync fails on every PR that
rebases onto next after #2938 (the regen produces a 3-line diff bringing three
RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET
— in sync with the prose already on next). Mechanical regen via
`node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the
#2647 rebase. No behavioral change.

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 14:45:51 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
81eeb8a53a docs(#2926): refresh ADR-1671 with findings re-verified on next (#2936)
Re-measured the Option-E prototype's reported figures against CONTEXT.md on
next (2026-07-31) and recorded the delta rather than overwriting the June
numbers:

- index counts 393/18 (2026-06-24) -> 416/20 today; CONTEXT.md gained the
  PROBE (11) and PROHIB (10) classes
- of the 3 duplicate predicate IDs, only RULESET.WORKFLOW_MARKDOWN.FENCES
  remains; the two RULESET.GEMINI.* went with the Gemini runtime removal
- gen-context-index.cjs --check exits 1 on next, so Phase 0's "--check green
  in CI" criterion is unmet (invisible to CI: the example sits outside tests/)

Adds Open question 4 (index keyed on baked line numbers re-drifts on any
CONTEXT.md line shift, which matters once Phase 1 promotes --check to a CI
gate), and records the reviewer-proposed eval-gate question as resolved by
the PROBE.*/PROHIB.* predicate classes (ADR-550 D4/D7, ADR-1606).

Also names both surfaces of the Windsurf 12 KB throw in Decision 2, since it
is duplicated byte-identically in bin/install.js and
src/runtime-artifact-conversion.cts.

Docs-only. No code, no runtime-loaded text, no behavior change.

Co-authored-by: sim <sim@local>
2026-07-31 10:52:52 -04:00
clezcoding
9bd0dbf0dd docs(#2534): rewrite your-first-project tutorial for beginners (#2569)
* docs(#2534): rewrite your-first-project tutorial for beginners

Adds a loop mental-model primer (Mermaid), per-step "what just happened"
callouts, a prerequisites flow, a glossary and a troubleshooting table.
Same commands, same .planning artefacts, same to-do CLI example.

Closes #2534

* docs(#2534): make the tutorial runtime-agnostic (all IDEs)

Adds a "Pick your runtime" section (Cursor, Claude Code, OpenCode, Codex,
Gemini CLI, Copilot, Windsurf, Kilo, Cline, Qwen, Antigravity, ...) with the
installer flag and command syntax per runtime (/gsd-*, /gsd:* colon form, and
Cline rules). Keeps the same guaranteed worked example and .planning artefacts.

Closes #2534

* docs(#2534): address review - drop gsd-cursor aside + dead hero comment

- Remove the '(pair with the gsd-cursor EoS ...)' parenthetical from the Cursor row.
- Remove the commented-out reference to a non-existent hero asset.
(Gemini CLI references retained: --gemini is still live in bin/install.js on next.)

Closes #2534

* docs(#2534): fix review defects (keep multi-runtime)

- Replace dead Gemini CLI / --gemini with its live successor Antigravity
  (#1928); remove the invalid --gemini row/flag everywhere.
- Replace fabricated Step 1 output with realistic installer lines
  (71 skills/commands + destination suffix; exact lines vary by runtime).
- Fix 'Skip research' -> choose 'No' on the real Research prompt.
- behaviours -> behaviors (2x).

Multi-runtime 'Pick your runtime' section retained per author intent;
scope re-approval on #2534 still pending.

* docs(#2534): scope tutorial back to single-runtime (Claude Code)

Per trek-e's 2026-07-27 review, resolve the multi-runtime blockers by
returning to the approved scope:

- Remove the 'Pick your runtime' table + per-runtime notes; leave a one-
  line pointer to docs/how-to/install-on-your-runtime.md (which already
  documents all runtimes) rather than duplicate it (avoids the drift).
  This kills Blocker 1 (Antigravity is slash-hyphen, not colon) and
  Blocker 2 (Codex is $gsd-*) at the source.
- Step 1 uses --claude concretely; config-dir prose is Claude-local.
- Step 5: fix singular researcher (plan-phase spawns one gsd-phase-
  researcher), and make the research choice consistent with Step 3
  (choose 'Skip research'); drop the RESEARCH.md artifact line.
- Glossary/troubleshooting/prereqs/Step 2 de-multi-runtimed.

Returns the PR to #2534's approved 'docs-only, same commands' scope.

* docs(#2534): correct tutorial prerequisites and outputs

* docs(#2534): match tutorial research prompts to workflow

* docs(#2534): complete tutorial step guidance

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 09:48:54 -04:00