Commit Graph

5715 Commits

Author SHA1 Message Date
Michel Moreira
ae40529d31 chore(#4394): lint allowed-tools parity — Bash without Grep (#4431)
gen-plugin-skills.cjs --check already guarantees skills/*/SKILL.md matches
what commands/gsd/*.md generates, so the two trees cannot silently diverge
FROM EACH OTHER. Nothing guarded the shape #3085 actually found: a command
shipping Bash without Grep purely by omission, identical in both trees and
therefore invisible to a parity check that only compares them to each other.
That drift ran until 29 of 71 skills lacked a tool most of their siblings
declared, and a manual audit — not a gate — is what surfaced it.

Detection only: the lint never edits a command's allowed-tools.

Two failure classes, not one. Violations are the rule itself. Stale
exemptions are the other half: an entry whose command is gone, or which no
longer declares Bash without Grep, fails just as loudly. An exemption list
that can only grow becomes a list of things nobody re-examined, and a
pre-forgiven command silently absorbs the next omission.

That check earned its keep immediately. #4394 named eight exemptions from the
#3085 review — the six ns-* dispatchers, help, and surface — and seven of
them do not declare Bash at all, so the rule never reaches them. Listing them
would have pre-forgiven seven commands for a condition none of them has. Only
surface needs an entry.

The tests drive synthetic fixtures rather than the live corpus: asserting
"the real tree is clean" would say nothing about whether the rule can detect
anything, which is the exact failure mode this lint exists to close. The one
live-corpus arm asserts no stale exemptions — a property of this script's own
list — and deliberately does not pin a violation count, which would make it a
baseline every fix has to update.

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 09:44:38 -04:00
Tom Boucher
0a0905705a fix(#4256): resolve todos from the root via todosDir everywhere (#4479)
* test(#4256): pin todos as root-scoped under workstreams (RED)

* fix(#4256): resolve todos from the root via todosDir everywhere

* chore(#4256): changeset fragment (pr number to backfill)

* chore(#4256): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 08:24:14 -04:00
Tom Boucher
6ebe6372ce fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474)
* fix(#4243): anchor stateReplaceProgressPercent bold form to line start

The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag,
so a bold percent-ish label quoted MID-SENTENCE inside prose — an
Accumulated Context bullet mentioning **Progress:** — captured the
machine-segment rewrite and destroyed the rest of its line, silently, while
the real Progress line stayed stale (and the frontmatter moved on without
it, breaking the #4213 surfaces-agree contract). Every caller
(cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation)
feeds the whole document, so all three were exposed.

Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the
exact idiom #4453 applied to stateReplaceField's bold branch (same-line
confinement per #4010: the leading class is [ \t]*, deliberately not \s*,
which can consume the newlines before the label into the match; $ is
explicit-and-inert and documents end-of-line).

#2177's recorded requirements all stand: frontmatter is stripped before
matching, the suffix-preserving machine-segment swap is untouched, and
bold-beats-plain priority now governs line-start forms, so an earlier
free-text plain Progress: line still cannot capture the rewrite ahead of the
real bold status line. Per the maintainer ruling (2026-09-07), #2177's
incidental bold-anywhere matching was not load-bearing.

* test(#4243): scope the C4 region check with splitLines, not a bare \n split

lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's
content.split(/\n## /)[0] — a bare \n split on readFileSync content is
CRLF-fragile under Windows autocrlf. Same scoping via splitLines()
(src/text-lines.cts), which splits on \r?\n.

* chore(#4243): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 04:27:35 -04:00
Tom Boucher
c4b6dbd486 fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468)
* test(#4247): failing-first regressions for checklist-form update-plan-progress

* fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry

* fix(#4247): single local source for the phase-heading anchor grammar

* docs(#4247): note the missing_phase_details refusal in cli-tools reference

* docs(#4247): backfill pr number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 02:20:47 -04:00
Tom Boucher
03c770be47 docs(#4463): record executor self-repair of worktree base as out-of-scope (#4470)
Denies reintroducing sub-agent-side `git reset --hard` recovery for a
worktree base mismatch — the exact primitive #48 removed for safety.
Keeps the fork-base measurement finding for future reference.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 00:52:49 -04:00
Tom Boucher
33e393ba4c fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453)
* test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip

* fix(#4243): anchor stateReplaceField bold form to line start

The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold
label quoted mid-sentence inside prose — the issue's **Status:** inside an
Accumulated Context bullet — captured the rewrite and destroyed the rest of
its line, silently, whenever a whole-body caller fed the function every
section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan
writes). The plain branch was always line-anchored; only the bold branch
lagged.

Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's
same-line confinement idiom for the leading class (deliberately not the
issue's suggested ^\s* — it can consume the newlines before the label into
the match) and #4186's recognition-by-anchoring discipline. Frontmatter
half of the issue (unknown-key drops, invented milestone defaults) is
already fixed on next by #2202/#3216/#4129; pinned here with the issue's
requested regression fixtures.

* test(#4243): pin survival contract, not derived percent, in frontmatter rows

Bench RED run caught two assertion defects in the pin rows: the unknown
progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a
declared derived subkey - omitted under the #3573 no-roadmap withhold,
recomputed when measured (#4129) - so pinning its value over-pins derived
semantics. The rows now pin what the issue demands: unknown/custom keys
survive, stored counters are kept under the withhold, milestone identity is
never reset to invented defaults.

* chore(#4243): changeset for the anchored bold-field fix

* chore(#4243): backfill PR number in changeset
2026-09-07 00:03:15 -04:00
Tom Boucher
8c8eda46b0 fix(#4225): scope the sibling-worktree phase-number horizon to the active workstream (#4450)
* test(#4225): failing-first matrix for phase.add --ws workstream-scoped numbering

Nine rows driven through the real CLI: the issue's verbatim topology
(root roadmap @39 committed, workstream @2, sibling git worktree carrying
the root roadmap), same-workstream sibling boundary, empty-workstream
first phase, coincidental root-maximum, no---ws control (the #3849
global horizon, byte-for-byte), cross-workstream isolation, sibling
lacking the workstream (fail open), add-batch parity, and a
next-decimal control. Rows 1/2/3/6/7/8 are RED on next @38e4ce5f62
(numbering computed from the sibling ROOT roadmaps: 40 instead of 3).

* fix(#4225): scope the #3849 sibling-worktree widening horizon to the active workstream

collectSiblingWorktreePhaseNums scanned each sibling git worktree's ROOT
.planning/ (phases/ dirs + ROADMAP.md headers) unconditionally. Under
--ws (GSD_WORKSTREAM), every local number source flows through
planningDir(cwd) and lands in the workstream scope, but the widening
horizon still merged the siblings' ROOT-roadmap numbers into it — so
phase.add --ws in a workstream at Phase 2 inside a project whose root
roadmap sits at Phase 39 minted Phase 40 (directory 40-<slug>, and a
Depends on: Phase 39 that does not exist in the workstream's numbering
universe).

The horizon now resolves each sibling's planning dir through the SAME
canonical resolver, planningDir(wt, ws), with the env workstream read
once via planningDir's own discriminator: a workstream-scoped allocation
scans the sibling's copy of the SAME workstream (a number taken by that
workstream on another branch is still taken — the #3849 widening
survives, scoped), and never the sibling's root roadmap or another
workstream's. No workstream active: ws is null and the root-scope
horizon is byte-for-byte the #3849 behavior. A sibling lacking the
workstream directory contributes nothing (fail open, unchanged).

phase.add and phase.add-batch share the helper; both scopes of both
verbs are covered by the matrix in the previous commit. Output shape
and the publishStateContract boundary are untouched — only the number
changes.

* fix(#4225): rename siblingPlanning -> siblingPlanningDir (review nit)

* chore(#4225): changeset fragment (pr number to backfill)

* chore(#4225): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 22:24:07 -04:00
Tom Boucher
f15887ebb1 feat(#4446): ban ad hoc timeout literals in tests, ship with full legacy allowlist (#4449)
Nothing enforced CONTRIBUTING.md's own stated preference ("A non-literal value
is trusted — that is the shape you should be writing") for timeout values in
tests. local/no-unbounded-spawn requires SOME bound but allows a bare literal;
no-magic-sleep-in-tests and no-elapsed-assertion cover different anti-patterns
entirely. This is exactly how PR #4428's Windows CI incident happened: two
independently-guessed 15000ms literals (one in production's check-latest-
version.cjs, one in this suite's own worker test) collided exactly and raced
two SIGKILLs against each other.

New rule local/no-adhoc-timeout-literal (eslint-rules/no-adhoc-timeout-
literal.cjs) flags a resolvable numeric timeout/timeoutMs literal; an
Identifier or MemberExpression value is trusted. No marker-comment escape —
the fix is always to extract a named constant, which is trivial.

Ships with a full legacy allowlist (126 files, 352 violations, generated
by running the rule with an empty allowlist against tests/) so it can go
live at error severity without breaking CI, mirroring how no-unbounded-spawn
itself was rolled out. Migration is tracked separately in #4445 (this PR
does not close it — only #4446, introducing the gate itself).

Documents the policy and the compliant shapes in TESTING-STANDARDS.md and
TEST-EXAMPLES.md.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 21:23:39 -04:00
Tom Boucher
ef30e59860 fix(#4448): stop io.test.cjs's in-process runMain calls from corrupting node:test's own fd-1 IPC (#4452)
tests/io.test.cjs's "review fix: pending-outcome cell lifetime" describe
block deliberately drives runMain()/output() in-process (needed to observe
a cross-invocation state leak) instead of via a subprocess. output() ends
with a raw synchronous fs.writeSync(1, ...) to the real stdout fd, and
Node's --test-isolation=process (default since Node 22) uses that same fd
for the file's own parent-child reporter protocol. The two writes racing
produced an intermittent "Unable to deserialize cloned data" that killed
the whole file — observed twice on next's macOS lane, most recently on the
commit that merged PR #4428 (unrelated to that PR's content; io.test.cjs
isn't part of its diff).

Empirically validated locally (gsd-test can't reach macOS): built a repro
loop running N parallel copies of `node --test tests/io.test.cjs` to
recreate CI-like contention. Baseline: ~13-17% of runs hit the corruption
(12/90, 15/90 across two samples). A first fix attempt wrapped the writes
in captureFdAsync (an await-aware twin of the existing captureFdSync,
added because runMain() defers main() through a microtask chain, so a
synchronous wrap restores before the real write fires) — but captureFdAsync
always forwards to the real fs.writeSync by design (matching captureFdSync's
"never swallow" contract, tests/helpers.cjs, #4306). Re-ran the same loop
against that fix: 15/90, statistically unchanged. Forwarding the write
doesn't stop it from reaching the fd node:test's own IPC also uses.

Replaced it with suppressFdAsync: a narrow, deliberate exception to the
never-swallow contract for a window the caller has verified is fully
controlled (a single runMain() call plus its promise-chain settling,
where nothing else can legitimately need that fd). It records the bytes
for the test's own assertions but never lets them reach the real fd.
Re-ran the loop: 0/300 across three samples (90+120+90), including one
round at 8-way parallelism. Also caught and fixed a real bug surfaced by
the same loop: the new regression assertion checked for compact-JSON
`"error":"x"` but output() pretty-prints, so it failed 100% of runs
deterministically until fixed to parse and check the structured value
instead (io.test.cjs, matching this repo's "assert on structured output,
not raw text" convention).

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 21:06:48 -04:00
Brenden Smerbeck
e54d3aa159 enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key

- Add compact_content: false to the nested workflow object in
  gsd-core/bin/shared/config-defaults.manifest.json
- Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts
  so an absent key resolves to false via config-get --raw
- validKeys entry in config-schema.manifest.json already present

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(#4401): behavioral and boundary tests for workflow.compact_content

- 19 behavioral tests covering config-set/config-get round trip, invalid-shape
  rejection (banana, 42, empty string), the corrected null-unset semantics
  (#2046), absent-key resolution against config-defaults.manifest.json,
  config-new-project wiring, and doc-row shape assertions
- Drops the install-tree fixture-parity block (and its docstring item) that
  asserted gsd-core/references/compact-content-gate.md and
  gsd-core/workflows/compact/map-codebase.md fixture entries — those paths
  belong to #4402 and do not exist on this filtered branch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(#4401): document workflow.compact_content in both config references

- One 4-cell row in docs/CONFIGURATION.md (workflow.* run)
- One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md
- Both cross-reference ADR-4139

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): add changeset

- Added-type fragment, pr: 4401 (issue number; backfill to the real PR number
  is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR-
  FIELD-DRIFT)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): backfill changeset pr field to #4441

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS

SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false
instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling
entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving
a single-source-of-truth drift risk: a future manifest-only edit to the
default could silently diverge from this literal, only caught later by
the D-03 test if it ever happened to manifest.

Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and
derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority
sibling pattern. Found during maintainer review (review-open-prs) of
this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP

The previous commit added compact_content to CONFIG_DEFAULTS in
src/config-loader.cts but missed the matching entry in
tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat
CONFIG_DEFAULTS keys to their namespaced doc form before checking
gsd-core/references/planning-config.md for a match. Without it, the
test looked for a bare `compact_content` doc reference instead of the
actual `workflow.compact_content` row, and failed:
"CONFIG_DEFAULTS keys missing from planning-config.md: compact_content".

Found by actually running gsd-test against the branch rather than
trusting the plausible-looking fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4401): register compact-content-4139 test in the docs-guard lane

tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md
directly (fs.readFileSync) to assert the workflow.compact_content doc row's
shape, which makes it a doc-reading test file under the #3753 docs-guard
lane. It was never added to scripts/docs-guard-registry.cjs's
DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so
tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed:
"compact-content-4139.test.cjs reads a docs/ path but is not registered in
the docs-guard lane and carries no docs-guard-exempt marker".

Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed
path it reads; gsd-core/references/planning-config.md is outside this
registry's docs/ scope, matching the sibling config-field-docs.test.cjs
entry's existing convention).

Found by actually running gsd-test against the branch — this gap predates
the maintainer's config-loader.cts fix and was already present in the
original PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
2026-09-06 19:52:59 -04:00
Tom Boucher
2cf119f57e fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends

* test(#4217): pin the completion-reconciliation contract

* chore(#4217): regen derived inventory and install-tree fixtures

* test(#4217): follow the #4003 anchoring pins into the reconciliation fragment

Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217)

* chore(#4217): add changeset fragment

* chore(#4217): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-06 18:57:38 -04:00
Michel Moreira
19b66c3ec8 fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working

An executor with recent RED/GREEN/REFACTOR commits and passing verification had
not yet written its SUMMARY because it was finishing closeout. The parent saw no
local OS test/build process, inferred an "idle tail", and sent "Finalize
immediately" into a working child; in CLI runs the same inference interrupted an
executor before GREEN, leaving a RED commit and an uncommitted edit.

The stall block said only "if no completion signal, no SUMMARY.md, and no
expected-branch commits appear for N minutes" — it never said what to do when
commits DO exist and only the SUMMARY is outstanding, never defined the
threshold as a period without progress rather than a total runtime, and never
ruled out a process listing as an idleness signal. Four rules close that:

- the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last
  sign of progress, not from dispatch — a long verification tail is not a stall;
- commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with
  steering, interrupting and re-dispatching each named and forbidden;
- urgency/finalization messages ("Finalize immediately" and family) are
  forbidden outright — they arrive mid-verification and truncate a correct run.
  The existing user-facing pause is the only sanctioned stop, and `kill and
  retry` is a clean restart, not a nudge;
- the absence of a local OS test/build process is NOT idleness: a native
  subagent runs in the runtime's own session, and an executor between two tool
  calls shows no process at all. Progress is judged only by the signals this
  workflow names.

Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on
next.

* fix(#4218): extract the progress policy to a step fragment

CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600
ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's
stated remedy, and this workflow already carries policy detail that way.

execute-phase/steps/executor-progress-policy.md owns the policy. The
worktree-recovery arm moved with it — `kill and switch to inline execution`
qualifies the stop this policy governs, so it belongs beside the rule about when
stopping is sanctioned at all, not stranded in the host. The #3212 recovery
OPTIONS stay in the host, where tests/config.test.cjs pins them.

The host keeps what must be read before the orchestrator acts: the verdict, the
threshold definition, and a pointer that fires before any message is sent to the
child. execute-phase.md is now 93475 bytes — 48 SMALLER than next.

* chore: add changeset for #4218

* chore(#4218): regenerate the inventory manifest for the new step fragment

docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind
INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear
there or gen-inventory-manifest --check reds the lint-tests lane.

* chore(#4218): restore the issue ref on the allow-test-rule marker

ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the
policy into the fragment dropped it.

* chore(#4218): regenerate the install-tree fixtures for the new step fragment

The fragment ships with the workflow, so every runtime's golden install tree
gains one path — gen:install-tree is the generator that owns those fixtures.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:51:34 -04:00
Tom Boucher
1c0acb2359 feat(#4422): block merging into next/main while the base branch's Tests run is red (#4428)
* feat(#4422): block merging into next/main while the base branch's Tests run is red

Adds a next-health job to test.yml that checks the base branch's own last
push-triggered Tests run via the GitHub API and fails the existing "Required
tests" required check when it's red, with a maintainer-applied "fix-next"
label as the explicit escape hatch for the fix-forward PR itself. No
branch-protection config change needed — it rides the already-required
check. The job is deliberately not gated behind preflight, same reasoning
as the changes job: a compute-free API read has nothing to save by waiting.

Documents the fix-next label in CONTRIBUTING.md and adds a property test
locking the CLEAN/RED/INDETERMINATE classification's iff-relationship.

This closes the second half of the 2026-09-06 RCA: three unrelated PRs
merged on top of an already-broken next before anyone noticed it was red.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: close two zero-margin CI timing gaps found while verifying #4422

Discovered while watching this branch's own CI, root-caused via /diagnose
rather than dismissed as Windows flakiness:

1. tests/gsd-check-update-worker-atomic-cache.test.cjs's outer timeout
   (15000ms) exactly matched the inner npm-view timeout the worker wraps
   (NPM_VIEW_TIMEOUT_MS, gsd-core/bin/check-latest-version.cjs). A slow
   registry response raced two SIGKILLs at the same instant, killing the
   worker before it could catch its own timeout and degrade gracefully.
   Windows's shell-wrapped npm subprocess made the race lose more often
   there, but the zero margin was platform-agnostic. Fixed by giving the
   test real headroom (+10s) beyond the named constant it wraps, plus an
   invariant test so the two values can't silently collide again.

2. scripts/run-tests.cjs's per-chunk weight budget (MAX_FILES_PER_CHUNK)
   let a Windows full-matrix chunk that was well under budget by the
   Linux/macOS-calibrated weight table (~32/60 units) still exceed the
   600s wall-clock backstop — codex-config.test.cjs's genuinely-measured
   weight (17.87) doesn't transfer 1:1 to Windows's slower install/
   subprocess overhead. Windows now gets its own lower cap (40 vs 60).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 17:29:54 -04:00
Tom Boucher
38e4ce5f62 fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin

Three defects from #4186:

1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the
   free-prose body Status field, so prose merely mentioning a status word
   was silently rewritten to a credible wrong token (a .planning/ path in
   Italian prose -> status: planning; verifica -> verifying; completezza ->
   completed). Recognition is now an ANCHORED whole-field match against a
   declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS,
   state-document.cts) — case/whitespace-tolerant, branch-order artifacts
   preserved (Planning complete -> planning; Phase complete — ready for
   verification -> verifying). The recorded lenient fallback (#3873 row 26)
   stands: unrecognized prose passes through verbatim. Read-side consumers
   (W011, statusline) ride the same function.

2. The progress recount skew (stray *-SUMMARY.md inflating
   completed_plans) is already dead on next via #1988/PR #2016
   (countMatchedSummaries pairs summaries to plans) — verified live and
   pinned with regression rows composed against the #4129/#4359 ratchet.

3. state record-session with no args executed and wrote STATE.md; it now
   errors like state update (stopped-at or resume-file required), handler-
   side so SDK callers are covered too. Four tests pinning the bare-call
   write are updated to the new contract.

* fix(#4186): update status pins to the anchored vocabulary contract

Bench round 1 follow-ups:

- Legacy bare 'Milestone complete' kept as reader-side vocabulary
  (ADR-2207 removed the writers, not recognition of legacy files).
- state.test pins updated: 'Paused at Plan 3' and round-trip
  'Executing Plan 5' were pins of the substring guessing itself —
  the round-trip now uses the real handler form 'Executing Phase 5'.
- record-session no-op/no-fields tests repurposed to the usage-error
  contract (CLI + SDK-level ExitError), byte-unchanged assertions kept.
- statusline tests repinned: vocabulary values collapse to keywords;
  narratives render the documented first-word fallback instead of a
  guessed token. Hook doc comment updated to match.
- docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md.
- docs/CLI-TOOLS.md: record-session signature notes the required flag.

* fix(#4186): repair a dangling sentence in the schema docstring

* test(#4186): bound the completed_plans scan regex (#2128 class)

* chore(#4186): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-06 17:08:24 -04:00
Dennis Alexis Valin Dittrich
47f83beb62 fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures (#4337)
* fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures

The runtime launcher's PATH-fallback arm probes `command -v gsd_run`
(renamed from gsd-tools in #3146), but every PATH-isolation fixture in
runtime-launcher-parity.test.cjs filtered on the pre-rename name. A real
installed gsd_run reachable on PATH survived the filter and got invoked
in place of the fixture's runtime-home stub, so negative tests passed
without proving PATH was actually empty and positive home-fallback tests
failed with "Unknown command" errors from the unrelated real CLI.

Adds (B1), a regression test that plants a sentinel gsd_run on PATH and
asserts the resolver still falls through to the HERMES_HOME stub instead
of invoking it.

* docs(#4205): fix stale gsd-tools references in PATH-probe doc comments

Addresses agy adversarial review nits on PR #4205: several doc comments
and JSDoc blocks still described the launcher's PATH-fallback probe as
`gsd-tools` after the filter fix. Updates them to `gsd_run` to match the
actual `command -v gsd_run` probe and the corrected filters. No test
logic changes.

* test(#4205): tighten (B1) assertions and dedupe rationale comments

Addresses opus critical-code-reviewer/ponytail findings on PR #27:
- (B1): split the collapsed && assertion into two, matching neighbor
  (B)'s style, for clearer failure diagnostics.
- (B1): drop the dead `if (nodeBinDir)` guard — cleanup() already
  no-ops on a non-string argument (tests/helpers.cjs:452).
- (B1): drop the dead backslash-path normalization — the test is
  win32-skipped, so stdout paths are always POSIX.
- Six near-identical "#4205: probe target is gsd_run, not gsd-tools"
  comments collapsed to pointers at the one canonical explanation in
  buildIsolatedPath().

No behavior change; 31/31 tests still pass.

* fix(#4205): scrub ambient config-dir env vars leaking into launcher fixtures

Same class of bug as the PATH leak this issue reports, different vector:
the resolver's runtime-home elif chain checks CLAUDE_CONFIG_DIR before
HERMES_HOME, CURSOR_CONFIG_DIR, CODEX_HOME, etc., but the fixtures
targeting those later arms never cleared the earlier ones from the
spread process.env. An ambient CLAUDE_CONFIG_DIR pointing at a real
install silently wins over the fixture's intended stub, exactly like
the reported gsd_run PATH leak. Confirmed with a real leaked install:
red on tests (D)/(H)/bug-211 (C)/(D)/(B1)/(B)/(C) without the fix,
green with it.

Also removes bug-891's (B) test, now a strict subset of (B1): once the
sentinel is filtered by buildIsolatedPath(), both tests exercise the
identical HERMES_HOME resolution with the identical script and env —
(B1) already asserts everything (B) did, plus the sentinel-not-invoked
check. Updated the block's header docblock to match.

30/30 tests pass (31 minus the removed duplicate).

* fix(#4205): scrub CODEX_HOME/XDG_CONFIG_HOME, restore (B) on Windows

Adversarial review of this PR found three more instances of the exact
leak class the PR exists to close.

(H) asserts the $HOME/.codex fallback but never cleared an ambient
CODEX_HOME, which overrides that default outright. Proven load-bearing:
with a fake install planted at CODEX_HOME the test fails without this
scrub and the leaked install's own output appears in stdout.

The two "every arm must miss" hard-error fixtures cleared all 16
config-dir vars but not XDG_CONFIG_HOME, which the resolver's opencode
and kilo arms fall back through as
${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode} — a
host XDG_CONFIG_HOME leaks past a fake HOME.

Restore bug-891's (B), deleted here as a subset of (B1). It is not one
on Windows: (B) ran cross-platform, (B1) is POSIX-only because it plants
an executable sh sentinel, so the deletion left the HERMES_HOME arm with
no Windows coverage. Restored with the CLAUDE_CONFIG_DIR scrub its
sibling fixtures already carry.

* docs(#4205): correct makeIsolatedPath's docblock and the bug-891 header

The doc-fix commit earlier in this branch rewrote makeIsolatedPath's
docblock from "strips gsd-tools" to "strips gsd_run". Both are false:
the function filters nothing, returns process.env.PATH whole, and the
"noToolsBin dir that shadows gsd_run with a sentinel" it describes does
not exist — every caller passes an empty directory. Its isolation comes
from resolution order, since the RUNTIME_DIR/.claude arm fires before
the PATH arm. Say that instead, and point anyone who needs the PATH arm
itself to miss at buildIsolatedPath(), which does filter.

Add (B) to the bug-891 asserts list; restoring it in 48cf0b917 left the
header describing a test set the file no longer has. Mark (B) and (B1)
cross-platform and POSIX-only respectively, which is why both exist.

Drop one more stale "remove gsd-tools" comment the doc pass missed.

* fix(#4205): derive the fixture env scrub, stop writing to CLAUDE_ENV_FILE

Two findings from review, both measured.

CLAUDE_ENV_FILE was never scrubbed. The snippet's tail appends
`export PATH='<dir>'` to it whenever it is set, so running this suite on
a host that exports it wrote 11 lines into the developer's real env
file, each naming a /tmp fixture directory the test had already deleted
— they accumulate per run and prepend dead entries to the PATH of every
later shell. The reported bug was fixtures READING developer state; this
was them writing to it. Now 0 lines.

The 17-key scrub list was hand-written, which tests/helpers.cjs already
warns against: "#2665: this list is DERIVED, not hand-maintained. A
hand-written list is exactly what reopened this bug twice". It was right
— the hand list here missed CODEX_HOME and XDG_CONFIG_HOME until review
caught them, and a 17th runtime home would have left it silently stale.
Replace both copies with TEST_ENV_BASE, derived from the registry the
resolver itself reads, applied at runBashFile/runResolver so every
fixture that sources the snippet is covered rather than the two that
remembered to ask. GEMINI_CONFIG_DIR is added explicitly: the runtime is
retired (#1928) so the registry no longer carries it, but the snippet
still probes its arm.

Verified by pointing all 16 config-dir vars plus XDG_CONFIG_HOME at a
real install tree: 30/30 pass.

Fold (B1) into (B). Reverting the filter under a clean PATH left the old
(B) green — it only caught the bug on an already-leaking machine — while
(B1) caught it anywhere but was skipped on Windows. One test now does
both: it plants the PATH sentinel on POSIX and still exercises the
HERMES_HOME arm on Windows. Mutation-checked both ways on a PATH with no
real gsd_run. Assert the hermes dir itself rather than "gsd-core/bin/",
which every resolver arm ends in and so cannot tell them apart.

* fix(#4205): scrub BASH_ENV and reject an empty RUNTIME_DIR in runResolver

BASH_ENV defeated the whole scrub. Non-interactive bash sources it
before the script runs, which is after the env: object is applied, so a
single inherited var re-injects any of the others. Measured: a BASH_ENV
exporting CODEX_HOME turned (H) red; blanked, 30/30.

runResolver passed RUNTIME_DIR: runtimeDir || '', and '' is
indistinguishable from unset to ${RUNTIME_DIR:-$(git rev-parse
--show-toplevel)} — an empty value falls back to the real repo root and
resolves its real install, the leak this issue is about. Both callers
already pass one, so require it rather than paper over it.

* test(#4344): plant the leaked gsd_run sentinel on Windows too

(B) planted its gsd_run sentinel only on POSIX, so the Windows shards
proved nothing about buildIsolatedPath()'s PATH filter — the exact gap
#4344 recorded. npm's global installs write an extensionless Bourne shim
beside gsd_run.cmd, and fs.constants.X_OK behaves like F_OK on Windows,
so the existing probe already sees the leak there; only the fixture was
POSIX-gated.

Plant the sentinel on every platform and assert that buildIsolatedPath()
strips its directory from the returned PATH. That assertion is red on
both platforms when the filter probes the wrong name, and unlike the
stdout assertions it does not depend on the MSYS mount's exec
heuristics. Restore process.env.PATH before the child spawns rather than
in a t.after hook: on Windows process.env spreads as 'Path', so a
still-live leak would compete with snippetEnv()'s 'PATH' override for
the casing the child receives.

Refs #4205

* fix(#4205): stop the host PATH leaking past snippetEnv on Windows

Adversarial review (agy, gemini-3.8-flash-high) found that the fixtures'
PATH isolation is defeatable on Windows regardless of which name the
filter probes. Windows environment variables are case-insensitive but a
spread of process.env is not: the host PATH enumerates as 'Path', so
'{ ...process.env, PATH: isolated }' yields both keys, and libuv's
make_program_env sorts the child's environment block case-insensitively
without ever dropping duplicates. The child could therefore resolve the
host PATH. snippetEnv() now drops every other casing whenever a caller
supplies its own PATH.

Also from that review:

- buildIsolatedPath() takes the PATH to filter as a parameter, so (B0)
  and (B) no longer mutate process.env.PATH and no longer need
  try/finally restores.
- The two loud-guard fixtures asserted 'not found' OR 'ERROR', which
  bash's own 'node: command not found' satisfies; they now assert the
  launcher's 'ERROR: gsd-tools.cjs not found'.
- Corrected a comment counting three scrub keys as two, and two comments
  crediting a removed env argument for clearing ambient config dirs
  rather than snippetEnv()'s derived TEST_ENV_BASE.

Refs #4344

* fix(#4205): make the fixtures' node shim work on Windows

bug-211 (C) located node with `which node` through the process seam.
`which` is not a Windows binary; the fixture only survived CI because
Git Bash ships one. process.execPath is the same answer without the
spawn, and (H) already used it.

Both fixtures then built their node shim with fs.symlinkSync, which
raises EPERM on Windows without developer mode or elevation — the same
reason buildIsolatedPath() skips its own symlink step there. The shared
linkNodeShim() helper hard-links instead on that platform (no privilege
required) and falls back to a copy across volumes.

Found by adversarial review (agy, gemini-3.8-flash-high).

Refs #4344

* fix(#4205): make the launcher PATH isolation extension-aware and Windows-safe

trek-e's review asks for a Windows-safe node fallback, an extension-aware
filter, and a Windows regression test, in that order: broadening the
filter first can strip the directory node itself lives in.

buildIsolatedPath() now always prepends a directory holding node, on
every platform, via the linkExecutable() helper (hard link on Windows,
where symlinks need elevation). nodeBinDir is no longer nullable and the
win32 early return is gone, so the fallback exists before the filter
widens. (B0), the co-location invariant, therefore runs on Windows
instead of being skipped on the one platform that had no fallback.

The filter probes every name the launcher's `command -v gsd_run` arm can
resolve. msys bash appends an executable extension during PATH lookup,
so a directory holding only gsd_run.exe is reachable on Windows although
gsd_run is absent. PATHEXT is folded in as well; it over-matches, which
costs nothing now that node is always supplied separately. (B0) asserts
every name in that set is filtered, and (B) plants a gsd_run.exe sentinel
on Windows beside the extensionless one npm installs.

The predicate lived in five hand-maintained copies — the drift that
caused #4205 in the first place, and four of the copies pointed readers
at a buildIsolatedPath() that was block-scoped out of their reach.
buildIsolatedPath() moves to module scope and the four inline copies call
it. The three shadow runBashFile() declarations this PR had to edit
identically go with them.

Red-proved both ways: probing 'gsd-tools' again turns (B0) and (B) red;
making the node prepend conditional turns (B0)(ii) red.

Refs #4344

* fix(#4205): drop empty PATH elements from the isolated PATH

A POSIX shell reads an empty PATH element as the current directory, so
an isolated PATH carrying one still lets the launcher's `command -v
gsd_run` arm resolve a gsd_run from the fixture's own working directory
— the leak class this file exists to close.

Two ways one appeared. An ambient PATH containing `::` survived the
filter, because `path.join('', 'gsd_run')` probes the working directory
rather than a directory entry, so hasGsdRun could not see what it was
admitting. And a PATH whose every entry was filtered joined to an empty
string, leaving the returned value ending in a delimiter, which means
the same thing.

Empty entries are now dropped alongside the gsd_run-bearing ones, and
the surviving directories are joined as a list, so a fully-filtered PATH
yields the node shim dir alone. (B0) asserts both cases.

Found by CodeRabbit on the fork rehearsal PR.

Refs #4344

* fix(#4205): keep only absolute PATH dirs, and assert the sentinel by basename

Adversarial review (agy, gemini-3.8-flash-high) on the previous head.

An empty PATH element was dropped, but `.` and any other relative entry
say the same thing explicitly and survived. hasGsdRun() cannot see what
it would admit either: `path.join('.', 'gsd_run')` probes the runner's
working directory, not the child's, so isolation also varied by where
the suite was started from. Only absolute directories survive now, which
can only tighten the isolation. (B0) asserts it over an empty element, a
`.`, a relative entry, and a fully-filtered PATH — the previous
empty-string assertion passed on the `.` case.

(B)'s GSD_TOOLS assertion compared an absolute os.tmpdir() path against
launcher output, which the file already documents as a mismatch on
Windows: git-bash prints /c/Users/... where Node gives C:\Users\....
It never matched there, so it asserted nothing on the platform it was
added for. It matches the mkdtemp basename now, which both path forms
share. (B) also plants ONLY gsd_run.exe on Windows: with an extensionless
sibling present, an extension-blind filter would strip the directory for
the wrong reason and pass.

snippetEnv() deduped case-variant keys for PATH alone. On Windows every
scrubbed key has the same exposure — an ambient `bash_env` reaches the
child beside the blanked `BASH_ENV`, and BASH_ENV re-injects the rest.
Every key the function sets now wins over other casings of itself; keys
it does not set are untouched, so a caller passing no PATH override still
gets the host PATH.

Refs #4344

* test(#4205): plant probe fixtures as files, not interpreter links

Review follow-ups on 776e9371f.

(B0)'s co-location fixture and its GSD_RUN_NAMES sweep only ever probe
the planted names with accessSync; nothing executes them. They used
linkExecutable, so on Windows the sweep hard-linked node.exe once per
PATHEXT entry — a dozen on a stock runner, and a full copy each when
os.tmpdir() and process.execPath sit on different volumes. plantExecutable
writes a zero-byte 0o755 file instead. linkExecutable keeps the two
callers that need a real executable: the node buildIsolatedPath prepends,
and the Windows sentinel.

The '/usr/bin:/bin' fallback formatted POSIX paths with the platform
delimiter, yielding '/usr/bin;/bin' on Windows, which path.isAbsolute
accepts and no Windows shell would ever produce. It was also unreachable:
basePath defaults to process.env.PATH. An unset PATH now yields the node
shim dir alone and fails loudly at spawn rather than being papered over.

(B)'s header said the sentinel is planted in both forms; the code plants
one per platform, and planting both on Windows is what the branch below
it exists to avoid. Also names which assertion carries the Windows
guarantee, since SENTINEL_INVOKED cannot fire there.

The shared helper's temp dirs were prefixed gsd-891-, attributing every
fixture's leftovers to one of the four bugs it now serves.

Refs #4344

* fix(#4205): model bash's PATH lookup, not cmd.exe's, and give (B0) its own oracle

Ponytail review on aec6012fc.

GSD_RUN_NAMES expanded PATHEXT, which describes cmd.exe rather than the
shell the launcher's `command -v gsd_run` arm runs under. The Cygwin/msys
rule is that .exe may be omitted from a command while '.bat and .com ...
you cannot omit the extension', so gsd_run.exe is reachable for a bare
gsd_run and gsd_run.cmd/.ps1 are not, whatever PATHEXT lists. The comment
claimed the resulting over-match was free. It was not:
buildIsolatedPath() restores node to the isolated PATH, but nothing
restores bash, which the fixtures spawn by name — so every extra dropped
directory was another chance to remove the one bash lives in and fail
with ENOENT instead of an assertion. Narrowed to gsd_run and
gsd_run.exe. (Aside, the PATHEXT default does not even contain .PS1.)

(B0)'s name-sweep took its list from GSD_RUN_NAMES, so it swept the
constant under test with itself and could only catch that constant being
deleted, never being wrong. Its relative-element case re-ran the
implementation's own filter predicate over that filter's output, which is
true for any predicate. Both now assert against written-out expectations:
the reachable names per platform, and the exact directories that must
survive each case.

Also: the test name covered two of its four assertions, and the case
table's prose counted three of its four entries.

Red-proved three ways: dropping the gsd_run filter, dropping the
absoluteness filter, and claiming a name the filter does not cover each
turn (B0) red.

Refs #4344

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:07:25 -04:00
Tom Boucher
f09e7ed08c fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375)
* test(#4137): keg-only Homebrew Cellar path falls back to raw execPath

Regression tests for the Homebrew branch of normalizeNodePath: the rewrite
to <prefix>/bin/node must be existsSync-guarded like the mise/volta
branches, falling through to the raw execPath when the keg-only formula
was never linked into <prefix>/bin. Also makes the existing #3181/#2185
Cellar assertions hermetic by injecting existsSync stubs (granting
existence to exactly the one candidate each asserts) so they no longer
depend on the runner machine's real /usr/local/bin/node.

* fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath

The Homebrew branch of normalizeNodePath returned <prefix>/bin/node
unconditionally — the only one of five runtime branches that never probed
its rewrite candidate. On a keg-only or versioned Homebrew install
(node@24 never brew-linked) that path does not exist, so every managed
hook command baked by resolveNodeRunner/buildBakedNodeToken/
buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh:
<prefix>/bin/node: No such file or directory.

Guard the rewrite with the already-injected existsSync exactly like the
mise and volta branches: when <prefix>/bin/node exists (linked formula)
the rewrite is byte-identical to today; when it does not, fall through to
the raw execPath — a working keg path instead of an immediately broken
one. Also drops two now-unused constants from the regression tests.

* test(#4137): make the #977 non-fnm Cellar assertions hermetic too

The Bug #977 folded block's two 'still maps to stable symlink' assertions
called normalizeNodePath without an existsSync stub, silently depending
on the runner machine's real /usr/local/bin/node (present on the Linux
bench image, absent for /opt/homebrew). With the #4137 guard these become
environment-dependent; grant each exactly the one candidate it asserts.

* chore(#4137): add changeset fragment

* chore(#4137): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-06 16:04:21 -04:00
Tom Boucher
2920bbc022 fix(#4421): rescind #494's macOS full-matrix skip on changed test files (#4427) 2026-09-06 15:42:23 -04:00
Michel Moreira
54085516c1 fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371)
* fix(#4211): materialize Kimi's agent tree recursively during surface apply

kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and
install copies that tree recursively (_copyStaged). Surface apply fell through
to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`:
the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was
written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name
that already carries it. A surface change could therefore corrupt Kimi's
installed artifacts while still reporting success.

Three divergences from the install path, all in src/surface.cts:

- _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped
  to exactly what install's _removeGsdEntries owns for this kind (the two root
  files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned
  and preserved.
- applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for
  an unmodified full profile, as it already does for the agents kind and as
  createRuntimeArtifactInstallPlan does for every kind — without it Kimi's
  generated subagents lost their path-prefix rewrites and attribution trailer,
  and an unmodified full profile staged only the skill-referenced subset.
- applySurface runs rewriteStagedSkillBodies for kimi-agents, which the
  install plan routes through it alongside skills.

* chore: add changeset for #4211

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 15:11:47 -04:00
Tom Boucher
acb3cc974b fix(#4197): dedup the update-context fast path against the selected global dir (#4413)
* fix(#4197): dedup the update-context fast path against the selected global candidate

The preferredConfigDir fast path derived scope from a cwd-relative match
alone, so a global install reported LOCAL whenever the shell sat in
$HOME — and run_update then drove the installer through its --local arm
(settings.local.json + the #338 relocation) against a global install.

Extract resolveGlobalCandidate (env candidates first, then $HOME-relative,
first hasInstall hit wins) and use it in BOTH paths: the fast path now
answers LOCAL only for a cwd-relative match that is not the selected
global dir, which is the same dedup the cascade applies at its isLocal
check. A preferred dir that is also the env-directed global now answers
GLOBAL on both paths (the cascade's answer), pinned by a parity test.

The discriminator is the selected global candidate, not the $HOME
pathname: with CLAUDE_CONFIG_DIR directing the global elsewhere,
$HOME/.claude probed from cwd === $HOME is a genuine local install, and
a pathname check would re-break parity (regression-pinned).

* chore(#4197): add changeset

* chore(#4197): backfill PR number in changeset

---------

Co-authored-by: agent-4197 <agent-4197@gsd.local>
2026-09-06 14:09:32 -04:00
Tom Boucher
ed133cc116 enhance(#3085): add Grep to allowed-tools for 21 skills (#4397)
* enhance(#3085): add Grep to allowed-tools for 21 skills

29 of 71 skills omit Grep from allowed-tools, forcing Bash grep for
structured search instead of the dedicated tool. Adds Grep to the
21-skill subset confirmed safe in prior review (excludes the 6
gsd-ns-* dispatchers, gsd-help, and gsd-surface, which have no
plausible structured-search need).

Hand-edits commands/gsd/*.md only; skills/*/SKILL.md is regenerated
via `npm run gen:plugin-skills` from that source. Updates the one
hardcoded copilot-install test assertion affected by gsd-health's
new tool order.

* chore(#3085): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#2618): assert the Needs clause on the structured field, not the path-length-dependent bullet

The bullet embeds the todo file's ABSOLUTE path, so its total length
varies by runner tmpdir: on macOS CI, /private/var/folders/… plus the
test harness's gsd-test-run-* wrapper pushed the full bullet to 244
chars — past renderPendingTodoBullet's intended 240-char cap, whose
documented first degradation step drops the Needs clause. The product
behavior is correct (#2618 design); the assertion was runner-dependent.
The extraction is now pinned on json.todos[0].needs; the rendered-bullet
shape stays covered by the path-independent title assertion and the
renderer's own unit rows.

Found blocking #4186's CI on the macOS shard (test landed 30 minutes
earlier in b7406b293f / PR #4384).

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 13:37:47 -04:00
Tom Boucher
b7917882bb fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure

The 240-char pending-todo bullet cap must be deterministic w.r.t. where the
repo is checked out. Deterministic long-base-path fixtures (a single 110-char
segment, no real macOS dependency) reproduce next's own macos shard 3/3
failure (run 34038716700) on every OS: with an absolute link the bullet
exceeds the cap and the documented needs-first truncation drops the
'Needs <solution>' clause. Rows cover the determinism property (byte-identical
bullets under short and long bases), the CLI surface, relative-path stability,
legacy no-projectRoot behavior, drop-order preservation, and adversarial
edges (outside-root, path===root, non-string path).

* fix(#4384): render the pending-todo bullet link repo-relative

renderPendingTodosMarkdown gains an optional projectRoot; when given and the
todo's path is absolute, the bullet's [todo file](…) target becomes
toPosixPath(path.relative(projectRoot, path)) — the idiom already used for
project_exists. cmdInitTodos passes cwd.

The JSON todos[].path field stays absolute (#2376). Only the rendered display
link changes: embedding the machine-variable absolute base let macOS's
/private/var/folders/… temp paths consume the 240-char budget and drop the
'Needs' clause on long-path machines only — next's own macos-latest shard 3/3
went red on exactly this (run 34038716700), Linux's short /tmp passed. The
240-char whole-bullet cap and the needs→title→area drop order are unchanged;
this matches PR #4384's own canonical example, docs, and unit tests, which all
show repo-relative links. Docs updated at all three surfaces that describe the
bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was
still pre-#4384 'count and reference' prose).

Fixes the macOS regression introduced by #4384; next is red on its own CI.

* test(#4384): fix substring false positive in the outside-root regression row

The ../-form relative link legitimately contains the absolute path as a
substring, so !line.includes(absolutePath) fired on correct output (caught by
the first remote verify run, linux-node24 44018/44019). Assert the property
itself instead: extract the link target and require it to be non-absolute and
not equal to the absolute path.

* chore(#4398): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 13:23:26 -04:00
Tom Boucher
aad04f4e9a docs(#4400): ADR-4139 — the compact-content seam (#4410)
* docs(#4400): ADR-4139 — the compact-content seam

Phase 0 of epic #4139. Locks the design before any code lands.

#4139's stated mechanism cannot reach the stream it exists for: 58 of 72
shipped commands deliver their whole workflow file through an eager
@-include, which the host expands before any project config is in context.
An in-content gate is evaluated after those bytes are already paid.

The ADR declines the obvious fix (convert the 58 execution_context blocks
to runtime Reads) because that removes the host guarantee for every user,
not only opted-in ones — a global install shares one skill tree, so an
@-include cannot be conditional. It instead keeps every @-include exactly
where it is and splits what sits behind them: the canonical path becomes a
runnable spine, elaborations move to a sibling detail file read at runtime.
A missed Read then degrades to "runs correctly with fewer tokens", never to
"runs with no instructions".

Also records: the rename to workflow.compact_content, the re-pitch onto
ADR-1610's context-rot argument rather than the cost argument ADR-1610
discounts, partition-not-duplication (which dissolves the dual-maintenance
cost the Feature Review called disqualifying), the guard-scope and
NEW_FILE_CAP mapping for the new subtree, and an argued reconciliation of
the acceptance criteria this design does not meet literally.

Corrects ADR-3646 §Context: it cites #3647 as open; #3647 closed
2026-09-01 as a duplicate of #3606. ADR-3646's Decision is unaffected —
it explicitly disclaimed any dependence on #3647's state. The residual
prose-dispatch variance named in #3647's own closure thread is unresolved,
and this ADR routes around it rather than assuming it away.

Refs #4139
Closes #4400

* docs(#4400): fold the two orthogonal review findings into ADR-4139

Code review (isolated context) and security review (isolated context) both
returned findings. Fixed here rather than carried.

Critical, from code review: the ADR repeated earlier research's claim that
discuss-phase, manager and pause-work all reach a workflow by runtime Read.
manager and pause-work carry plain eager @-includes and are inside the 58,
not outside. discuss-phase is the only precedent, and it is one file. The
Open Questions section is corrected with it.

The NEW_FILE_CAP mapping was wrong in a way that changes the layout. It
lives at tests/helpers/emitted-diff.cjs:96, not in workflow-size-budget,
and its own doc comment records that it is a hard cap, not ack-able, and
NOT tier-exemptible -- the pre-#2724 test-file version was. So a single
detail.md holding plan-phase.md's elaborations is blocked outright with no
exemption path. Detail content is now one or more parts under
workflows/<name>/detail/, each below the cap, named by the spine in the
dispatch-table shape discuss-phase.md already uses.

commit-files-pathspec is in scope and earlier research called it
irrelevant. Per CONTRIBUTING.md:1164-1170 it sweeps every .md under
gsd-core/workflows/ for unscoped commit-seam invocations. Added to the
guard table.

Two byte figures were inherited rather than measured, against this ADR's
own evidence note. Templates and agents re-measured; the table now carries
the method and the exact numbers.

From security review: "a spine that has shed a protected-content marker
fails" never defined what a marker was, leaving the strongest check in the
set resting on a prose-category judgment. Protection is now a literal
greppable sentinel in the existing gsd: comment namespace, and the guard
rule has no discretion in it. Also added: an explicit statement that the
detail path is never user- or project-supplied and cannot be shadowed by a
project-local file, and an exact-version pin commitment for gpt-tokenizer.

Code review also found a real hole in the central fail-safe argument:
spine sufficiency is verified once at split time and never again, so
load-bearing procedural text carrying no sentinel could later drift into a
detail part with every check green. Sufficiency is not machine-decidable,
so a fifth ongoing check is added -- a spine that loses lines which
reappear in its parts fails unless the PR declares the boundary move. The
ADR now states plainly that this is authoring discipline with a forced
checkpoint, not a structural invariant, and that the partition relocates
the Feature Review's cost rather than fully eliminating it.

Refs #4139
Refs #4400

---------

Co-authored-by: sim <sim@local>
2026-09-06 12:24:33 -04:00
Tom Boucher
708d9a0b82 fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file (#4388)
* test(#4187): bare VERIFICATION.md regression matrix for the status surface

Both query verbs must agree on every row: bare file, suffixed variants,
missing file, other-dir placement, and the staleness seam. Row 1 is the
failing-first regression from the issue repro.

* fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file

readVerificationStatus and its internal staleness check
(findStaleVerificationSummary) called the shared resolver without
allowBare, so a phase whose only report was a bare VERIFICATION.md read
as missing and was told to re-run execute-phase while
verification.resolve-file, determinePhaseStatus, and both init
verification_path projectors all resolved the same file. Both call
sites now pass allowBare: true, matching the other five; tier order
(dashed > bare) is unchanged, so only bare-only directories change
behavior.

* fix(#4187): correct call-site counts in allowBare docblocks

Adversarial review caught the comments claiming five of six call sites
opted in; the current tree has six call sites with four previously
passing allowBare — the two module-internal status-path sites were both
holdouts, not one.

* chore(#4187): changeset for the bare VERIFICATION.md status fix

* chore(#4187): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 11:54:03 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
b7406b293f enhance(#2618): render pending todos as one bounded bullet per todo (#4384) 2026-09-06 08:06:39 -04:00
Tom Boucher
66e4034fe4 fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380)
* test(#4138): failing-first regression — begin-phase without --phase must fail closed

* fix(#4138): begin-phase without --phase exits non-zero and writes nothing

* chore(#4138): changeset fragment for begin-phase arg validation

* chore(#4138): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 07:03:39 -04:00
Tom Boucher
03738824de enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) 2026-09-06 05:46:49 -04:00
Tom Boucher
0aa4202f6a fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376)
* test(#4135): regression rows for pristine regen coverage collapse

RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a
null-returning stub (wired into verifyFile after the #4145 orphan tier,
behavior-neutral) so the git-history rows fail behaviorally, not at require
time. Failing-first rows: baseline_covered aggregate on a 1-of-13
multi-version fixture, coverageHeadline typed renderer, the opt-in
--min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass
and malformed-value boundaries), git-history baseline recovery (dropped-line
catch + surviving-line verify + older-commit hop), findPristineInGit unit,
Step 5a workflow headline contract, and the installer-side
describeBaselineCoverage honest N-of-M summary with the collapse disk-state
pinned. Negative-space rows pin today: non-git ok_no_baseline posture,
no-match-no-adoption, #3657 drift never rescued, canonical precedence, and
no git tier without --pristine-dir.

* fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening

The #3407 promotion rule regenerates gsd-pristine/ baselines from the
INCOMING release source and keeps only candidates byte-identical with the
OUTGOING recorded hash — correct in isolation, but on a multi-version jump
the surviving set is precisely the files upstream did NOT change. The
verifier then reports ok_no_baseline (advisory, exit 0) for everything
else, and no surface distinguishes a 12-of-13-unverified green run from a
fully-verified one: the human summary printed Checked/Failures only, the
JSON had no coverage aggregate, and the installer's update output gave
per-bucket counts without N-of-M framing.

All three issue directions, none exclusive:

- Report coverage prominently: --json gains an additive baseline_covered
  aggregate; the human summary leads with 'Baseline coverage: N of M
  file(s)...' on every run plus an advisory section naming each skipped
  file and reason; the installer prints an honest covered-of-modified line
  via the exported describeBaselineCoverage helper (typed return, exact
  contract); workflow Step 5a computes and prints the headline before any
  pass/fail framing.
- Fail louder on low coverage: opt-in --min-baseline-coverage <0..1>
  exits with new documented code 3 when coverage falls below the
  threshold (>= semantics; empty run vacuously passes; content failure
  exit 1 outranks it; malformed values are usage errors, exit 2).
  Default posture unchanged — no_baseline stays advisory per #934.
- Widen the promotion rule (its only trustworthy form): when no baseline
  resolves under gsd-pristine/ and a hash is recorded, the verifier now
  recovers the baseline from the config dir's own git history — the
  workflow's documented Option A — anchored by the same authority every
  tier trusts, exact pristine_hashes sha-256 equality. Read-only
  (git log/git show, windowsHide per #685), bounded (100 commits/file,
  10s/subprocess), null-on-any-failure so ok_no_baseline remains the
  universal fallback. Tier order: canonical join -> #4145 orphan scan ->
  git history -> OK_NO_BASELINE; #3657 drift and canonical precedence
  untouched.

Hash validation in saveLocalPatches is NOT relaxed — the collapse is
legitimate conservatism; hiding it was the bug. Measured on the issue's
shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report
baseline_covered 1/13 with the headline and can gate at exit 3; a
git-managed config dir with the outgoing bytes in history verifies 13/13.

Review fixes folded in: workflow headline derives the unverified count
from checked - baseline_covered (not the drift+no_baseline sum), and the
new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on
the marker line.

Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it.

* chore(#4135): backfill PR number into changeset fragment

---------

Co-authored-by: agent-4135 <agent-4135@gsd.local>
2026-09-06 05:26:05 -04:00
Tom Boucher
c95b734145 fix(#4136): compute the Incorporated status; stop re-grafting superseded patches (#4373)
* test(#4136): failing-first rows for the unreachable incorporated status

Folded block bug-4136-reapply-incorporated-status locks the --classify
contract (incorporated / needs_merge / unknown, never-incorporated
guards, cycle end-to-end) plus the workflow-contract rows; REASON gains
OK_UNVALIDATED_BASELINE in both shape-locks; the #2994 invocation count
moves 1 -> 2 (classify + gate). All red until the verifier grows
--classify and the workflow consumes it.

* fix(#4136): compute the incorporated status; stop re-grafting superseded patches

Add --classify pre-merge mode to the deterministic verifier: with a
hash-validated pristine baseline, a file whose every significant
user-added line is already present verbatim in the freshly installed
version is classified incorporated (new frozen CLASSIFICATION enum,
structured --json report, always exit 0 — the binding gate stays the
post-merge run). Drifted (#3657), absent (#934), unvalidated (new
OK_UNVALIDATED_BASELINE), and no-baseline runs classify unknown — a
false incorporated silently retires a live customization, so only a
confirmed baseline may ever confirm adoption.

The baseline-resolution block moves out of verifyFile into a shared
resolvePristineBaseline so the gate and the classifier cannot drift on
what counts as a usable baseline; gate behavior is byte-identical.

reapply-patches.md step 4 gains the pre-flight classifier invocation
and the not-re-grafted contract: incorporated files are left exactly as
shipped (their hash then re-converges with the manifest, ending the
backup cycle), statuses feed steps 3/7, and the merge rules gain the
already-present-verbatim arm.

Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract

* fix(#4136): address review findings on the workflow contract

Drop the unused INCORPORATED_COUNT shell variable (standards pass) and
close the all-files-incorporated gap in the hunk-table guidance: emit a
header row plus a note line so the step 5b absent-table halt is not
tripped when nothing was merged (spec pass).

Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract

* chore(#4136): backfill changeset fragment with PR 4373

* test(#4136): lock that a 4145-recovered baseline can confirm incorporation

The hash-first orphan recovery merged with next (PR #4364) lands in the
shared resolvePristineBaseline as a validated resolution; this row pins
the composition so a future change cannot quietly downgrade recovered
baselines to unknown and silently disable incorporated detection for
prefix-less installs.

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:41 -04:00
Tom Boucher
7bb366e836 fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening

Block A (flag): check decision-coverage-plan --context <path> must route
identically to the positional form; flag wins over positional context;
valueless --context falls through to the #2770 fail-closed caller error;
verify keeps its positional surface (flag is plan-only). RED on base:
the flag token lands in the args[2] phase slot (false uncovered) or the
args[3] context slot (silent CONTEXT.md-missing skip).

Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper
(?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the
no-adjacent-overlap property; a differential property compares the module
against a frozen copy of the pre-hardening grammars (reference validated
against the base build: 60k generated lines, 0 mismatches); 40k cliff
shapes assert correct outcomes with no wall-time asserts (repo rule).

A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags.

* fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions

(A) check decision-coverage-plan --context <path> — sibling convention
(check predicate, #2008): --flag value pairs parsed by the new shared
partitionPredicateArgs (parsePredicateFlags reimplemented as its flags
half — one parser, cannot diverge), the flag winning over a same-purpose
positional, positionals kept (no sibling deprecates them; the plan-phase
workflow caller passes positionals), valueless --context falls through
to the #2770 fail-closed caller error. Repair of the routing accident
where --context landed in the args[2] phase slot (false uncovered) or
the literal token in the args[3] context slot (silent green skip).

(B) parseDecisions regex seam hardened, byte-identical on all legal
inputs: the three bullet grammars consume the ID atomically via the
(?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split,
~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to
[^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group
indices unchanged (handlers untouched). Pinned by regex-lattice tests,
a differential fast-check property vs the frozen pre-hardening grammars,
and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo
rule — no deterministic engine step counter exists in Node).

* docs+test(#4130): document --context invocation; harden lattice test tooling

- docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan
  gate directly' block documenting both the positional and --context
  forms, flag precedence, and the valueless-flag fail-closed semantics
  (same place the gate's behavior is documented; sibling check predicate
  documents its flags the same way).
- Two changeset fragments per the maintainer brief (Added: flag; Fixed:
  hardening), PR numbers to be backfilled.
- tests/decisions.test.cjs review fixes: readRegExpTemplate template
  escaping (bare ')' SyntaxError), range-aware lattice checker with
  backreference skip and template unescape, honest A1 contract, lint
  escape warning.

* fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context

Suite-caught fixes from the first verify run:
- cmdDecisionCoveragePlan now refuses a flag-shaped token as the
  positional context path: a bare valueless --context stays a positional
  (sibling parser semantics, unchanged) but reading it as a PATH would
  turn a caller mistake into a silent 'CONTEXT.md missing' green skip —
  exactly what #2770's fail-closed law forbids. Now falls through to
  the missing-context-argument error, as documented.
- A8 test compares decoy-positional+flag against flag-with-phase (phase
  held constant) so the row isolates WHICH context was read; the old
  form compared against a no-phase invocation that could never match.

* chore(#4130): backfill PR number in changeset fragments (PR #4374)

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:17 -04:00
Tom Boucher
6adf3098ac fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines

RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a
null-returning stub so the new rows fail behaviorally, not at require time.
Failing-first rows: verifier resolution (no_baseline must drop to 0 when an
exact-hash orphan exists), findPristineByHash unit row, and the two
saveLocalPatches relocation rows. Negative-space rows pin today's behavior:
missing baselines still report ok_no_baseline, mismatching orphans are never
adopted or deleted, canonical precedence and the #3657 drift posture are
untouched.

* fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans

Both pristine readers joined the manifest-keyed path strictly, so a snapshot
stored without the gsd-core/ prefix (an earlier release's writer) was reported
as ok_no_baseline by the verifier and pushed into regeneration by
saveLocalPatches — where incoming-release candidates can never satisfy the
recorded outgoing hash, leaving the correct baseline permanently unconsumed.

- src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash —
  deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the
  recorded pristine_hashes entry (the same authority the #3657 drift guard
  trusts), symlink-skipping, canonical path excluded via skipRel.
- verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded
  hash, adopt byte-identical content found anywhere under gsd-pristine/ before
  reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and
  the frozen REASON/report shapes are untouched; the verifier stays read-only.
- install.js saveLocalPatches(): preserve-check rescue — relocate a
  hash-matching orphan to the canonical path (copy, hash-verify, then remove
  the orphan) so the state self-heals on the next update instead of repeating
  forever. Honest accounting: new non-overlapping rescued counter.
- Workflow doc: one-sentence note on hash-based snapshot resolution.
- Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore
  entries for the compiled artifact, seedFixture mkdir fix in the new rows.

Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145)

* fix(#4145): review follow-up — orphan scan never consumes a canonical path

Adversarial review finding: with two modified files sharing byte-identical
outgoing content, recoverOrphanedPristine could adopt the OTHER file's
canonical pristine as its rescue source — relocating it (copy + delete at
its home path) and ping-ponging the single baseline between the two files
across updates. findPristineByHash's skip parameter now accepts a Set, and
saveLocalPatches passes the normalized manifest keys so every canonical
path is excluded; only genuine non-canonical orphans are eligible for
removal (no strict-join reader ever consults those). Adds the
canonical-theft regression row, a Set-skip unit assertion, and tightens the
workflow doc sentence the same pass flagged as overstated.

* fix(#4145): INVENTORY roster row + symlink-fixture correction

Two leftovers from the ab17b7a1e5 bench run, both root-caused:
- docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs
  (#3762 gate: every manifest entry carries a row).
- The findPristineByHash symlink unit fixture placed its symlink target
  INSIDE the scanned root, so the walk legitimately matched the real target
  file. The implementation skips the symlink itself; the fixture now keeps
  the target outside the scanned tree so the assertion tests what it claims.

* changeset(#4145): fixed fragment for pristine baseline hash resolution

---------

Co-authored-by: gsd-agent <agent@gsd.local>
2026-09-06 02:04:18 -04:00
Tom Boucher
d5a85da8ab fix(#4363): bump download-artifact and setup-node off node20 runtimes (#4365) 2026-09-06 00:03:05 -04:00
Tom Boucher
c3e2da153b fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names

A first-milestone ROADMAP.md H1 that puts the version after the name
(# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's
own version token, which the ADR-3180 §7.2 pinned name rule returns as a
COMPLETE-scope milestone name. Failing-first coverage:

- getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only
  fallback) must yield TRUNCATED {version, name: null}, never ')'
- the refusal is level-agnostic (H2/H3)
- punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only)
- listMilestoneHeadings enumerates the heading with name: null
- init manager CLI reports milestone_name: null and no lone ')' anywhere
- property (seed 20260905, 300 runs): a word-char remainder is always a
  name, a punctuation-only remainder never is
- negative space: canonical delimiter forms, parenthetical names (#3171),
  trailing markers, digit-only names, CRLF headings, version-last-no-parens
  control

* fix(#4134): refuse punctuation-only milestone heading names

extractMilestoneHeadingName returns everything after the heading's own
version token as the name (ADR-3180 §7.2 pinned rule), which assumes
version-then-name. A name-then-version heading — the H1 a first-ever
ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves
exactly ')' after the token, and that fragment was returned as a
COMPLETE-scope milestone name, propagating into init.* JSON output and
buildStateFrontmatter's STATE.md writes.

A remainder with no letter or digit anywhere (any script) is heading
structure, not a curated name: refuse it as name: null so callers report
the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names
that merely contain punctuation are unaffected — '(' stays an ordinary
name character (#3171) — and digit-only names qualify.

Also closes the template gap that lets the shape occur: the roadmapper
agent's output_formats now templates the version-free canonical H1
('# Roadmap: [Project Name]', per templates/roadmap.md) instead of
leaving a first milestone's title line to invention. The new section
shifts the file's existing bare-gsd-tools prose mention from line 647
to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it.

Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134)

* chore(#4134): add changeset

* chore(#4134): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 23:31:35 -04:00
Tom Boucher
e6d047decc fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359)
* test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment

* fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet

Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/):

P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at
the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the
one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch
that owns the denominator. A completed phase whose verification routes stale
(#2348 clean-commit-time drift) or is missing no longer under-counts forever.

P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on
a measured scan: totals derived both directions (#2440), completed counters
up-only (#2969 — the schema-declared progress-ratchet, now enforced on the
write path like the read path always has), percent recomputed from the merged
counters. The #3756 unmeasured guard and the #3242 explicit-progress contract
are unchanged.

P3 — phase complete's atomic 3-file commit passes the post-completion
ROADMAP-derived counters through the #2736 authoritativeFm seam (new object
direction for the progress key; completedOnlyRaise at the post-preservation
re-assert), because the transaction's disk scan reads the pre-completion
ROADMAP and failed to increment on the completing phase's own write.

* fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites

The pre-preservation merge could lower a correctly-higher disk-derived
counter (a verification-passed phase whose ROADMAP table row drifted behind
the disk signal). completedOnlyRaise now governs both application sites: the
intent and the derivation agree on direction (up), never on subtraction.

* fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal

The re-parsed derived block carries string scalars ("2") while the curated
snapshot carries numbers (2); substituting the curated spelling over an
equal derived one was a no-op in substance but a shape churn the ADR-3473
§8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived
warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater
curated counter replaces the derived value now; percent gets the same
verbatim rule.

* changeset(#4129): backfill PR 4359

---------

Co-authored-by: sim <sim@local>
2026-09-05 23:02:07 -04:00
Tom Boucher
bd75d42f52 Merge pull request #4362 from open-gsd/chore/backmerge-main-to-next-b67f6028c
chore: back-merge main → next (b67f6028c)
2026-09-05 22:11:10 -04:00
github-actions[bot]
0293afd108 chore: back-merge main into next (b67f6028c) 2026-09-06 02:10:44 +00:00
Tom Boucher
519bb60263 Merge pull request #4361 from open-gsd/chore/sync-next-version-1.13.0
chore: sync next package version to 1.13.0
2026-09-05 22:10:36 -04:00
github-actions[bot]
b3906c66f6 chore: sync next package version to 1.13.0 2026-09-06 02:10:29 +00:00
Tom Boucher
b67f6028ce Merge pull request #4360 from open-gsd/release/1.13.0
chore: merge release v1.13.0 to main
2026-09-05 22:10:26 -04:00
github-actions[bot]
b5b9814f03 chore: promote CHANGELOG for v1.13.0 2026-09-06 02:09:37 +00:00
github-actions[bot]
d0bf2c5165 chore: finalize v1.13.0 2026-09-06 02:09:28 +00:00
Tom Boucher
06eba5fdb0 fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs

Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags,
discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and
parity properties (well-formed digit-prefixed ids parse to their exact id;
a non-digit injected into the prefix fails loud). Update the #2347
non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and
graduate the representative d5-prefix corpus fixture from could-not-parse
to parsed-but-uncovered.

All new rows are RED against origin/next; they go green with the parser
fix in the next commit.

* fix(#4130): parse phase-prefixed decision IDs (D4-01)

The three declaration grammars, the parse-miss guard, the #3939 join
regexes, and the token evidence all anchored on the literal 'D-' (or
'**D-'), so an ID carrying a digit-run phase prefix between the leading
letter and the hyphen matched nothing — while the #2347 shape detector
correctly called those bullets decision-shaped, collapsing the whole
CONTEXT.md to could-not-parse with 0 extracted instead of a coverage
verdict.

Derive the extractor ID grammar from one shared DECISION_ID_SOURCE
('D[0-9]*-' + the existing alnum tail, full id captured), widen the
guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial
prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose
like 'Deferred-until' stays none-present), and align the bare-token
evidence. Both gates and the gap-checker share the parser, so all three
surfaces read phase-prefixed decisions now; the gate messages name the
accepted forms including the phase-prefixed one.

* docs(#4130): document the phase-prefixed decision identifier form

The canonical CONTEXT.md reference said decisions carry 'a sequential
D-NN identifier' with no mention of the optional phase-number prefix the
parser now accepts (D4-01) or the alphanumeric tail it always accepted
(D-INFRA-01). Name both in the Decision identifier format section, EN
and ja-JP.

* chore(#4130): changeset

* chore(#4130): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 21:05:19 -04:00
Tom Boucher
48789fe9a9 fix(#4355): add --merge-async to test:coverage:unit:raw (#4356)
* test(#4355): assert test:coverage:unit:raw carries --merge-async

tests/c8-merge-async-flag.test.cjs asserted the opposite based on a
disproven assumption that `--reporter none` skips c8's merge phase --
it does not (Report.run() computes the merge unconditionally before
consulting the reporter list). This is the failing-first assertion for
the fix in the next commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4355): add --merge-async to test:coverage:unit:raw

c8's Report.run() computes the full coverage merge unconditionally,
even with --reporter none -- it only skips the final text/json output,
not the merge dispatch (verified against node_modules/c8/lib/report.js).
Without --merge-async this used the synchronous _getMergedProcessCov(),
loading every raw per-process V8 coverage dump into memory at once and
OOM-crashing release.yml's finalize-test job (run 33997057100) with a
silent exit 1 and no diagnostic ~21s after the test suite itself
finished cleanly ("# fail 0"). Same class already fixed on the other
three coverage scripts via #4068/#4172.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(changeset): add changeset for #4355 coverage:unit:raw fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(changeset): backfill PR number for #4355 fix

pr:0 -> pr:4356

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 20:28:51 -04:00
Tom Boucher
a65cb291e8 fix(#4105): park the #3889 hang fixture on a settling timer (#4349)
* test(#4105): guard the #3889 hang fixture — must genuinely hang by itself and self-terminate

RED at this sha: against the current never-settling-promise body the guard
fails on the matrix line (Node 24: the unheld promise never self-terminates,
ceiling expires) and off it (v22-class runtimes: the child exits rc=1 after
~60ms, never reaching the still-hanging checkpoint). Same shape as the #4104
self-exit regression: spawn the exact served body, observe liveness past the
chunk bound and natural exit — no elapsed-value assertions.

* fix(#4105): park the #3889 hang fixture on a settling timer

The never-settling promise held no libuv handle, so the hang T1/T4 rely on
was a property of the runtime's test-runner shutdown behavior, not of the
fixture: v24/v26 happen to hold the loop open; v22-class runtimes exit rc=1
after ~60ms (# cancelled 1), so the chunk never reaches the timeout path and
the two timeout assertions assert nothing. Park on a settling 10s timer (the
#4104 idiom): an explicit handle makes the hang the fixture's on every Node
line, 10s >> the 2000ms chunk bound (margin asserted structurally in the
#4105 guard), ~0% CPU while parked, and guaranteed self-termination if a
kill orphans it. Behavior on the Node 24 matrix line is unchanged — the
chunk is still killed by the harness timeout (~2006ms) with the identical
diagnostic.

* test(#4105): drive the fixture guard off the child's exit event + runner timeout

Review-driven restructure (Memtrace flaky_test_fixed_sleep on the 200ms poll
interval): the guard now waits on the child's natural 'exit' event — no
polling interval, no hand-rolled watchdog setTimeout. The immortal-body bound
is the node:test per-test { timeout: 2 * HANG_PARK_MS } backstop, the
health-validation #663 house pattern and the no-elapsed-assertion-compliant
form. t.after still reaps the child on every path. Same failing-first arms:
still-hanging checkpoint, natural-exit (no signal), exit code 0.

* changeset(#4105)

* changeset(#4105): backfill PR number

---------

Co-authored-by: sim <sim@local>
2026-09-05 19:39:35 -04:00
Tom Boucher
0be5bf865a enhance(#3783): audit-uat summary segments current-milestone vs archived debt (#4336)
* test(#3783): add failing coverage for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3783): segment audit-uat summary into current_milestone and archived buckets

Additive: current_milestone/archived are new; total_items, total_files, parse_gap_files, by_phase, and by_category are unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): add changeset fragment for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3783): allowlist the new audit-uat-summary-segmentation test file

lint-test-file-count.cjs baselines the "audit" module (keyed off bin/lib/audit.cjs)
at 6 pre-existing files; this adds the new dedicated suite as a 7th, matching the
module's existing one-file-per-feature-slice precedent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3783): fix phase/file number mismatch in the mixed-milestone fixture

The active phase fixture used dir "02-current" with file "01-UAT.md" — a
cross-phase stray per phase-id.cts's isPhaseArtifact/scopeToPhase (#3511),
so the file was silently excluded from the scan and current_milestone read
{files:0, items:0} instead of {files:1, items:1}. Confirmed by direct CLI
run against a hand-built fixture before recommitting. Renamed the file to
02-UAT.md to match its directory's phase number, matching every other
fixture in this suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): backfill changeset PR number to 4336

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 19:03:14 -04:00
Tom Boucher
f9f72cb54c enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch

Extracts and executes the real bash blocks this PR is about to add to
plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and
the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the
new planning.chunked_parallel key. Expected RED against the current shipped
workflow text — the extraction anchors do not exist yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave

Adds opt-in planning.chunked_parallel (default false, byte-identical to the
existing serial loop). When true and the runtime's negotiated dispatch
capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's
per-plan Tasks that share one outline Wave are issued together instead of
one at a time; a later Wave still waits for the current one to be verified
on disk and committed. A host with no declared maxConcurrency (most
non-Claude runtimes today) stays serial regardless of the setting.

Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md
itself (gated on the section's own CHUNKED_MODE skip-check) rather than in
plan-phase.md, so a non-chunked run pays no extra gsd_run calls.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3777): repoint extraction at chunked-planning-mode.md after the move

CHUNKED_PARALLEL resolution moved out of plan-phase.md into
chunked-planning-mode.md itself (see the preceding commit); update the
test's extraction path and header comment to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): relocate the canonical runtime-launcher preamble before its first use

The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in
the file than the sole existing preamble (in the commit step), which
tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every
gsd_run call in the file. Move the preamble (not duplicate it) to the top of
the resolution block; the commit step's fenced block now just calls
gsd_run directly.

Caught by the GREEN checkpoint gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): strip the canonical preamble from the extracted resolution block

The CHUNKED_PARALLEL resolution fence now carries the relocated
runtime-launcher preamble as its first line (previous commit). Extracting
the whole fence and running it after the test's own gsd_run stub let the
embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before
reaching the resolution logic, since no real gsd-tools.cjs exists in the
temp script dir — every test calling runChunkedParallelResolution() failed.

Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh,
the same file scripts/sync-runtime-launcher.cjs treats as canonical) before
splicing in the stub, so this suite tests only the resolution logic it is
actually about.

Caught by the post-rebase gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): add the How-To page the phase gate requires

Enablement is 2 commands (config-set, then --chunked), which this repo's
own doc-quadrant gate flags as how-to-owed: a reference table cannot carry
a sequence. Covers enablement, the dispatch-capacity gate's honest
"most runtimes today: no effect" case, and the two accepted trade-offs.

An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before
this commit) had incorrectly claimed #3034 shipped with no equivalent
how-to page, as precedent for skipping one here. That claim was false —
docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The
phase gate caught the omission before merge; corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:57:17 -04:00
Tom Boucher
1db726ebbf feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract

Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md,
and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions
Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only
supersession rule. These fail until the canon and its two references are added.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3806): canonize the Review Dispositions Ledger contract

Promote the existing planner-reviews.md Step 4 return-payload tables
(Review Feedback Addressed/Deferred) into a canonical `## Review
Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md
and referenced (not restated) from plan-phase.md's
<review_incorporation_contract> and gsd-plan-checker.md's Review
Incorporation dimension. Adds round-scoping (`### Round {N} —
{REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md
reference survives the file being rewritten each round, and an
append-only supersession rule. Scoped to part 1 only per the
maintainer's approved-feature verdict — the deterministic lint/check
verb (part 2) is explicitly deferred to a follow-up.

Also: ADR-3806 recording the decision, a docs/features/ fragment
(FEATURES.md is generated), and a changeset fragment.

Closes #3806

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fenced-example count bug and lint findings from review

- tests/plan-review-convergence.test.cjs: the "heading exactly once"
  test counted the canonical heading text globally, so it also matched
  the illustrative fenced-code example in planner-reviews.md that shows
  the same heading as sample content, always failing 2 !== 1. Rewritten
  as a bounded line scanner that skips fenced blocks (found by an
  isolated adversarial review pass). Also bounded an unbounded regex
  quantifier over readFileSync content flagged by
  local/no-unbounded-quantifier.
- docs/features/review-dispositions-ledger.md: match house fragment
  style (bold-lead paragraphs, not #### headings) per the Standards-axis
  review; regenerated docs/FEATURES.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fit reference-cite fix within size hard caps; ack growth

Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a
single short clause pointing at gsd-core/references/planner-reviews.md
(also fixes the bare `references/planner-reviews.md` cite the #3576
shipped-reference-cites gate rejects), bringing both files back under
their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline.
Both files still grow slightly versus origin/next, acknowledged below
per ADR-2719's emitted-drift-ack contract.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block

The previous commit's two Emitted-Drift-Ack-Growth trailers were
separated from the Co-Authored-By trailer by a blank line, so git's
own trailer parser (which tests/helpers/emitted-runtime.cjs reads via
`%(trailers:key=...)`) only recognized the last contiguous block
(Co-Authored-By) and treated the Ack-Growth lines as ordinary body
text — invisible to the emitted-attribution gate, not malformed data.
Restating them here as one contiguous trailer block, git log over the
PR range aggregates trailers from every commit, so this is additive.
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): isolate the ack-trailer paragraph as its own trailer block

Git's trailer parser requires the trailer paragraph to be the message's
final paragraph, preceded by a blank line, and to contain nothing but
trailer-shaped lines. The prior commit's blank line before the trailer
lines was missing, which folded the leading Emitted-Drift-Ack-Growth
lines into an ordinary prose paragraph.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3806): backfill PR #4345 into changeset and ADR

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:50:27 -04:00
Tom Boucher
ea91268d02 ci(#4335): shard release.yml rc/finalize unit-suite tests (#4338)
* ci(#4335): shard release.yml rc/finalize unit-suite tests

The finalize job's unsharded unit-coverage step outgrew the 30-minute job
timeout that was already raised once for this exact symptom (#2280): run
33988966357 finished all tests with 0 failures at 28m26s, then got cancelled
~80s into the post-test coverage merge — a phase that historically completes
in 54-101s. The suite's wall-clock time, not a hang, ate the budget.

test.yml already fixed the identical cliff for its own full-scope lane
(#2952, #3057) by sharding the unit suite 3 ways with a separate merged
coverage-gate job. Apply the same pattern to rc and finalize (rc has the
byte-identical unsharded shape and would hit the same wall next): each gains
a `*-test` matrix job (raw coverage only, no report/gate) and a
`*-coverage-gate` job that merges the shards' raw V8 dumps before enforcing
the existing gsd-core/bin/lib coverage floor. rc/finalize now depend on their
gate job instead of running the suite inline.

Updates release-coverage-scope.test.cjs's exact-count assertion for the new
command surface and adds release-shard-lane-sharding.test.cjs to pin
shard-set completeness and gate wiring, mirroring ci-full-lane-sharding.test.cjs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Potential fix for pull request finding 'CodeQL / Cache Poisoning via execution of untrusted code'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Potential fix for pull request finding 'CodeQL / Cache Poisoning via execution of untrusted code'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix(#4335): close CodeQL cache-poisoning and missing-permissions findings

CodeQL flagged the PR (10 actions/cache-poisoning/poisonable-step errors, 4
actions/missing-workflow-permissions warnings) on release.yml.

Remove `cache: 'npm'` from every actions/setup-node step in the file (7
occurrences, not just the 4 newly-added jobs the alerts pointed at) —
restoring an npm cache before running install/build code in a
write-permissioned job is exactly the shape this query targets, and the
same pattern was already present unchanged in create/rc/finalize. These are
short CI/release jobs; losing npm's install cache costs a few seconds per
job, closing the finding everywhere it appears in this file rather than
only where the alert happened to land on a changed line.

Add explicit `permissions: contents: read` to rc-test, rc-coverage-gate,
finalize-test, finalize-coverage-gate — the four new jobs had no
permissions block at all and inherited the ambient default. Matches
validate-version's existing least-privilege pattern; create/rc/finalize
keep their own broader write/publish scopes unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-09-05 18:49:35 -04:00
Tom Boucher
c20675cc4d fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode

The pre-commit protected-branch assertion in the executor agent (#2924)
only fired inside a Claude Code worktree and matched a hardcoded
five-name branch list. It never ran in an ordinary checkout and never
covered this repo's own default branch ("next"), so gsd-executor could
commit planning-repo documents directly onto a shared checkout's
default branch with no PR ever created.

Widen the guard to run in every isolation mode, and resolve the
protected branch via the repository's actual default branch (with the
existing five-name list retained as a fallback when the resolver
itself cannot be invoked) plus any configured git.protected_branches.
Add a git.allow_default_branch_commits escape hatch for projects that
intentionally execute on their default branch. Also point the
separate <final_commit> commit helper back at the same guard, so it
cannot be sidestepped by that path.

Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3819): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:42 -04:00
Tom Boucher
4e1c449281 enh(#3811): add hooks.commit_types config surface to gsd-validate-commit (#4340)
* enh(#3811): add hooks.commit_types config surface to gsd-validate-commit

Extends the opt-in Conventional Commits hook with a hooks.commit_types
config array that adds project-specific types to the 10 built-ins
without replacing them. Configured values pass a safe-token filter
before reaching the compiled regex, so a config entry can never alter
the pattern's structure. The regex alternation, the human-readable
error text, and a new typed valid_types JSON field all derive from one
list instead of the two hand-synced copies this replaces.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3811): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:09 -04:00