Commit Graph

5777 Commits

Author SHA1 Message Date
Tom Boucher
1ec4b38bd3 test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch (#4300)
* test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch

Epic #4272 Phase 5's own checklist named this deliverable ("the same class
of coverage loop-walk.cjs gives the loop") separately from #4268. Adds
tests/qa/tdd-walk.cjs, extracting and REALLY EXECUTING (via a real `bash -c`
subprocess against a real temp fixture project) the shipped bash resolution
snippets from both TDD dispatch backends — never reimplementing or
grep-simulating the predicate.

Proves, by execution rather than text-shape assertion: the CLI predicate and
both backends agree for a type: tdd plan and a plain plan; the worktree
backend's fail-closed guard genuinely halts (non-zero exit, FATAL stderr) on
a missing plan file; and the tdd.md embed ternary's condition tracks the
real resolved value (#3800). This is exactly the class of proof #4264/#4265
(unassigned/divergent predicate) and #4268 (static-shape checks can't see
backend divergence) could not provide.

Extraction uses indexOf/slice on fenced-code markers only, never a
backtracking regex over whole-file text (per the #4228 incident this repo's
tests already document).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4298): scrub ambient env, tighten fail-closed assertion, fix comment

Standards+Spec review found: (1) executeBackendScript spread raw process.env
unfiltered into the spawned bash subprocess, unlike tests/helpers.cjs's
runGsdTools, which deliberately scrubs SESSION_IDENTITY_ENV_KEYS +
config-location env vars before spawning (#2665) — an ambient developer/CI
override could silently change what phase.tdd-applicable resolves to in a
way a gsd-test bench container won't reproduce; (2) the row-5 fail-closed
test asserted only `stderr.includes('FATAL')`, which would also pass if the
file's unrelated ISOLATION fail-closed guard fired instead of the TDD one;
(3) a docstring called the worktree backend's first fenced block a "shim
preamble" when it's actually the whole ISOLATION-resolution block.

Fixes: spread the exported TEST_ENV_BASE (every scrub-listed key set to '')
before the two intentional RUNTIME_DIR/GSD_TEST_MODE overrides; assert the
exact TDD-applicability FATAL text; correct the docstring. Re-verified by
direct execution against real fixtures — all three precedence-tier cases
and the fail-closed case behave identically to before the fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 20:17:36 -04:00
Tom Boucher
747a3730d4 fix(#4268): harden tdd-single-statement.test.cjs against reworded restatements and backend divergence (#4297)
* test(#4268): harden reworded-restatement and backend-predicate-divergence detection

tests/tdd-single-statement.test.cjs's restatesCycle() keyed on the exact
literal `commit: `test({phase}-{plan})`` substring, so a reworded restatement
of the RED/GREEN/REFACTOR procedure shipped green. Adds
restatesCycleStructurally(), a structural (span + list-marker) detector that
stays linear-scan (per the #4228 catastrophic-backtracking incident this must
not reintroduce) and is proven, empirically, to flag a paraphrased multi-step
fixture while not flagging the real compact citations in execute-plan.md and
gsd-executor.md (#4267's legitimate pointers).

tests/tdd-backend-wiring.test.cjs never compared the two dispatch backends'
`gsd_run query phase.tdd-applicable` calls against each other, so a one-word
divergence between them (e.g. a changed --pick flag in only one backend)
shipped green. Adds a byte-identity assertion on the command-substitution
content (normalized for the two backends' differing variable-name prefixes),
proven to have teeth via a RED-first mutation check before asserting it
against the real files.

The third gap in #4268 (nothing proves TDD_APPLICABLE has a real definition)
was already covered by this file's existing assertTddApplicableIsComputed
(epic #4272 Phase 2, #4266) — verified by inspection, no new test needed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4268): redesign restatement detector around deferral, not length

Standards+Spec review of the prior commit proved by execution that a
compact, no-list-marker restatement (under the 200-char span threshold)
sails past the span/list-marker-only signal. Redesigns the primary check:
the actual invariant is deferral, not length — a legitimate RED/GREEN/
REFACTOR mention always names tdd.md as the authority nearby, a restatement
never does. Flags when no tdd.md/canonical reference appears within a
500-char trailing window past the cycle mention, regardless of length or
list-marker shape; keeps span>200 and three-distinct-list-marker-lines as
secondary defense-in-depth OR-conditions. Verified independently against
both real files (execute-plan.md span=10, gsd-executor.md span=14, both with
a nearby deferral marker at +228/+82 chars) — no false positive, and the
reviewer's exact gap class (a 189-char no-citation paraphrase) is now
flagged.

Also fixes: boundary coverage at the span threshold (199/200/201, isolated
via a factored-out measureCycleSpan() helper), a fast-check property test
proving the fix holds for arbitrary filler text, and a fragile line-match in
tdd-backend-wiring.test.cjs that happened to work only because a FATAL echo
message containing the same substring came later in document order than the
real assignment line.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 19:48:14 -04:00
Tom Boucher
5214ad5802 fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication (#4295)
* fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication

execute-plan.md and gsd-executor.md each cited "Red-Green-Refactor Cycle"
for three facts (commit-scope contract, fail-fast rule, error handling),
but only the commit-scope contract lives there. Fail-fast is in tdd.md's
"Fail-Fast Rules" subsection (under "Gate Enforcement Rules") and error
handling is in tdd.md's "Error Handling" section — cite each correctly.

gsd-executor.md's "Plan-Level TDD Gate Enforcement" section also fully
restated the gate-sequence rules tdd.md's "Gate Enforcement Rules" already
owns (and covers more thoroughly, including the actual git-log validation
script). Collapse it to a short pointer, matching the treatment already
used by the cycle-steps pointer immediately above it.

Adds tests/tdd-reference-correctness.test.cjs asserting the pointer text
cites the correct section names, that those sections actually carry the
guidance, and that the old gate-sequence restatement is gone from
gsd-executor.md.

Closes #4267
Closes #4269

* docs(#4267): add changeset for tdd.md pointer correctness fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4267): acknowledge execute-plan.md growth

execute-plan.md grew 95 bytes (39766 -> 39861) from the corrected
three-section citation in the #3990/#4267 cycle-steps pointer.

Emitted-Drift-Ack-Growth: execute-plan.md — net +95 bytes from citing the "Fail-Fast Rules" and "Error Handling" sections by name instead of a single mis-scoped section (#4267).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4267): update PROSE_ALLOWLIST line numbers shifted by the pointer-citation edit

This branch's edits to agents/gsd-executor.md and gsd-core/workflows/execute-plan.md
shifted line numbers, leaving tests/no-bare-gsd-tools-command-position.test.cjs's
PROSE_ALLOWLIST pointing at stale lines. Update both entries to their new correct
lines (811 and 419 respectively) without changing the underlying prose.

* fix(#4267): restore INVALID_RED citation, fix allowlist line shift after #3770 rebase

The rebase onto next picked up #3770's already-merged fail-fast update to
gsd-executor.md's plan-level gate section, which this branch's own commit
collapses into a pointer. The conflict resolution kept the pointer but
dropped the literal "INVALID_RED" term that tests/tdd-red-evidence.test.cjs
requires gsd-executor.md to name — restored it. Also updates
no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST line number
for gsd-executor.md, shifted again by the rebase.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4267): fix non-matching regression-guard regex in tdd-reference-correctness

The fail-fast regression guard asserted gsd-executor.md no longer contains
"If a test passes unexpectedly during the RED phase" — but the actual old
prose (removed by this branch's pointer-collapse) read "If a test passes
unexpectedly during RED, STOP". The regex never matched the real old text,
so the assertion would have passed even against the unmodified pre-change
file. Caught by an isolated orthogonal review pass. Fixed to match the
actual removed wording, and confirmed (via a direct grep) it is genuinely
absent from the current file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4267): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 19:13:56 -04:00
Tom Boucher
2e056488d9 fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED)

Five-case matrix: decline on unsummarized plans (regression), fire on
fully-summarized phase, fail-open on unresolvable phase dir, idempotent
decline, normal advance untouched.

* fix(#4067): derive advance-plan phase-complete from disk, not the plan counter

The phase-complete branch of state.advance-plan was decided purely by
STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale
counter carried into a newly planned phase, or a counter raced by
wave-parallel executors, let 'Phase complete — ready for verification'
land while sibling plans were still executing.

cmdStateAdvancePlan now re-decides that branch from disk before the
write: every plan in the Current Position phase's directory must have a
SUMMARY.md (scanPhasePlans single owner, the same source
state.update-progress recalculates from). Outstanding plans decline the
entire write byte-identically (idempotent, concurrency-safe, counter
stays display-only); an unavailable disk answer fails open to the
counter-derived decision.

* fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs

#3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed
(current-milestone) lookup first so an archived milestone's stale dir
cannot shadow the live one; unscoped retry when the window cannot
answer. Also restore the transform's undefined-data error semantics and
extract scanOutstanding.

* chore(#4067): add changeset fragment

* chore(#4067): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-04 17:49:01 -04:00
Tom Boucher
d29b50d696 fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md

* fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding

* fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test

Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch.

* chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names

* fix(#4051): review fixes — em-dash description style, split audit-fix route

* chore(#4051): sync skill mirrors of execute-phase/phase descriptions

* chore(#4051): add changeset (pr backfill to follow)

* chore(#4051): backfill PR 4289 in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 17:19:16 -04:00
Tom Boucher
8249ebcf6e fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate

RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do
not exist yet; every row fails on require. Per #3770 only an intentional
target-test failure may authorize GREEN; zero-test discovery, fixture crashes,
unrelated failures, and unexpected green are INVALID_RED.

* fix(3770): require intentional RED evidence before GREEN

Only an intentional failure of the TARGET test (distinctly named, TAP-reported
assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test
discovery, fixture/load crashes (file-named failures), nonzero exits without a
failing test, unrelated failures, unexpected greens, and malformed/missing
records are INVALID_RED and block GREEN.

- src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses
  the prohibition-enforcement TAP primitives; fail-closed, never throws)
- check tdd-red-evidence <record.json>: validates the persisted record
  (command, exit code, failing test, expected, actual)
- gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now
  requires the evidence record + gate verdict, not a nonzero exit or a RED: tag

* chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs

* fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib

- gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B <
  49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid)
- tests: the row-6 fixture used String.replace (first-occurrence), so the
  `not ok` line still named the target test and the classifier was right to
  accept it; replaceAll makes the failure genuinely unrelated
- eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs
  (lint the src/*.cts source, per ADR-457 migration rule)

Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line

* chore(3770): add changeset

* chore(3770): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:44:38 -04:00
Tom Boucher
2f4f7538e9 fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) 2026-09-04 16:08:42 -04:00
Tom Boucher
580059251a fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing

Red: init.progress/init.resume/init.new-project payloads carry no
partial-init discriminator, and progress.md/resume-project.md/
new-project.md route an interrupted bootstrap (.planning/PROJECT.md +
config.json only) to Route F / STATE reconstruction / a hard error.

* fix(#4040): route partially-created .planning to initialization recovery

A bootstrap interrupted after .planning/PROJECT.md (but before
REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways:
progress.md read it as between-milestones (Route F) or 'no planning
structure', resume-project.md offered STATE.md reconstruction, and
new-project.md errored 'already initialized' — a routing loop with no
recovery exit.

Add a shared buildInitCompletenessFields discriminator
(planning_exists / requirements_exists / milestones_exists /
init_incomplete) to the init.progress, init.resume and init.new-project
payloads, and branch on init_incomplete in progress.md, resume-project.md
and new-project.md BEFORE the legacy branches. MILESTONES.md presence
excludes the archival between-milestones state, so Route F and the
STATE-reconstruction path keep working.

Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes.
Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch.
Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring.

* chore(#4040): add changeset fragment

* chore(#4040): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:03:42 -04:00
Xiangfang Chen
b1b7cabfb5 docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry

Adds one `type: "eos"` entry for a Qoder host integration and regenerates
docs/registries/eos-registry.md.

Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop).
The integration depends on @opengsd/gsd-core, negotiates the ADR-1239
host-integration handshake, and projects GSD's agents, skills, and hook
scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the
China edition), merging GSD's lifecycle hooks into settings.json.

Every axis is sourced from Qoder's own docs per the never-infer rule.
`dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a
frontmatter-declared, per-agent-definition property, and GSD's two
isolation negotiation models both assume a per-dispatch injection point
Qoder does not expose.

Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in
favor of the EoS path.

Closes #4123

* docs(#4123): backfill changeset pr field
2026-09-04 15:09:10 -04:00
Tom Boucher
75ee7b0214 enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate

One query verb computes TDD-applicability for a plan (CLI flag, plan
type: tdd frontmatter, a task's tdd="true" attribute, or the
workflow.tdd_mode config default), mirroring phase.mvp-mode's
precedence-cascade shape. Foundation for epic #4272 Phase 2, which
wires both dispatch backends to consume it instead of restating the
predicate independently.

Also fixes workflow.tdd_mode, workflow.research, and
workflow.nyquist_validation, which never reached
cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone
because loadConfig() never populates config.workflow — a dead
accessor found while wiring this verb's own config read, fixed inline
per the no-defer rule rather than left alongside it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4273): document phase.tdd-applicable's FEATURES.md entry

Add a docs/features/ fragment for the new phase.tdd-applicable query
verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left
untouched: it documents /gsd-* slash commands only, and the sibling
verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI
reference entry anywhere in docs/ either -- only inline prose mentions
in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md
precedent to extend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests

Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): stop whitelisting capability-owned config keys centrally

workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are
each already owned by their own first-party capability's federated config
schema (the tdd/research/nyquist capabilities declare them under their own
capability.json `config`), resolved via isCapabilityConfigKey. Adding them
to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as
the prior commit in this branch did (mirroring workflow.mvp_mode, which
genuinely is central-only), declares the same key in two places at once.
That collision breaks capability-loader.cts's loadRegistry composition:
gsd-test caught this as 84-85 unrelated failures across
capability-cli/capability-command-dispatch/capability-lifecycle test files,
every one showing "unknown capability: <id>" for a freshly-installed
third-party capability that should have resolved fine.

Verified directly (not asserted): reverting only this file, keeping the
config-loader.cts tdd_mode/research/nyquist_validation flattening and the
init.cts call-site fixes from the prior commit, and re-running the exact
capability install + capability set repro from
tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the
failure with the whitelist entries present and clears it without them.
loadConfig() still surfaces all three flattened values correctly with no
central whitelist entry (confirmed directly against the compiled module) —
the whitelist additions were never required for the #4273 fix to work; they
were an incorrect over-application of the mvp_mode precedent to keys that
aren't central.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file

Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4273): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 14:14:56 -04:00
Adnan
f4bf449296 fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat

cmdAuditUat admits `human_needed` OR `gaps_found`, but
parseVerificationItems had a body only for the first and returned an
empty array for the second — standing on a comment deferring to
`plan-phase --gaps`, a different command audit-uat never reaches. Since
cmdAuditUat pushes a file into `results` only when `items.length > 0`, a
`gaps_found` report did not under-report: it vanished, taking its
phase's `by_phase` row with it, so a clean-looking total gave the reader
no cue anything was skipped.

Eligibility now has one owner (the caller) and parseVerificationItems
reports what the file says.

The closed-entry filter could not be built on extractFrontmatter: its
array-item parser keeps only each `- ` entry's FIRST line and has no
notion of nested key/value objects, so an entry's `status:`/
`resolution:` siblings never reach its output and a closed entry is
indistinguishable from an open one downstream. Rather than grow a
competing object-list parser — or change extractFrontmatter, whose blast
radius is every frontmatter consumer in the repo — this reads the raw
segment BEFORE the flattening, via the existing anchored
sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps`
machinery that already parses exactly this `- `-opened, indentation-
continued shape.

The human_needed path is byte-for-byte unchanged: same reader, same
display names, same numbering, no resolved-entry filtering — pinned by
a test and verified by identical CLI output on base and head.
parseGapsItems keeps its narrower `status: resolved` rule so no
*-UAT.md behaviour moves.

Closes #3850

* chore(#3850): backfill changeset pr number for #3879

* fix(#3850): one parse per entry, one fence parser, one resolved-entry rule

Adversarial review on #3879: B1, B2, M3, m5, m8 and n9.

B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence
regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell
5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found`
report vanished from the audit exactly as it did before this fix — this issue's
own symptom, on a platform the repo already has a named defect class for.
`extractFrontmatter`'s BOM+fence logic is now factored out as
`frontmatterRegion` and shared. One fence parser, not two.

B2 — the resolved-entry skip paired two DIFFERENT parsers by array index:
`parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A
block sequence written at its key's indent — ordinary, legal YAML — makes them
disagree about entry count, and from the first disagreement every index names a
different entry, so an OPEN entry inherits a CLOSED one's resolution and is
silently dropped. That is the defect this PR exists to fix, reintroduced inside
the fix. Display name and sibling fields now come from ONE parse of the raw
slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as
`parseYamlRegion` does, so the string is byte-identical to what
`extractFrontmatter` produced. The flattened array remains the #2286 GATE, but
is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the
LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment.

M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both
readers use, rather than two copies differing only in `result`.

m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an
acceptance criterion #3850 does not contain: the issue has no AC section, and
its suggested fix (2) states the skip unconditionally, naming a file with 14 of
16 entries resolved. That file is `human_needed`, so the asymmetry left the
reporter's own scenario over-reporting by 14.

m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and
says the column-0 boundary rule is now a cross-module contract.

n9 — the vestigial bare block is gone and its body de-indented.

Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF
fixture (M4 — it survived by accident, now pinned) and the unified skip rule.
Fail-first verified by running the new tests against the pre-fix build: the BOM,
nested-sequence and unified-skip cases are red there.

* fix(#3850): read the entries as objects, not as re-parsed display text

Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1
(#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml:
`parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry
now flattens to `test: A, resolution: R` rather than to its first line.

The original mechanism existed ONLY to work around that lossy first-line
flattening — it sliced the raw frontmatter segment and re-parsed each entry by
hand so a `resolution:` sibling was visible at all. With a real parser upstream
that workaround is obsolete, so it is deleted rather than repaired:
`sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the
`splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are
all gone.

`frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` —
the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence,
same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping
one step before the display flattening. `flattenObjectListItem` is exposed
alongside it so a caller deriving a display name produces the byte-identical
string `extractFrontmatter` would have.

That collapses the review's blockers into properties of the parse rather than
things this fix has to get right:

- B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI.
- B2 (index pairing) — there is no second reader. Display name and sibling
  fields come from one object.
- M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers.
- M4 (CRLF) — js-yaml's, not ours; verified through the CLI.

Also confirmed on the rebased base, per review: #3850 still reproduces on `next`
after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found`
fixture), so this PR is still doing work #3707 did not do. Nothing was dropped
as redundant.

One behaviour note: `entryField` returns a present value verbatim and treats
only whitespace-only as absent. Trimming would rewrite an author's `truth:` on
its way to becoming the display name.

* fix(#3850): keep every frontmatter list entry at its own row

Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result
to objects, and filtering COMPACTS: `parseHumanVerificationItems` then
numbered the survivors by their position in the compacted array. On a list
mixing object and non-object entries the non-object rows disappeared outright
and the rest were renumbered — #3850's own vanishing-row defect, reached
through entry SHAPE instead of file STATUS. Base never had it: it walked the
display array, so every row surfaced at its own position.

Renamed to `frontmatterListEntries` and it no longer filters (the name now
matches what it returns). Deciding what a non-object entry MEANS is a
caller's judgement; dropping it is nobody's.

Both readers now walk the DISPLAY array — one element per row, the array
#2286 already gates on — and consult the parsed array only for "does this
entry carry a closure field?". `parsedEntriesFor` owns that pairing and
checks the two lengths agree before trusting an index; all-null is the
correct degradation, since over-reporting a closed row is recoverable and
closing the wrong one is not. Names stay byte-identical to base for every
entry shape, including a nested sequence (`[nested]`, not `["nested"]`).

Same class closed in the gaps reader: a non-object `gaps:` entry surfaced
nothing at all and now surfaces as `unknown`, which is this module's
documented fail-safe direction (`parseGapsItems`) on a false-negative bug.

Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase
dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule
inlined twice; `extractFrontmatter` now routes through it, so "one fence
parser" is enforced rather than asserted in a comment.

Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and
its "shared by both readers" comment corrected — it has one call site, and
the two readers differ deliberately, each mirroring its own established
sibling (`parseGapsItems` vs #2286). Documented at the divergence.

Tests: `B2` asserted a name substring, so it passed while the row was
mis-numbered and would have passed through outright loss; it now asserts
positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim,
B2c the survivors' file positions across skipped rows, B2d the gaps reader.
All four fail-first against the reviewed head; 332/332 green with the fix.

* fix(#3850): make status authoritative, and let the two gaps readers agree

Round 4 review, all five findings.

Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as
closure regardless of `status:`, so `status: failed` + `resolution:
"attempted retry, still failing"` vanished from the report — the
silently-vanishing-item defect #3850 exists to close, reached by field
combination instead of file status.

Closure is now per key, because the two keys have different conventions
and one rule cannot serve both:

  `gaps:`               `status: resolved` only, byte-identical to the
                        rule `parseGapsItems` applies to a `## Gaps`
                        markdown section, so one authored entry cannot
                        read closed in one reader and open in the other.
  `human_verification:` a bare `resolution:` still closes, since that is
                        how verifier-written entries record it — but a
                        readable `status:` that contradicts it wins.

A single unified rule was the first draft and is wrong: it closes a
frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which
`parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring
claims it mirrors that reader's fail-safe status handling.

The contradiction guard is not a judgment call about YAML. It is the rule
this codebase already applies to the same field pair: `validateResolution`
(probe-core.cts) rejects a populated `resolution:` on a non-resolved status
outright — "a populated payload is an authoring mistake ... Reject it so
the mistake surfaces." A reporter cannot throw, so it surfaces the item.

Minor 1. Direct unit tests for `frontmatterListEntries` and
`flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that
historically co-changes with `frontmatter.cts`. They were reachable only
through `uat.cts`' readers before.

Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted
directly. Verified unreachable through content rather than assumed: both
readers enter through `frontmatterRegion`, `extractFrontmatter`'s only
extra argument gates a warning, and `normalizeParsedValue`'s `value.map`
is 1:1. It is a drift alarm for a future edit to either parser, so the
helper is exported for tests rather than left as the one unpinned branch.

Minor 3. The vestigial `const skipResolved = true` and its dead
conditional are gone.

Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:`
entry has no `test:` in its vocabulary — the template's entries carry
truth/status/reason/artifacts/missing — so it was speculative support for
a field the shape does not have, and it collided with the 1..N row numbers
`parseHumanVerificationItems` assigns by array position. Not reading it
makes the collision impossible; an offset would have rewritten an authored
value, against `entryField`'s verbatim contract.

Docs, changeset and the dispatcher docstring all stated the unconditional
rule and are corrected — three prior rounds here were comment/code drift.

Fail-first proven: restoring the universal rule reddens all three new unit
tests and both rewritten properties.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 17:47:01 +00:00
Tom Boucher
585a8b7f1b fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity

* fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path

* fix(#3747): scope doc evidence claim to skills discovery per adversarial review

* chore(#3747): add changeset

* chore(#3747): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 11:48:38 -04:00
Dennis Alexis Valin Dittrich
75bad7aedd test(#3936): tighten quick researcher regression coverage (#4169)
* test(#3936): tighten quick researcher regression coverage

* test(#3936): restore adjacent quick dispatch coverage

Assert the default researcher model and bind the executor persona check to its Agent payload.

* test(#3936): make parse-list assertion wrap-safe

Bound the researcher-model check to the full parse paragraph so formatting-only line wraps do not fail the regression test.

* test(#3936): isolate Windows model defaults

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 10:22:10 -04:00
Tom Boucher
7eefad6f92 docs(#4260): record npm-audit retry/backoff enhancement as out-of-scope (#4271)
Co-authored-by: sim <sim@local>
2026-09-04 10:20:38 -04:00
Tom Boucher
18e5cfff8a fix(#4250, #4260): distinguish a timed-out npm audit from a JSON parse failure, retry with backoff (#4251)
* fix(#4250): distinguish a timed-out npm audit from a JSON parse failure

npm-audit-baseline.cjs's runPackageLockAudit, and the near-identical
auditProductionVulns helper in npm-integrity-gate.test.cjs, both grabbed
e.stdout whenever an npm audit child process exited non-zero -- without
checking whether the process was actually killed by its 180s timeout.
A timeout-killed process's stdout is truncated mid-write, not complete
JSON, so JSON.parse threw a misleading "Unexpected end of JSON input"
instead of naming npm's registry timeout as the real cause.

Root-caused live during a CI investigation: npm's own status page
reported degraded service, and the registry's bulk-advisories endpoint
was returning 503/hanging, causing npm audit to sit until the timeout
fired.

Adds a shared isTimeoutKill(error) predicate (checks execFileSync's
documented killed/signal fields) and checks it first in both catch
blocks, throwing a clear, actionable error before ever reaching
JSON.parse. The pre-existing "non-zero exit with complete JSON"
recovery path is unchanged and still covered by regression tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4250): add changeset for npm-audit timeout fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4250): share the timeout-kill error message and cover auditProductionVulns

Two independent review passes (standards + spec) on the first commit found
real gaps: the timeout-kill error message was duplicated verbatim between
runPackageLockAudit and the near-identical auditProductionVulns helper in
tests/npm-integrity-gate.test.cjs (this repo's own Generative Fix Divergence
anti-pattern -- shared logic across parallel surfaces with no parity check),
and auditProductionVulns picked up the same production fix with zero test
coverage of its own.

Extracts buildTimeoutKillError(cwd), used by both callers so the message
cannot independently drift. Gives auditProductionVulns the same injectable
execFileSyncImpl seam runPackageLockAudit already had, and adds the matching
regression tests (timeout-kill throws the clear error; the pre-existing
non-zero-exit-with-complete-JSON path still recovers correctly).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4250): backfill changeset PR number to #4251

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* diag(#4250): surface captured stderr in the timeout-kill error

The killed child process's stderr is buffered in-memory by execFileSync
and attached to the thrown error, but nothing surfaced it -- the timeout
message named the timeout but discarded the one piece of data that could
show WHY npm was still running when it fired (DNS stall, TLS handshake
stall, a registry-side retry loop, all look identical without it).

buildTimeoutKillError now takes the killed error and includes its stderr
(or an explicit 'no stderr was captured' note) in the message. This is a
diagnostic improvement for the next CI occurrence, not a behavior fix --
local reproduction has directly ruled out npm version (installed the
exact CI-bundled 11.17.0 and ran it against this repo: 0.49s, clean),
general npm registry reachability (0.4-1.4s locally, repeatedly), and
npm ci speed (2m, succeeded) as explanations for the 180s CI hangs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4260): bounded retry with backoff for npm audit calls, finish the extraction

The audit backend has real, independent latency variance from the rest of
the npm registry -- measured (see #4260): a bulk-advisories POST took
43.41s vs 0.20s for a plain registry fetch on the same host, and the same
endpoint returned no response at all (000) twice in the same window,
while status.npmjs.org reported fully operational throughout. Against
that, runPackageLockAudit and its near-duplicate auditProductionVulns
each made exactly one attempt with no retry -- any single bad moment
failed a REQUIRED CI gate on a transport hiccup, not a real advisory.

Replaces the single 180s attempt with runNpmAuditWithRetry: up to 3
attempts at 60s each (comfortably above the worst measured working
latency) with exponential backoff between them. Only a confirmed
timeout-kill is retried; a genuine non-timeout failure still fails
immediately, and exhausting all attempts still fails the gate -- per
#4260's own caveat, silently disarming a required security check on a
transport error is worse than occasionally re-running CI.

Also finishes the extraction #4260 flagged as stopped halfway:
auditProductionVulns (tests/npm-integrity-gate.test.cjs) duplicated
runPackageLockAudit's entire candidate loop, recovery branch, and timeout
classification, differing only in npm args and precondition check. It is
now a thin wrapper delegating to the newly-exported runInstalledTreeAudit,
which shares runNpmAuditWithRetry with runPackageLockAudit -- one
implementation instead of two that could independently drift.

buildTimeoutKillError now reports attempt count and still surfaces
captured stderr from the last kill.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4260): update changeset for retry/backoff scope

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4260): budget for two sequential retry-audit calls, close coverage gaps

Two review passes on the retry/backoff commit found real gaps:

- TEST_TIMEOUT_MS budgeted only one retry-audit call's worst case (210s),
  but checkTreeAgainstBaseline makes two sequential calls (HEAD tree via
  auditProductionVulns, baseline tree via runPackageLockAudit) -- combined
  worst case is ~372s. If both genuinely exhausted retries, node:test's
  own timeout would fire first and mask buildTimeoutKillError's clear
  message, undercutting #4250's own fix in that edge case. Recomputed
  using the same backoff formula the production code uses, so it can't
  independently drift.

- buildTimeoutKillError's default-attempts(1) singular-phrasing branch had
  zero direct test coverage (nothing calls it with a single attempt
  anymore) -- a real mutation-testing risk. Added direct tests for both
  phrasing branches plus the no-error-object case.

- runInstalledTreeAudit's null-guard skip paths (missing package.json,
  missing node_modules) had no tests, unlike runPackageLockAudit's
  matching paths. Added for parity.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 10:02:17 -04:00
Tom Boucher
97ce61dee2 fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally

* fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds

Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD

* chore(#3990): changeset for the single-statement TDD cycle

* chore(#3990): backfill changeset pr number

* fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap

* test(#3990): allowlist pin tracks the rebased line

* fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError

* fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError

Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with
empty stdout; the empty string survived the recovery path and surfaced as
'SyntaxError: Unexpected end of JSON input', hiding the captured error. The
recovery path now requires non-empty stdout, and an empty result throws with
the captured stdout/stderr/message so the actual error is on the record.
Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only
stops masking it.

---------

Co-authored-by: sim <sim@local>
2026-09-03 21:41:14 -04:00
Rezolv
590edec7a7 fix(#3956): require positive evidence for verify artifacts/key-links pass (#4004)
* fix(#3956): require positive evidence for verify artifacts/key-links pass

An all-string or path-less must_haves.artifacts / key_links block is
item-by-item skipped, leaving zero checked results, yet the pass verdict
was computed as `passed === results.length` (0 === 0), so all_passed /
all_verified read true with status valid and exit 0: a silent false GREEN
over zero acceptance evidence.

Add a positive-evidence floor (results.length > 0) to both verdicts,
mirroring the no-vacuous-pass rule at src/uat-predicate.cts. A well-formed
block, the fully-empty-block error, the parser's string tolerance, and
key-links pending (#1202) semantics are all unchanged.

Governing: ADR-3473 section 8 / 37C (absence, emptiness and failure must
not encode as success) and Decision 3 (failure is a value).

* chore(#3956): add changeset for verify vacuous-pass fix

* test(#3956): add mixed-block coverage and correct the key-links vacuous-pass comment

Addresses review on #4004:
- Correct the cmdVerifyKeyLinks positive-evidence-floor comment: only bare-string
  items are continue-skipped; a from:-less object is NOT skipped (it falls through
  to a verified:false hard failure), so it was never part of the vacuous-pass
  surface. The prior comment overclaimed symmetry with the artifacts side.
- Add a mixed-block regression test per verb (one bare-string prose bullet + one
  well-formed entry): the string is skipped, results.length === 1 > 0, and the
  verdict follows the single real entry — pinning that the floor does not
  over-reject a partial block.
- Tighten the changeset wording to match (all-bare-string, not "no path:/from: key").

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-03 21:46:50 +00:00
Rezolv
a788afb120 fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives (#4021)
* fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives

stateReplaceField's bold and plain patterns used `\s*` for the label-to-value
gap, which matches the newline after an empty field; `(.*)` then captured the
following line and the rebuild discarded it -- silent STATE.md data loss on any
`state update` against an empty body field (Status:, Stopped at:, Paused at:),
with exit 0 and no warning.

Confine the gap to same-line whitespace (`[ \t]*`), mirroring the already-correct
read side (stateExtractField, src/state-document.cts:404/:409), and pin the
label-value separator to a single space when the label line had none, so an empty
field yields `**Status:** value` rather than a glued `**Status:**value`. Non-empty
and pipe-table replacements are byte-identical to prior behaviour.

ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; this aligns
the writer to it. Regression test fails before / passes after and covers bold and
plain shapes, LF and CRLF, the non-empty byte-identity guard, and an end-to-end
transitionCore characterization at the consumer (ADR-3180 Decision 4(c)).

* chore(#4010): add changeset for the stateReplaceField empty-field fix

* test(#4010): add boundary and property coverage; scope the changeset's unchanged claim

Addresses review on #4021:
- Add boundary tests for the shapes the example tests missed: an empty field at
  end-of-document (no following line, bold + plain), two consecutive empty fields
  (only the target is filled, the other empty field's line survives), and an empty
  new value on an empty field (joinFieldReplacement synthesizes no dangling
  separator and the following line is preserved).
- Add a fast-check property over the bold/plain branches and joinFieldReplacement:
  for any field name, any values (empty fields included), and any new value,
  replacing one field changes only its own line and never the total line count —
  the invariant #4010 violated, now guarded directly.
- Scope the changeset's "unchanged" claim to ordinary space/tab separators (an
  exotic vertical-tab/form-feed separator, which no GSD template emits, now
  normalises to a single space).

* test(#4010): pin glued-separator non-empty field, scope joinFieldReplacement JSDoc

Round-3 review carried forward a Minor finding: joinFieldReplacement's JSDoc
still claimed non-empty replacements are unconditionally "byte-identical to
prior behaviour", but a non-empty field written with no label-to-value
separator (**Status:**value) gains a single inserted space under the narrowed
[ \t]* gap. Round 2 scoped only the changeset prose; the source JSDoc was left
making the false unconditional claim.

- Scope the JSDoc's byte-identity claim to ordinary space/tab separators and
  name the no-separator normalization as the one intentional exception.
- Add a test pinning the glued-separator case (**Status:**Planning): exactly
  one space inserted, following line survives, not byte-identical.

Emitted .cjs is gitignored (class-1), so no emitted-drift-ack applies.
build:lib clean; 74/74 state-document tests pass.

Claude-Session: https://claude.ai/code/session_01Mzmut6aeqZ1APfUBAkBZTR

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-03 17:14:12 -04:00
Tom Boucher
02c6955162 chore(#4241): add merge_group trigger to test.yml (#4242)
* chore(ci): add merge_group trigger to test.yml (#4241)

GitHub's merge queue fires the `merge_group` event for the temporary
merge-group commit it creates when a PR is added to the queue - not
`pull_request` or `push`. Without this trigger, `required-tests` (the
registered "Required tests" branch-protection check) never schedules
for a queued PR, permanently stalling the queue on a check that never
runs. This is workflow-side prerequisite wiring only; enabling the
merge queue itself is a separate manual branch-protection step.

* fix(#4241): pin AUDIT_BASELINE_REF for merge_group events too

Code review on this branch caught that AUDIT_BASELINE_REF's ternary
only branched on pull_request/push, so a merge_group run silently fell
through to '' -- scripts/npm-audit-baseline.cjs's resolveBaselineRef()
documents its origin/next live-tip fallback as unreachable from CI
specifically because AUDIT_BASELINE_REF is "always set by test.yml".
Reopens the exact race #4196 fixed, but only for merge-queue runs.

Extends all three AUDIT_BASELINE_REF pins (test, test-inert, test-full)
to also branch on merge_group, using github.event.merge_group.base_sha
(confirmed against GitHub's own webhook payload schema: "the SHA of the
merge group's parent commit") -- the base tip the temporary
merge-group commit was built against.

Adds a regression test asserting every AUDIT_BASELINE_REF pin branches
on merge_group with the correct field.

---------

Co-authored-by: sim <sim@local>
2026-09-03 15:00:04 -04:00
Tom Boucher
1fe85cd43e chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk

Repo-wide sweep (ahead of adding lint rules for these exact bug classes)
found both incident patterns still live and unfixed on `next`:

- scripts/run-tests.cjs's sweepProtectSet walk stopped on
  `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel.
  win32 dirname('D:\') is a fixed point (length 3, never satisfies
  `> 1`... wait, it does satisfy length>1), so a selected file living
  outside runTempRoot (the common case) spins the walk forever on
  Windows. Extracted a pure, exported computeSweepProtectSet helper
  that terminates on dirname(cur) === cur instead, with in-process
  RuleTester-style coverage for both win32 and posix paths.

- tests/run-tests-temp-root.test.cjs's own #4020 regression test set
  only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never
  reads TMPDIR on Windows (only TEMP, then TMP), so the redirect
  silently no-oped there — masked because Windows CI died in the
  dirname-walk hang above before ever reaching this test.

- tests/config-schema.property.test.cjs's fallow config-set test had
  the same TMPDIR-only pattern, direct process.env assignment this
  time, restored in its own finally block.

Origin: #4220 and its shared root cause #4020.

* feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules

Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug
class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY
catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for
either shape.

- local/require-full-tmpdir-triad: flags a TMPDIR environment override
  (direct process.env.TMPDIR assignment, or a TMPDIR property in a
  spawn-like call's env: object literal) not accompanied by TEMP and TMP
  in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows.
  Registered on tests/**/*.cjs, matching the require-userprofile-with-home
  precedent.

- local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning
  from dirname() with no fixed-point termination guard
  (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a
  no-op at the platform root, but the value differs by platform
  (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped
  length/equality bound never fires on Windows. Registered on BOTH
  tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in
  scripts/run-tests.cjs, not tests/.

Both rules join the zero-escape-hatch discipline already established for
this catalog (no bespoke comment marker; PROTECTED_RULES in
tests/portability-rule-disable-ban.test.cjs independently bans
eslint-disable of either). ADR-1703 and its two companion contributing
docs get an amendment documenting the mechanism, code examples, and the
repo-wide sweep (three live instances found and fixed in the prior
commit; no others found). CI test-scope selection updated so an edit to
either rule or to scripts/run-tests.cjs re-runs the right suites.

* fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too

checkWhile bailed out early unless node.test was a LogicalExpression,
so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } --
was silently skipped and never reported. That is the EXACT minimal
shape of the original #4020/#4220 bug, and it is literally the shape
used by this rule's own shipped RuleTester fixtures (the "equality-only
bound" invalid cases), which were failing (0 errors reported, 1
expected) until this fix -- confirmed by running RuleTester directly
against both fixtures, not just via a passing test-runner exit code.

The conjunct-collection helper already handled a non-LogicalExpression
test correctly (it pushes a single node as the sole conjunct); only the
early-return gate needed to stop requiring a compound && / || test.

Verified: RuleTester run directly against both previously-broken
fixtures plus two new sanity cases (a guarded single-condition loop
stays valid; an unrelated single-condition loop stays silent), and a
fresh `npx eslint .` across the whole repo remains clean (no other
single-condition dirname-walk shape exists in the tree).

* fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call

isSpawnLikeCallee only recognized a MemberExpression callee
(child_process.spawnSync(...)) or a bare identifier in
ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare --
const { spawnSync } = require('child_process'); spawnSync(...) -- has an
Identifier callee named "spawnSync", which matched neither branch, so
the whole env-literal check was skipped. gsd-test caught this: both
"invalid: child_process.spawnSync with TMPDIR-only env" cases in
tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors
reported, 1 expected).

Widened the bare-identifier branch to also match any of the known
ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same
lightweight convention this repo's other eslint-rules/*.cjs use (e.g.
no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow
tracing.

Verified: RuleTester run directly against all 11 cases in
tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that
were failing), all pass; a fresh npx eslint . and npm run lint:ci
across the whole repo remain clean.

* fix(#4244): correct a stale escape-hatch reference in a test comment

The comment on the "length comparison against another expression's
length" case referenced a "// allow-dirname-walk marker" that doesn't
exist -- the rule has zero comment-based escape hatches by design
(ADR-1703), and an earlier draft's marker mechanism was removed before
this branch's first commit. Spec-axis review caught the stale
reference. No behavior change; comment-only.

* chore(#4244): backfill changeset PR number (pr:0 -> pr:4246)

---------

Co-authored-by: sim <sim@local>
2026-09-03 14:14:09 -04:00
Tom Boucher
456136659d fix(#4220): terminate the Windows temp-sweep ancestor walk (and two bugs it unmasked) (#4245)
* fix(#4220): terminate the temp-sweep ancestor walk with a fixed-point check

scripts/run-tests.cjs's sweepProtectSet block walked each selected test
file's ancestor directories, stopping on `cur !== runTempRoot &&
cur.length > 1` — a POSIX-only sentinel. path.posix.dirname('/') === '/'
(length 1) correctly stops, but path.win32.dirname('C:\\') === 'C:\\'
(length 3) never satisfies the length check, so the walk spun forever on
Windows whenever a selected file lived outside runTempRoot (the common
case). This has hung every Windows CI shard since #4207.

Extract the walk into a pure, exported computeSweepProtectSet(selected,
runTempRoot, dirnameImpl) helper and replace the length sentinel with a
fixed-point check (stop when dirnameImpl(cur) === cur), which terminates
correctly on POSIX, Windows drive roots, and UNC roots alike with no
platform branch.

* fix(#4220): repoint TEMP/TMP alongside TMPDIR in run-tests-temp-root test child env

Node's os.tmpdir() on Windows never reads TMPDIR, only TEMP/TMP. The
test's runNode child-process env override only set TMPDIR, so on a
real Windows runner nested inside a run-tests invocation the child
inherited the outer process's already-repointed TEMP/TMP and its
mkdtempSync(os.tmpdir()) landed under the outer run's temp root
instead of the test's intended `outer` directory. This was masked on
gsd-test's benches and locally because Windows CI always died in the
#4220 infinite loop before reaching this test.

* fix(#4220): stop the ancestor walk from protecting the filesystem root itself

computeSweepProtectSet added `cur` to the protect set before checking
whether dirname(cur) === cur, so on the terminating iteration it
protected the filesystem root (posix `/`, and analogously a win32
drive root) instead of stopping before adding it. Caught by the
existing posix-parity regression assertion
(`!protectSet.has('/')`) on the linux-node24 gsd-test bench. Reorder
to compute the parent and check the fixed point before adding.

* fix(#4220): backfill changeset pr number to 4245

---------

Co-authored-by: sim <sim@local>
2026-09-03 13:42:14 -04:00
Tom Boucher
515191f07d feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only)

* test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED)

Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/
merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and
confirms the prior research pass's Open Question 1: a coordinator crash
between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge)
leaves BATCH.json at "pending" with no STATE.md row yet (only written in
Step 9), so --resume's eligibility re-derivation would dispatch a second
executor into a new worktree for the same item, orphaning the first.

This test asserts worktree-dispatch.md's Step 6 excludes an item whose
SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's
existing PLAN.md-existence check one layer earlier. Fails against the
current worktree-dispatch.md, which has no such guard.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the full trace and fix-location rationale.

* fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN)

worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round
via the same quick-batch resume call resume-mode.md uses, but had no check
for "did this item already finish executing" the way planner-wave.md
already checks "did this item already get planned" (PLAN.md existence)
before re-planning. A coordinator crash between Step 6 (executor commits,
SUMMARY.md written) and Step 7 (merge) left the item eligible for a second
dispatch on --resume, orphaning the first worktree's real, already-
committed work and silently losing it once the second executor's SUMMARY.md
write clobbered the first at the same item_dir path.

Adds a SUMMARY.md-existence exclusion before spawn-plan is computed,
symmetric to planner-wave.md's PLAN.md check. The excluded item is not
lost: merge-wave.md's own mergeable-wave criterion (status=pending,
SUMMARY.md on disk, not yet merged) already picks it up independently of
this eligible/spawn list.

Workflow-prose-only fix — touches no already-merged/reviewed .cts module.
See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the fix-location rationale (why not resumeBatch itself).

* test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules

Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677,
epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership
attempts", "scope drift", "submodules"):

- Arbitrary-worktree ownership tampering: a manifest entry naming a
  non-agent branch is silently dropped at normalization before any git
  subprocess runs; a manifest entry naming a plausible agent-branch that
  was never actually created by this repo's own worktree.create (a
  genuinely foreign repo/branch) is blocked via base_mismatch. Both leave
  the foreign location and repoRoot's HEAD provably untouched.

- Advisory scope drift: a committed path outside declared files_modified
  still merges successfully (advisory, never blocking) while surfacing a
  scope_out_of_declared warning naming the drifted path; an exact
  declared-scope match produces zero warnings (boundary case).

- Real .gitmodules submodule integration: a repo containing a real local
  git submodule merges cleanly through executeWorktreeWaveCleanupPlan for
  an unrelated plan; a real gitlink pointer bump (declared) merges cleanly
  with the superproject tree reflecting the new pinned commit; an
  undeclared bump is advisory-only and surfaces a scope warning naming
  vendor/sub, same as any other undeclared modification.

No src/*.cts changes — all three gaps were coverage-only; the underlying
primitives already behaved correctly (independently verified against real
git subprocess output before writing each assertion).

* docs(#3677): document how to diagnose a preserved quick-batch worktree

Extends the one-sentence "worktree is preserved (never deleted)" mention
into a concrete diagnosis procedure: where the preserved directory is, how
to read the executor's real commits/diff against the plan's declared
files_modified, how to read the item's own SUMMARY.md independent of merge
outcome, how to manually merge-and-clean-up or discard, and how to re-run
--resume afterward. Also documents that a SUMMARY.md-written-but-still-
pending item (the crash-window case fixed in this same PR) needs no manual
intervention — --resume routes it straight to the merge step.

* chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only)

* fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable

Orthogonal review (Spec finding): the crash-window regression test added
earlier this phase only asserted readStep('worktree-dispatch.md') + regex
matches against the markdown prose — proving the DOCUMENTATION says the
right thing, never that the runtime condition (pending status + on-disk
SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own
"Alternatives considered" explicitly rejects "document recovery without
fault injection" for exactly this reason.

Extracts the filtering decision into a pure, independently testable
function, filterAlreadyExecuted(eligibleIds, executedIds) in
src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed`
CLI verb (src/quick-batch-command-router.cts) — the same pure-decision-
then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already
establish. worktree-dispatch.md now calls this verb explicitly instead of
only describing the decision in prose. A genuine fixture-based test in
tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch),
writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the
REAL resumeBatch, and proves both that resumeBatch alone still reports the
item eligible AND that filterAlreadyExecuted (fed a real filesystem check)
correctly excludes it. The prior prose-assertion tests are kept — they now
prove the workflow markdown is correctly WIRED to the verb — but are no
longer the only proof.

Self-discovered defect while building that fixture (fixed inline, not
deferred): tracing merge-wave.md against /gsd:quick's own prior art
(QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed
$QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed
coordinator correctly does not re-dispatch an already-executed item (this
fix), but nothing durably recorded that item's worktree_path/branch/base
either — Step 7 in the resumed process would have had no data to build its
cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/
dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT
a reuse of the pre-existing `worktree` field, whose loadBatch validation
requires the path to exist on disk (verified empirically: reusing it made
the batch permanently unloadable the moment a legitimately-merged worktree
was removed). worktree-dispatch.md persists the triple once a worktree is
created; merge-wave.md falls back to it when the ephemeral manifest lacks
an entry, clears it after a successful merge, and fails closed rather than
guessing if no record exists anywhere.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1
and §9.3 for the full trace, empirical verification notes, and rejected
alternatives (reusing `worktree` directly).

* test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees

Orthogonal review (Security finding): the two existing ownership-tampering
tests didn't test ownership — one was trivially rejected by
WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch-
NAME filtering, not ownership), the other pointed at a wholly separate,
never-linked foreign repo, so merge-base failed immediately because the
branch didn't exist as a ref at all. Neither exercised the real scenario:
a manifest entry whose worktree_path/branch are swapped to point at a
DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot,
with a branch name passing the shape check and a base in allowed_bases.

Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts)
directly: this is NOT a reachable gap. Git enforces branch-per-worktree
uniqueness, so a swapped-in entry.branch can only match worktree_path's
ACTUAL checked-out branch if it names that sibling's own real, uniquely-
generated branch name — which manifest tampering confined to one batch's
own record has no way to know (branch names are
agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is
collision-checked GLOBALLY across every existing quick task and batch, not
merely within one batch).

Adds a stronger test that empirically proves this: two REAL, concurrently-
alive sibling worktrees of the same repo (both via real `git worktree add`,
both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with
worktree_path/branch swapped between them in both directions. Both attempts
are blocked via branch_mismatch; both real worktrees, their branches, and
one sibling's real uncommitted-to-main commit survive completely untouched.
Supplements (does not replace) the original two tests, which still prove
distinct, real boundaries.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2
for the full trace, including the one explicitly-documented (not fixed)
trust boundary this investigation surfaced: the primitive defends against
fabricated data, not a caller bug that misattributes a real-but-wrong
item's own triple to a different item.

* chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only)

* docs(#3677): add changeset for PR 4240

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 09:47:22 -04:00
Tom Boucher
8f013983e5 fix(#3751): provision the claude CLI in CI and cover agents/ in the plugin-validate fixture (#4229)
* test(#3751): the validation fixture must cover agents/ and CI must provision the CLI

* fix(#3751): cover agents/ in the plugin-validate fixture and provision the claude CLI in CI

* fix(#3751): wire the strict live-config guard into the plugin-validate job

* fix(#3751): job-level strict-guard env, where the guard derivation reads it

* chore(#3751): changeset for the CI-provisioned plugin-validate gate

* chore(#3751): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-03 06:59:30 -04:00
Tom Boucher
3639ab0431 fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated

* fix(#3968): measure commit claims at all three surfaces

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)
Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument
Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain

* fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin

* fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)
Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument

* fix(#3968): hold the gsd-executor size cap with a compact ledger contract

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)

* fix(#3968): allowlist pin and HALT regex track the final prose

* chore(#3968): changeset for measured commit claims

* chore(#3968): backfill changeset pr number

* fix(#3968): quote the BASE expansion (SC2086)

* ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap)

---------

Co-authored-by: sim <sim@local>
2026-09-03 05:43:16 -04:00
Tom Boucher
d5f8191f66 fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical

* fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick

* fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter

Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run)
Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction

* chore(#3730): backfill changeset pr number

* fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch

---------

Co-authored-by: sim <sim@local>
2026-09-03 01:32:41 -04:00
Tom Boucher
114dfcb739 fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones (#4214)
* fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones

The #3588 gate failed on ANY advisory in the production tree, regardless
of whether the PR/push actually introduced it. Because npm's advisory
database updates continuously and independently of repo state, a commit
could pass this gate at merge time and fail it minutes later on the
identical tree -- proven on PR #4188/dce40eeb6, which passed on all 3
OSes at 15:57-16:27 and failed the same assertion at 16:19-16:30 on the
unchanged commit, purely because GHSA-jqff-g426-hqxp was disclosed for
fast-uri in the interim.

scripts/npm-audit-baseline.cjs diffs the head tree's vulnerable-package
set against a resolved baseline (the PR's target branch, or the prior
commit on a direct push) and blocks only newly-introduced advisories.
When no baseline can be resolved, falls back to the original
zero-tolerance behavior -- fail-closed, never silently weaker.

* fix(#4196): pin the npm-audit baseline instead of using a drift-prone ref

Two orthogonal reviews found the same class of bug this repo already
fixed once for a different gate (see GSD_EMITTED_BASE's own incident
comment in test.yml): origin/<branch> is live under fetch-depth: 0 and
can advance mid-run, so resolveBaselineRef()'s fallback to
origin/${GITHUB_BASE_REF} could silently disagree with the tree
ci-rebase-check.cjs actually merged. Wire AUDIT_BASELINE_REF from the
workflow to github.event.pull_request.base.sha / github.event.before,
the same pinned values GSD_EMITTED_BASE already relies on.

Also: HEAD~1 assumed exactly one commit per push, which this repo's
allow_rebase_merge:true setting can violate (a rebase-merged PR lands
as several discrete commits in one push) -- github.event.before is
git's own record of the correct pre-push state, not an assumed offset.
HEAD~1 remains as a documented last-resort fallback for out-of-band
invocations (e.g. gsd-test) that don't set any of the above, alongside
a new local-branch fallback for gsd-test's local `next` (not
origin/next) sandbox shape.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:37 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
7d6d788b51 fix(#4020): bound the test run's temp footprint with a swept run-scoped root (#4207)
* test(#4020): the runner must bound and sweep a run-scoped temp root

* fix(#4020): bound the run's temp footprint with a swept run-scoped root

* test(#4020): isolate the env-mutating rows in child processes

* fix(#4020): gate root removal on ownership so nested runners spare the outer root

* test(#4020): pass the probe file via --files, the runner's explicit-file flag

* test(#4020): resolve the probe by basename, as --files matching requires

* test(#4020): assert root survival, not content survival, in the nested-row

* chore(#4020): changeset for the run-scoped temp root

* chore(#4020): backfill changeset pr number

* fix(#4020): the sweep spares ancestors of the runner's own selected files

* fix(#4020): TMPDIR precedence — an operator redirect beats inherited TEMP/TMP

* fix(#4020): only the root's owner sweeps — a nested runner spares live sibling fixtures

---------

Co-authored-by: sim <sim@local>
2026-09-02 20:47:34 -04:00
Tom Boucher
858bb89769 ci(#4196): exempt dependabot[bot] from issue-link, title, and unsolicited-PR gates (#4203)
Dependabot has no mechanism to link a PR it opens to a repo issue -- its
alerts live in the Security tab, not as issues -- so require-issue-link,
pr-title-validator, and auto-close-unsolicited-prs all rejected its PRs
by design (confirmed live on #4193: auto-closed for "no pre-approved
issue", then flagged again by the title gate on reopen). Exempt by
authenticated author login (github.event.pull_request.user.login /
context.payload.pull_request.user.login), which GitHub attributes and a
crafted title or branch name cannot forge -- scoped narrowly to
dependabot[bot] only, no other author gets this treatment.

Co-authored-by: sim <sim@local>
2026-09-02 15:56:43 -04:00
Tom Boucher
91ed46882a feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives

Adds the full behavioral (tests/quick-batch.test.cjs) and property-based
(tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core
primitives per the phase's 35-row test matrix — task-list parsing (inline +
--file, with path-confinement/symlink-escape/non-regular-file rejection),
collision-safe quick-id preallocation under withPlanningLock, BATCH.json
schema/validation/resume, dependency-DAG + partitionByFileOverlap wave
construction, and exactly-once STATE.md completion (including the
STATE-row-written-but-manifest-not-yet-updated crash window).

The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a
not-yet-written src/quick-batch.cts) does not exist yet — every test in both
files fails at the top-level require() before any assertion runs. Five
fast-check properties cover collision-freedom under lock contention, resume
idempotency, exactly-once STATE completion, wave totality, and DAG-respecting
wave order, per the design doc's property-based-coverage requirement.

* feat(#3675): implement quick-batch core primitives

Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to
gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core
primitives per the phase design lock — pure/state primitives and
CLI-testable core operations only, no agent dispatch, no worktree creation,
no user-facing command (Phase 4/#3676's job):

- parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list
  parsing (>=2 items required) and a --file variant strictly confined to the
  planning workspace root via requireSafePath, rejecting non-regular-file
  targets.
- allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id
  preallocation under withPlanningLock, checked against both on-disk
  .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json
  manifests (never on-disk-only, which would miss another in-flight batch
  that hasn't dispatched any real quick directory yet) — replicates
  cmdInitQuick's own grammar rather than delegating to it (that function's
  2-second granularity is not batch-safe).
- computeWaves: deterministic wave construction combining dependency-DAG
  layering with partitionByFileOverlap (#3674), called per DAG layer over
  path-separator-normalized planned_files — normalization happens at this
  module's boundary, never inside the Phase 2 helper.
- loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated
  JSON, wrong types, missing fields, out-of-batch dependency references,
  dependency cycles, a worktree path absent from disk).
- resumeBatch: skips complete items, never auto-retries failed items,
  propagates/reverses blocked status along the DAG to a fixed point, and
  detects a STATE.md row that already exists for a non-complete item (the
  "STATE written, BATCH.json not yet updated" crash window) — completing it
  without re-appending. Idempotent across repeated calls.
- completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion —
  appendQuickTaskRow (unmodified) is called at most once per quick id, gated
  by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks
  Completed" table, since appendQuickTaskRow itself carries no idempotency.

BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling
of .planning/quick/ — never inside it, so scanQuickTasks never misreads a
batch manifest as a broken quick task.

* docs(#3675): register the new quick-batch module

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row plus the regenerated
docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary
entry matching the convention set by the sibling File Overlap Partitioner
Module (#3674) entry it sits beside.

NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically
sorted single-entry insertion into families.cli_modules, matching the
existing file's structure) rather than via
`node scripts/gen-inventory-manifest.cjs --write` — this session's
MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script
path from Bash, with no available Memtrace tool to route through instead.
The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs
--check` to confirm this hand-edit is byte-identical to the generator's own
output before merging.

* fix(#3675): resolve lint findings in quick-batch primitives and tests

Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color
array, two unnecessary `as string[]` casts TS 5.5's inferred type
predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the
Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch`
import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a
CONTEXT.md glossary illustration that looked like a real file reference.

* feat(#3675): close acceptance-criteria gaps found in review

Standards- and spec-axis review (plus a self-caught race) surfaced real
gaps against issue #3675's own acceptance criteria and this repo's test
conventions:

- BATCH.json was missing options, base_revision, per-item wave, and
  per-item commit — the issue's AC explicitly lists all four as things
  the manifest must track. Added them: createBatch persists caller-supplied
  batchOptions/baseRevision verbatim and assigns each item its computed
  wave index; completeQuickItem now persists the commit onto the item,
  not just the STATE.md row. All four are backward-tolerant on load (an
  older/hand-built manifest without them still validates).
- resumeBatch had no "incompatible base divergence" check at all, despite
  the AC and the ADR's own "Base divergence" section requiring one. Added
  an opt-in currentBaseRevision comparison that fails closed with a
  recoverable diagnostic on mismatch, and touches nothing on refusal.
- resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the
  only durable write path in this module that wasn't lock-protected,
  a real lost-update race against a concurrent completeQuickItem or
  another resume. Now runs inside the same lock createBatch/
  completeQuickItem use.
- loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no
  size cap (security review, Low/informational); switched to the
  existing safeJsonParse (1MB cap) for defense-in-depth.
- Parser (parseTaskList) had only example-based tests; CLAUDE.md requires
  a fast-check property test for parsers. Added one plus a companion
  reject-property for <2 items.
- The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at
  any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for
  direct limit-1/limit/limit+1 testing without needing 46k fixture dirs.
- Issue AC explicitly asks for prompt-injection-payload test coverage,
  distinct from the existing shell-metacharacter test; added one.
- Test row 9 (FIFO skip) silently returned instead of calling t.skip(),
  so an unsupported platform would report a pass rather than a documented
  skip; fixed to bind the test-context param and skip properly.
- Extracted toWaveInput to remove a 2-site production duplication of the
  QuickBatchItem -> computeWaves reshape (Standards-axis smell).
- Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/
  is user-facing even though the compiled .cjs is gitignored).

* fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason

gsd-test caught this: switching loadBatch to safeJsonParse changed the parse-
failure message shape ("... parse error — ...") without preserving the
"not valid JSON" substring row 27's own test asserts on. Re-wrap
safeJsonParse's error into the original diagnostic phrasing regardless of
which of its three failure modes fired.

* docs(#3675): backfill changeset pr number to 4190

* fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one

CI caught this on windows-latest: row 9's platform-skip only caught mkfifo
throwing (command not found). On this runner mkfifo resolves to something
that exits 0 without creating a file (NTFS has no FIFO concept), so
execution fell through to parseTaskListFromFile against a path that
doesn't exist, producing an ENOENT stat error instead of the expected
"not a regular file" rejection. Check the artifact actually exists before
trusting a zero exit code, and skip with a documented reason either way.

---------

Co-authored-by: sim <sim@local>
2026-09-02 15:38:27 -04:00
Tom Boucher
1c4a00244f ci(#4196): auto-merge Dependabot patch/minor bumps once required checks pass (#4200)
Dependabot already opens a correct fix PR within minutes of a new
advisory (e.g. #4193 for GHSA-jqff-g426-hqxp), but nothing merged it --
it sat until a human noticed `next` had gone red and an unrelated PR
tripped over the same npm-audit gate. Auto-approve + auto-merge closes
that gap for patch/minor bumps; major bumps still need a human.

Co-authored-by: sim <sim@local>
2026-09-02 15:38:09 -04:00
dependabot[bot]
0598a2cf2c chore(deps): bump the npm_and_yarn group across 1 directory with 2 updates (#4193)
Bumps the npm_and_yarn group with 2 updates in the / directory: [browserslist](https://github.com/browserslist/browserslist) and [fast-uri](https://github.com/fastify/fast-uri).


Updates `browserslist` from 4.28.2 to 4.28.8
- [Release notes](https://github.com/browserslist/browserslist/releases)
- [Changelog](https://github.com/browserslist/browserslist/blob/main/CHANGELOG.md)
- [Commits](https://github.com/browserslist/browserslist/compare/4.28.2...4.28.8)

Updates `fast-uri` from 3.1.5 to 3.1.7
- [Release notes](https://github.com/fastify/fast-uri/releases)
- [Commits](https://github.com/fastify/fast-uri/compare/v3.1.5...v3.1.7)

---
updated-dependencies:
- dependency-name: browserslist
  dependency-version: 4.28.8
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: fast-uri
  dependency-version: 3.1.7
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-02 15:37:37 -04:00
Tom Boucher
7c52344284 fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope

* fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag

Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself

* test(#4003): align shape assertions with the implemented gate text

* fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next)

* fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps

* test(#4003): the gate pin tracks the anchored scope grep

* fix(#4003): trim the gate rationale to hold the 93400 margin ceiling

* test(#4003): the RED-grep pin tracks the milestone-bounded invocation

* chore(#4003): changeset for the anchored resume-gate scope

* chore(#4003): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 13:45:14 -04:00
Tom Boucher
dce40eeb6e fix(#4002): rewrite zcode command @-refs to the zcode runtime home (#4188)
* test(#4002): zcode commands must rewrite at-refs to the zcode home

* fix(#4002): add the missing zcode case to the runtime rewrite engine

* chore(#4002): add ZCode to the bug-report runtime dropdown and drop the changeset

* fix(#4002): attribute zcode command and skill ripples to the rewrite engine

* fix(#4002): attribute zcode nested-skill ripples to the rewrite engine

* chore(#4002): backfill changeset pr number

* fix: bump qs past GHSA-x5fp-wj9c-mxmx (transitive, advisory reddened next)

---------

Co-authored-by: sim <sim@local>
2026-09-02 12:19:41 -04:00
Tom Boucher
acb903c2e8 enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable

Add `workflow.code_review_point` (`execute:post` default, or
`execute:wave:post`) so a multi-wave phase can run code review once per
wave instead of once at the end, scoped to what changed since the phase's
prior review.

The code-review capability now declares its step at both loop points via a
new generic `pointFrom` step field: `pointFrom` names an enum config key,
and the step is only active at its own `point` when that key resolves to
a matching value. `_resolvePointGate` (capability-activation.cts) is the
single shared implementation consumed identically by loop-resolver.cts and
capability-state.cts, and capability-validator.cjs enforces that `pointFrom`
references an enum key whose values cover the declaring step's own point.

code-review.md's manual-invocation gate now reads `workflow.code_review`
directly instead of probing registry presence at the hardcoded execute:post
point (so manual `/gsd-code-review` keeps working regardless of which
automatic point is configured), and its file-scope tiers narrow to what
changed since the phase's last review commit when one exists.

execute-phase.md's wave-post step dispatch gets a small, precedented
carve-out so the code-review skill still receives its required phase
argument when dispatched generically (caught by the isolated spec review).

Closes #3661

Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers.
Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill.

* docs: backfill changeset PR number for #3661 (#4159)

* fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd

Five fault-injection mocks in the "bug #1008" describe blocks intercepted
every fs.writeSync call regardless of file descriptor, and several threw or
truncated unconditionally on the first call. This surfaced as an
intermittent macOS CI failure: node:test's own IPC channel back to the
parent process (which also goes through fs.writeSync internally) could get
a bogus injected error or truncated write if node's internal machinery
called it while one of these mocks was active, corrupting the message
frame the parent tried to deserialize ("Unable to deserialize cloned
data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file
IPC crash, not a test assertion failure).

Root cause confirmed by a working counter-example already in the same
file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection
and were never implicated. Applied the same fd-scoped pattern to the five
unscoped mocks (four output()-targeting tests gate on fd 1, one
error()-targeting test gates on fd 2), and added a regression test proving
an unrelated fd passes through untouched while the fault-injection mock is
active.

Found while verifying #3661; unrelated to that change's own diff.

---------

Co-authored-by: sim <sim@local>
2026-09-02 11:01:54 -04:00
Tom Boucher
6fdac3947b fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell on the watermark (#4184)
* test(#3996): antigravity diagnostic must carry stderr and gate the stall tell

* fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell

* fix(#3996): drop the stall token from the session-started sentence

* fix(#3996): decide session-started from watermark growth, type the failure mode

* chore(#3996): changeset for antigravity diagnostic stderr carry

* chore(#3996): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 09:46:21 -04:00
Tom Boucher
77dcdda534 enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4)

* fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4)

* test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift

src/core-utils.cts's new #4014 import block shifted every subsequent line by
6, moving generateSlugInternal's real closing brace from line 193 to 199.
tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that
line number to plant a synthetic violation immediately after the function's
real body; the guard script itself locates the boundary dynamically via
brace-matching and needed no change.

* docs(#4014): document the unreadable-directory scope signal and add changeset

* docs(#4014): backfill changeset PR number to #4163

* test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff

---------

Co-authored-by: sim <sim@local>
2026-09-02 08:12:57 -04:00
Tom Boucher
2131fe13f3 enhance(#3464): exec() detection widening, citation-debt cleanup — Phase 8 (#4171)
* feat(#3464): widen no-source-grep to detect regex.exec() on tracked text

Adds an execCall kind alongside the existing regexTest detection --
regex.exec(tracked) was invisible to the rule while regex.test(tracked)
was already caught, despite both reading a source-derived string through
a regex. Measured: 4 previously-invisible sites across 2 files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3464): migrate 4 sites newly flagged by the exec() widening

docs-hooks-table-parity.test.cjs's three regex-extraction loops are
site-scoped marked (source-text-is-the-product) -- the dynamic
preToolEvent/postToolEvent dialect branching they mirror is explicitly
documented as not statically parseable, so a literal-pattern mirror is
the practical minimum-cost check.

no-bare-gsd-tools-command-position.test.cjs's readRouterVerbs() now
requires HOST_COMMAND_ROUTERS directly instead of regex-walking
gsd-tools.cjs's source text -- the same accessor three other suites
already use.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3464): pay down 6 grandfathered uncited allow-test-rule markers

Two were genuinely load-bearing (suppressing a real detected violation)
and just needed a citation added -- phase6-capstone-conformance.test.cjs,
runtime-name-policy.test.cjs, both now (#3464).

Four were dead-weight file-header markers suppressing nothing -- each
file's real effective sites are covered by separate, already-cited
markers elsewhere in the same file. Deleted outright rather than cited,
per Phase 1's own precedent (remove non-load-bearing markers instead of
grandfathering them forever) -- codex-config.test.cjs (two copies),
gsd-check-update-worker-platform-gate.test.cjs, orphaned-hooks.test.cjs,
settings-jsonc.test.cjs.

allowlist.json: 134 -> 128 entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3464): re-baseline effective-exemption ceiling to 84

The exec() widening's 3 newly-marked sites are now suppressed and
counted; ceiling rises 81 -> 84, the exact measured high-water mark.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3464): correct citation and restore a wrongly-deleted marker

Two review corrections, both found by the orthogonal review pass:

- docs-hooks-table-parity.test.cjs's 3 new exec() markers cited #3464
  (mechanically "the phase that widened the rule") when the file's own
  established, correct reference is #3839 (the issue this whole test
  exists to enforce, already cited in its file header) -- fixed to match.

- gsd-check-update-worker-platform-gate.test.cjs's deleted file-header
  marker was NOT dead weight: its codeOnly() helper wraps readFileSync
  and is called inline as an assert argument, a genuine source-grep
  pattern on real .cjs/.js source that the rule cannot currently see
  (helper-function indirection is a distinct blind spot from anything
  Phase 7/8 measured) -- CONTRIBUTING.md is explicit that "unverified"
  is not the same as "vestigial." Restored, site-scoped this time
  (directly above codeOnly(), not as an inert file-header comment) and
  cited (#3103, the issue the file's own docstring already references).

codex-config.test.cjs's two deletions and orphaned-hooks.test.cjs's /
settings-jsonc.test.cjs's deletions were independently re-verified and
stand: their flagged lines read generated .toml/.json OUTPUT, not
source, or have no residual pattern at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 08:11:23 -04:00
Carlos Cativo
9b77320580 fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh (#4092)
* fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh

gsd-node-runner.sh was registered in MANAGED_HOOKS but shipped without a
gsd-hook-version header, so gsd-check-update-worker.js always classified it
as 'definitely stale' (a missing header is indistinguishable from a
pre-version-tracking file). Every install on an otherwise up-to-date
version showed a permanent, unclearable '⚠ stale hooks — run /gsd-update'
warning naming this one file.

Root cause: the build-hooks.js comment claimed the file is 'not a
registered hook' and 'staged verbatim — no templating', but it IS in
MANAGED_HOOKS (managed-hooks-registry.cjs:34) and install.js already
stamps {{GSD_VERSION}} into every .sh hook unconditionally, gsd-node-runner.sh
included. The comment contradicted both the registry and the installer's
actual behavior, and the header line itself was simply never added.

Fix: add the header (matching every other managed .sh hook's format) and
correct the comment so it no longer asserts the opposite of what the
registry and installer actually do.

Adds a regression test that iterates every MANAGED_HOOKS entry and asserts
it carries a header matching the worker's own detection regex, so a future
hook added to the registry without one fails CI instead of shipping
silently.

Fixes #4076

* chore(#4076): add changeset fragment for PR #4092

* fix(#4076): address review nits — drop unneeded exemption, fix blank line

Per @trek-e's review on #4092:
- tests/managed-hooks.test.cjs:96: the readFileSync call uses a loop
  variable (entry-derived hookPath), not a literal path, so
  local/no-source-grep's static literal-path detector never flags it —
  the allow-test-rule exemption comment was unnecessary. Replaced with a
  plain note explaining the source-read rationale.
- tests/managed-hooks.test.cjs:121-122: dropped a stray extra blank line
  before the bug #2136 section divider.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-02 08:04:37 -04:00
Tom Boucher
b0572c0108 feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner

Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted
output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage
golden script) as a regression safety net ahead of extracting
partitionStages into a standalone module. Also adds the new module's
unit and property tests (test matrix rows 1-11) against its expected
public API, which does not exist yet and is added in the next commit.

* feat(#3674): extract file-overlap partitioner into a shared, generic module

Moves partitionStages' greedy first-fit file-overlap algorithm into a new,
dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap),
generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's
Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own
Plan[] shape onto the generic input and back — behavior-preserving, no dependency
ordering, no path normalization, no filesystem access moved or added. Enables a
future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive
without pulling in orchestration internals.

* docs(#3674): register the file-overlap-partitioner module bookkeeping

New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained
registrations beyond the code itself: .gitignore (compiled artifact),
eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs),
docs/INVENTORY.md's CLI Modules roster row (regenerated via
gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry
matching the convention set by similarly-scoped leaf modules
(text-lines.cts, plan-dependency-graph.cts, spec-section.cts).

* fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct

- docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write,
  produced no diff (manifest is keyed by content, not row order)
- tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is:
  tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to
  "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs
  file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier.

* fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction

The `no two plans in the same stage share a modified file` property
reconstructs which physical item produced each output id via
`remaining.findIndex(r => r.id === id)`. Under duplicate ids (an
explicitly-supported input shape for `partitionByFileOverlap`) that
reconstruction can pick the wrong physical occurrence, producing a
false-positive overlap failure (observed counterexample: p0(f1),
p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but
misread by id-order as [[p0,p208#1],...], which do overlap).
Properties (a) determinism and (b) totality already exercise
duplicate ids correctly and are left unchanged; only this property's
generated items are now constrained to unique ids via
`fc.uniqueArray`, where the reconstruction is unambiguous.

---------

Co-authored-by: sim <sim@local>
2026-09-02 07:33:42 -04:00
Tom Boucher
fa107c0461 fix(#4172): add --merge-async to test:coverage:scripts-floor (#4173)
* fix(#4172): add --merge-async to test:coverage:scripts-floor

The "Coverage gate (merged shards)" test.yml job has OOM-crashed (exit
134, SIGABRT) on every push to next since 4dfc46b. test:coverage:scripts-floor
was the only c8 coverage-merge invocation in package.json still missing
--merge-async: c8's default sync merge path (Report._getMergedProcessCov)
loads every raw V8 coverage file from the merged 3-shard coverage/tmp
directory into memory as one array before merging, instead of folding
them in one at a time, and now blows through the job's 8192 MB heap
ceiling.

This is the same bug class as #4068 (fixed in 4d70b4dc4), which added
--merge-async to test:coverage:unit and test:coverage:report -- the two
other scripts sharing this merge/report code path -- but missed this
third sibling, which reads the exact same merged coverage/tmp data in
the same job. check-coverage and report both dispatch through c8's
shared getCoverageMapFromAllCoverageFiles()/mergeAsync branch, so the
fix is identical in shape to #4068's.

Extends the existing #4068 regression guard (tests/c8-merge-async-flag.test.cjs)
to also assert test:coverage:scripts-floor carries --merge-async, closing
the coverage gap #4068 left on this sibling script.

Fixes #4172

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4172): backfill changeset PR number (pr:0 -> 4173)

* fix(#4172): route scripts-floor through `c8 report`, not `check-coverage`

The prior commit (1a3d6b75e) added --merge-async to
test:coverage:scripts-floor, but the Coverage gate job still OOM-crashed
identically on PR #4173's CI run
(https://github.com/open-gsd/gsd-core/actions/runs/33587951779/job/100118358975).

Root cause of that miss: c8@11.0.0's `check-coverage` CLI subcommand
handler (node_modules/c8/lib/commands/check-coverage.js) never forwards
argv.mergeAsync into the Report constructor -- only the `report`
subcommand's handler (node_modules/c8/lib/commands/report.js) and the
default command (which also calls into report.js) do. Verified directly
by constructing Report the same way each handler does: the check-coverage
path yields report.mergeAsync === undefined even with --merge-async on
the command line, while the report path yields true.

Fix: route test:coverage:scripts-floor through `c8 report --check-coverage`
instead of `c8 check-coverage`. report.js's outputReport() calls the same
checkCoverages() threshold-checking helper when --check-coverage is
truthy, so behavior (and exit code on threshold failure) is unchanged --
only the code path taken to get there now actually honors --merge-async.

Extends the tests/c8-merge-async-flag.test.cjs regression guard with an
explicit assertion that scripts-floor invokes `c8 report`, not
`check-coverage`, so this can't silently regress back to the broken form.

Fixes #4172

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4172): correct changeset to describe the actual root cause (check-coverage vs report subcommand gap, not just the missing flag)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 06:39:15 -04:00
Tom Boucher
5c7243e54b fix(#3995): derive the review diff base from the phase directory (#4181)
* test(#3995): diff base keys on the phase directory, not commit subjects

All three derivation sites (Tier 3, spawn_reviewer, fallow pre-pass)
must anchor on the phase directory's first commit; the milestone-blind
repro (an archived milestone's same-numbered phase commit capturing
the base) is the failing-first row. #3191/#3503 rows reworked to the
directory-anchor contract; T6 docs-parity forbids any remaining
phase-scope message-grep site.

* fix(#3995): derive the review diff base from the phase directory

A phase number is unique within a milestone, not a repository; the
message grep had no milestone bound and tail -1 deliberately selected
the oldest same-numbered subject, dragging archived milestones phases
into the scope (7 files to 3388 plus the >50 depth downgrade). All
three lockstep sites now anchor on the first commit that added anything
under the phase own directory — the same anchor class
git-base-branch phaseStartCommit uses. ShellCheck baseline gains the
escaped fragment shifted parse signature.

Emitted-Drift-Ack-Growth: code-review.md — phase-directory anchor replaces the message-grep derivation at both sites (#3995)

* chore(#3995): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 06:03:17 -04:00
Tom Boucher
647365faf1 fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection

Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as
a gate condition, the end-of-phase escalation must not require MVP, the
executor agent's gate section triggers on TDD_MODE alone, and the gate
semantics reference loads without MVP_MODE.

* fix(#4011): key the TDD runtime gate on TDD_MODE alone

The RED-commit gate shipped as #76's MVP slice kept the paired
invocation's conjunct, so workflow.tdd_mode=true was silently inert on
every non-MVP phase, contradicting references/tdd.md's own contract.
Drops the MVP conjunct from the per-task gate and the end-of-phase
review escalation; rescopes execute-mvp-tdd.md's load condition,
gsd-executor's gate section, and mvp-concepts' intersection claim.
MVP remains free to imply TDD; the file is not renamed (stated
assumption in the PR body).

* test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing

Review follow-ups: the detector now only inspects if/[ condition lines
so explanatory prose mentioning both flags cannot trip it; remaining
'under/outside MVP+TDD' phrases in execute-phase.md, the gate
reference, and docs/INVENTORY.md now describe TDD-mode semantics.

Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011)
Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011)

* chore(#4011): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 04:24:52 -04:00
Tom Boucher
383c2f6b34 fix(#3982): strip closed-milestone details from the current window (#4177)
* test(#3982): archived details must not leak into the current-milestone window

Parser-level regression (newest-first layout, two archived details
blocks) plus the issue's end-to-end phase.complete fixture: completing
phase 20 must advance to 21, never backwards into the archived range.

* fix(#3982): strip closed-milestone details from the current window

The heading-located window stripped <details> archives from the
preamble but not from currentSection; on newest-first roadmaps the
archived titles sit in summary tags rather than headings, so the
section walk reached end-of-document and the window swallowed every
collapsed archive below the active milestone. The strip is gated on
isClosedMilestoneHeading over each block's summary — the issue's
prescribed narrow form — so the active milestone's own collapsed
blocks (#1341) survive instead of trading this bug for the
phase_count: 0 class (#557/#2947).

* chore(#3982): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 02:33:54 -04:00
Tom Boucher
9eef1b791b fix(#3981): give blocking guards a host-stall-proof timeout budget (#4175)
* test(#3981): blocking-guard timeout budget and 5→120 migration

Fresh registration must carry a host-stall-proof 120 s budget on the
six blocking PreToolUse guards; existing managed timeout:5 entries
are migrated; non-managed entries and advisory budgets are untouched.

* fix(#3981): give blocking guards a host-stall-proof timeout budget

Claude Code treats a timed-out hook as non-blocking, so the 5 s
budget on the six blocking PreToolUse guards silently dropped the
gates exactly when the host stalled under load. Registers them at
120 s (covers every observed stall, max 84.3 s) and migrates existing
managed timeout:5 entries in place, context-monitor-backfill shape.
Advisory hook budgets are unchanged.

* chore(#3981): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 01:26:37 -04:00
Tom Boucher
7960374d15 fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git

Behavioral coverage: extract the trailers:key token from ship.md and
prove a real git commit carrying that trailer reads back via
%(trailers:key=token,valueonly). gate_status contains an underscore,
which git's trailer machinery cannot tokenize, so the audit read was
structurally empty.

* fix(#3962): rename the TDD-Audit trailer token to gate-status

Underscore is not a valid git trailer token character, so
%(trailers:key=gate_status,...) could never match. Renames the token
at the read, the documented aggregate write, and the section's
prose/table header. Self-suppression semantics (#2431) unchanged.

* test(#3962): compare outcome against the seam's literal, fix doc token

Review follow-ups: the round-trip test compared outcome to 'EXITED'
but the process seam emits 'exited'; and the trailer rename is
propagated to docs/ship-pr-body-sections.md's read/write examples.

* chore(#3962): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 00:01:55 -04:00
Tom Boucher
10593e2cf2 fix(#3959): give review-prompt plans citable path anchors (#4170)
* test(#3959): plan copies carry source names, prompt carries path anchors

Regression coverage: budget copies named gsd-review-plan-<plan-id>.md
(not a bare padded index), no bare-index copy survives, and the Plans
to Review template instructs a per-plan repo-relative #### path header.
Also corrects the stale -00 assertion to the source-named copy.

* fix(#3959): give review-prompt plans citable path anchors

The Plans to Review template now instructs a per-plan repo-relative ####
path header, and the budget copies are named gsd-review-plan-<plan-id>.md
instead of a bare padded index, restoring provenance for prompt-budget's
per-plan headers while keeping the trim glob intact.

Emitted-Drift-Ack-Growth: review.md — per-plan path-header instruction in Plans to Review + provenance-named copy loop (#3959)

* chore(#3959): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-01 22:40:23 -04:00
Tom Boucher
91d5fdff6f chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces

Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse
advisory output alongside the existing additionalContext prose:

- gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName
- gsd-context-monitor.js: severity ('warning'|'critical')
- gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS
  + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523
  pattern)
- gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its
  findings array already existed from #3523)
- gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg,
  distinct from the existing force-add block leg's code

additionalContext stays byte-identical in every hook (verified per-hook
against the pristine HEAD version across a spread of payload shapes).

Migrates all 20 assertion sites named in the issue off
additionalContext.includes(...)/assert.match(...) substring-matching
onto the new typed fields, per CONTRIBUTING.md's prohibition on raw
text matching on test outputs.

Closes #3546

* test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper

Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called
gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so
each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's
own doc comment says is sized for plumbing reads (rev-parse/branch/log),
not write-heavy add/commit spawns. That file already documents the exact
same defect class from a prior incident (PR #3323) and exports
GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites -
commitN just wasn't using it. Observed failure: `git commit -m filler 9`
timed out under normal bench load, unrelated to any of this PR's own
diff (hooks/*.js + 5 other test files).

Not a flake: root-caused to the timeout bound being sized for the wrong
call class, per this repo's no-flakes rule.

* chore(#3546): backfill changeset PR number (#4167)

---------

Co-authored-by: sim <sim@local>
2026-09-01 22:32:54 -04:00
Tom Boucher
f16ff7d1b3 enhance(#3545): widen no-source-grep with fold+hooks, migrate 76 sites (#4161)
* feat(#3545): widen no-source-grep with one-hop path-fold and hooks dir

Resolve a readFileSync() path argument that is a bare Identifier one hop
back to its VariableDeclarator initializer before classification, and
recognize `hooks` as a source directory alongside bin/lib/gsd-core/src.

Measured (epic #3464 phase 7): fold+hooks together newly flag 76
unsuppressed sites across 18 files that were previously invisible to
identifier-indirected or hooks/-rooted source reads. Neither widening
alone is sufficient — hooks-only surfaces 0 new sites, confirming #3520's
prior finding that the identifier-indirection gap must close first.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3545): migrate 76 sites newly flagged by the fold+hooks widening

Per-site classification: rewrite behaviorally (require() the real module,
assert on its actual exported behavior) wherever the read was a proxy for
code behavior; add a site-scoped `// allow-test-rule: <reason> (#3545)`
marker only where the raw source text genuinely is the product under test
(codex-config.test.cjs's adapter-header-contract checks, install.js
structural-wiring guards with no exported symbol, AST-parse fixture
inputs, etc.) — each marker cites an existing repo-sanctioned category
from CONTRIBUTING.md's allow-test-rule exception table.

Also converts two try/finally test bodies (introduced during this same
migration) to the required t.after() cleanup pattern per CONTRIBUTING.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3545): re-baseline effective-exemption ceiling to 81

The fold+hooks widening's own newly-detected sites are now suppressed by
site-scoped markers, moving them from invisible into the tightly-ratcheted
effective-exemption count. Ceiling rises from 10 to 81 (the exact measured
high-water mark, grace unchanged at 2) — a deliberate, measured re-baseline
per the widening working as intended, not an ordinary ceiling bump.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3545): use canonical allow-test-rule category tokens

4 markers added during migration cited an issue ref correctly but didn't
use one of CONTRIBUTING.md's seven recognized category tokens, unlike
every other marker in this change. Cosmetic only — same suppression
lines, same effective/live counts (81/81, 0 live).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3545): correct stale phase-artifact path in test comment

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:40:38 -04:00
Tom Boucher
ff0361071d feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane

cursor-agent exposes --model (204 selectable models) but the cursor lane
declared modelArg: null / modelConfigKey: null, so review.models.cursor
was rejected as an unknown config key and the #1517 reviewer-instances
escape hatch silently discarded a configured model at modelExpansion.

Wire the lane the same way codex already is: inject {{model}} into args
right after -p, set modelArg to --model, and declare modelConfigKey as
review.models.cursor plus its config schema entry. An unconfigured lane
still invokes byte-identically to today.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3653): add changeset for review.models.cursor

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3653): update co-change surfaces that assumed cursor has no model key

gsd-test surfaced three surfaces still hardcoding "cursor declares no
modelConfigKey", broken by wiring review.models.cursor:

- tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797
  invariant test listed cursor among lanes that must own no model key.
- tests/settings-integrations.test.cjs: the #3651 keyless-lane test
  listed cursor as keyless, including a live config-set assertion that
  now correctly succeeds instead of failing (swapped to qwen).
- gsd-core/workflows/settings-integrations.md: the settable-keys
  enumeration and two prose call-outs still named cursor as keyless.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3653): acknowledge deliberate growth of settings-integrations.md

settings-integrations.md grew 4 bytes because it now enumerates
review.models.cursor as a settable key alongside the other reviewer
lanes, matching the modelConfigKey wired for cursor in this PR.

Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer

The previous commit's trailer was separated from Co-Authored-By by a
blank line, splitting it into an earlier, non-trailer paragraph — git's
trailer parser only recognizes the last contiguous block. Restating it
here immediately adjacent to Co-Authored-By so both parse as trailers.

Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3653): backfill changeset pr number

pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:40:20 -04:00