Commit Graph

2671 Commits

Author SHA1 Message Date
Tom Boucher
8ecbac0350 fix(#3189): drop non-id prose from phase_req_ids before gap check (#3438)
* fix(#3189): drop non-id prose from phase_req_ids before gap check

* chore(#3189): pin changeset fragment to pr 3438

* fix(#3189): use hyphenated IDs in e2e fixtures (parseRequirements requires hyphen)

* fix(#3189): assert e2e item set via table, not rows (check query has no rows)

* fix(#3189): check prose fragments against table rows, not the table heading

---------

Co-authored-by: sim <sim@local>
2026-08-14 01:10:48 -04:00
Tom Boucher
dbbcb8f736 fix(#3191): anchor remaining diff-base greps, portably (#3437)
* fix(#3191): anchor remaining diff-base greps, portably

The #2989 fix anchored only the Tier-3 grep, and did so with \b — not a
POSIX ERE token, so on macOS regex(3) it silently matches nothing and
Tier 3 always fails closed. spawn_reviewer's agent-context DIFF_BASE and
the fallow structural pre-pass's --changed-since base each still ran the
original unanchored --grep="${PADDED_PHASE}", whose oldest substring
match is routinely a version-string/date commit from months before the
phase existed — feeding the reviewer agent a bogus diff_base exactly
when files: is empty, and widening fallow's changed-files scope.

All three derivations now use the same anchored, POSIX-portable
'[Pp]hase N([^[:alnum:]_]|$)' with --extended-regexp; spawn_reviewer
also gains Tier-3's parent-exists guard so the two computations are the
same algorithm. Behavioral regression tests execute the shipped bash
extracted from the workflow files against a git fixture on every
platform, so the macOS \b hole is covered, not just the Linux CI view.

* chore(#3191): backfill changeset PR number 3437

* fix(#3191): scope fallow test snippet past the gsd-tools resolver

The CI runners have no installed gsd-tools, so executing the resolver
line that precedes FALLOW_SCOPE_ARGS in the extracted fence exits 1
before the derivation under test ever runs. Slice the snippet to start
at FALLOW_SCOPE_ARGS=() — the resolver is orthogonal to the base
derivation the regression test binds.

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:53:13 -04:00
Tom Boucher
483083aea6 fix(#3194): verify source-grounded lane evidence from review output (#3436)
* fix(#3194): verify source-grounded lane evidence from review output

* chore(#3194): fill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:29:42 -04:00
Tom Boucher
1d5d77951c fix(#3190): commit review.md in --auto loop; fix report env var (#3434)
* fix(#3190): commit review.md in --auto loop; fix report env var

Three coupled defects in gsd-core/workflows/code-review-fix.md:

- The --auto re-review loop overwrote REVIEW.md each iteration but the
  single docs commit staged only REVIEW-FIX.md, so the committed REVIEW.md
  stayed at iteration 1 and contradicted the committed REVIEW-FIX.md. The
  --auto commit now stages the converged REVIEW.md alongside REVIEW-FIX.md
  (guarded on AUTO_MODE; non-auto single-pass runs unchanged).
- The two inline frontmatter validators (HAS_STATUS, FIX_FRONTMATTER)
  exported REVIEW_PATH into a node -e body that reads process.env.
  FIX_REPORT_PATH, so the status check was always empty and REVIEW-FIX.md
  was never committed. Both now export FIX_REPORT_PATH.
- On successful convergence the spent .iterN.md backups are removed so the
  phase directory is clean; they are retained on degradation for post-mortem.

Regression test: tests/code-review-fix-pipeline-regression.test.cjs.

* chore(#3190): set changeset pr to 3434

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:26:07 -04:00
Tom Boucher
ff2d08d453 fix(#3193): tolerate attributes on plan-task child tags (#3433)
* fix(#3193): tolerate attributes on plan-task child tags

* chore(#3193): add changeset

* chore(#3193): set changeset pr to 3433

---------

Co-authored-by: sim <sim@local>
2026-08-14 00:14:07 -04:00
Tom Boucher
0d4d78550f fix(#3188): null absent planning-doc paths in init phase queries (#3430)
* fix(#3188): null absent planning-doc paths in init phase queries

The init phase-op / plan-phase / execute-phase queries emitted a non-null
absolute requirements_path / state_path / roadmap_path even when the named
file did not exist — built with a bare path.join and no existence check,
unlike the conditional sibling fields (patterns_path, context_path, ...) in
the same payload. Consumers (e.g. ultraplan-phase.md:104 'requirements_path
is not null') therefore read missing files.

Each of the three reading sites now returns null when the file is absent and
its absolute path when present. The project/milestone-bootstrap and doc-ingest
emitters that use these paths as write-targets for not-yet-created files are
intentionally unchanged.

* chore(#3188): backfill changeset pr: 3430

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:50:22 -04:00
Tom Boucher
a3b72d9071 fix(#3171): init execute-phase emits display name, not directory slug (#3429)
* fix(#3171): init execute-phase emits display name, not directory slug

When a phase directory already exists on disk, the disk-lookup path
(searchPhaseInDir) derived phase_name from the directory-name remainder --
itself an already-slugified value (phase.add writes ${num}-${slug} dirs) --
so phase_name and phase_slug came out byte-identical. The execute-phase
workflow forwards phase_name into 'state begin-phase --name', which wrote
that raw slug into STATE.md's current_phase_name on every phase start.

cmdInitExecutePhase now prefers the ROADMAP's curated display name
('### Phase N: <Name>') for phase_name, matching the no-disk fallback path
that already did this correctly. phase_slug is unchanged (it feeds
branch-name construction). The state.begin-phase authoritativeFm override
(#2821/#2736) is untouched; the correction is in the value fed into --name.

The milestone_name half of #3171 was subsumed by #3216 / PR #3226; this
fixes the remaining current_phase_name half.

* docs(changeset): backfill pr 3429 for #3171

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:49:23 -04:00
Tom Boucher
4792bc02f5 fix(#3165): recover phase_count when scoped window is truncated (#3428)
* fix(#3165): recover phase_count when scoped window is truncated

* chore(#3165): add changeset for roadmap.analyze phase_count recovery

* chore(#3165): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:47:23 -04:00
Tom Boucher
7976b1ca0d feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing

Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing.

- src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint)
- agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names
- gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json)
- execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling)
- execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree)
- config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator)
- docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests)

* chore(#1689): backfill changeset PR number (#3417)

* chore(#1689): regenerate install-tree fixtures for new workflow fragment

* chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing)

* test(#1689): SPAWN contract allows parameterized subagent_type placeholder

agent-frontmatter's spawn-type checks scanned subagent_type="..." as a
concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}"
(a runtime placeholder resolved via resolve-agent, default gsd-executor).
Skip {TOKEN} placeholders in both the known-type and <available_agent_types>
checks; execute-phase still lists the built-in roster incl. gsd-executor.

* fix(#1689): CI conformance for the routing fragment

- per-plan-executor-routing.md: add the canonical runtime-launcher preamble to
  its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments.
- agent-install-check.cts: drop a literal ~/.claude/agents path from the
  resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs
  (cline install leak guard).

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:10:52 -04:00
Tom Boucher
470389f3a2 chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169

Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes
hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"):
tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and
indentWidth (bullet-nesting depth).

git-cmd.js migrates onto tokenizeShellLike with zero behavior change
(parity-asserted against every existing #3129 fixture in
tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases
1-3 (env-prefix skip, executable check, global-option consume) extracted
into skipToSubcommand, shared with the new extractBranchArgument (git
checkout -b / git branch <name>) — a new capability exercising the seam
on the domain the ADR names, not a migration of existing duplicated logic
(none existed).

Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish
a cross-reference bullet nested under an open decision from a fresh
malformed declaration attempt. An earlier bold-run-content-classification
design was tried and disproven against the repo's own existing FIX-B
fixtures (D-02, "no colon no dash") before being adopted — both have
identical shape under any content-only rule. Nesting depth (via
indentWidth) is the actual distinguishing signal: a bullet indented
deeper than the currently-open decision's own bullet is elaboration,
folded into its text like a continuation line, never tested against the
parse-miss guard. A bullet at the same-or-shallower indent is unchanged.

Scope-narrowing disclosed, not silent: of the ADR's four named bugs
(#3197, #3169, #2570, #2528), three no longer need this phase's work.
were independently fixed and closed since the ADR was authored — #2570's
fix is already a correctly-bounded regex per the ADR's own decidability
test (no scanner needed); #2528's fix is a deliberate, twice-reviewed
non-scanner design (its own code comment records a scanner-based attempt
that regressed a symmetric case and was reverted) that this phase does
not disturb. Only #3169 required new work.

get_impact: isGitSubcommand CRITICAL/196 affected symbols,
parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence).

Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md,
docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary.

Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md
Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3414): add required fast-check property tests per code review

TESTING-STANDARDS.md:169 requires at least one fast-check property test
for any module that implements parsing — src/token-scanner.cts had none,
an orthogonal Standards-axis review finding. Adds two seeded property
tests (mirroring Phase 1/2's fast-check-setup.cjs convention):
indentWidth counts exactly a generated leading-space run; tokenizeShellLike
round-trips a generated array of whitespace/quote-free words joined with
single spaces.

The design doc's own "no property test needed" rationale was wrong — it
argued no algebraic law applied, but the standard is unconditional for
parsing modules regardless of whether one "feels" applicable. Corrected
in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md.

Also fixes two Spec-axis wording drifts the same review found between
the design doc and the shipped code (doc-only, no behavior change):
extractBranchArgument's documented signature dropped an unused
subVariants parameter that was never implemented, and the #3169
fail-first fixture description corrected from "15-decision plan via
cmdDecisionCoverageVerify" to the actual compact 3-decision analog via
the real blocking gate, check.decision-coverage-plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): add changeset for #3169 fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): backfill changeset pr number to 3424

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-13 23:08:34 -04:00
Tom Boucher
b9adedbc86 fix(#3151): stop emitting effort: into skill frontmatter (cache invalidation) (#3425)
Claude Code applies SKILL.md effort: as output_config.effort; any change from
the session baseline invalidates the prompt cache at BOTH scope boundaries
(entry + exit, the latter often machine-fired via subagent-completion
notification). The reporter's owned measurement confirms it: /gsd-progress
(effort:low) in a medium session → cache_creation 63,404 (entry) + 18,589
(exit), while a no-effort skill shows none. ~76% of invocations paid in full.

Fix (trek-e AC#2/AC#4): convertClaudeCommandToClaudeSkill no longer emits
effort: into Claude-runtime skill frontmatter (src/runtime-artifact-conversion.cts
+ duplicated bin/install.js). normalizeClaudeSkillEffort removed (dead). The six
declaring skills (plan-phase/execute-phase/autonomous/next/progress/stats) no
longer carry effort. Source command files keep effort (input, used elsewhere);
the separate agent-effort surface (#3160) is untouched.

Tests: install-runtime-artifacts #769 block flipped to assert effort is ABSENT
from installed SKILL.md + converter output (the AC#4 behavioral coverage).

Co-authored-by: sim <sim@local>
2026-08-13 22:00:26 -04:00
Tom Boucher
d30c99bc92 chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference

* test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md

* chore(#1892): reword retired-workflow mentions for removed-but-needed lint

* test(#1892): correct stale surface labels in retargeted suites

* docs(#1892): add verifier-phase-gates row to locale inventories

* chore(#3421): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-13 21:22:03 -04:00
Tom Boucher
b77b7f8e56 fix(#1526): delegate auto-chain post-completion to transition workflow (#3419)
* fix(#1526): delegate auto-chain post-completion to transition workflow

execute-phase's auto-chain completion called phase.complete then a light inline
set (partial PROJECT.md update + offer-next) and never invoked the transition
workflow, silently skipping graduation scan, session-continuity, project-reference,
accumulated-context, and current-position updates — so a phase completed via
auto-chain left different project state than a normal transition.

Fix (delegate, user decision 2026-08-13): replace execute-phase's update_project_md
+ offer_next with a delegation step that @-includes transition.md in post-completion
mode. Add a post_completion_mode step to transition.md that skips verify_completion
+ update_roadmap_and_state (phase.complete already ran; avoids double-write) and
begins at evolve_project. Standalone transition (mode 1) is unchanged.

Regression: tests/auto-chain-transition-delegation.test.cjs (source-text-is-the-
product) asserts the delegation, the skip-set, the removed inline step, and the mode.
Ack fragment 1526 covers execute-phase.md + transition.md growth (spent 2930 fragment
removed — same-path owner conflict, like #3025/#3024).

* docs(#1526): backfill changeset PR number (#3419)

---------

Co-authored-by: sim <sim@local>
2026-08-13 20:31:55 -04:00
Tom Boucher
0624c5da6f chore(#3212): src/text-lines.cts is the sole owner of line-terminator handling — Phase 2 (#3420)
* test(#3413): failing-first suite for the line-terminator seam

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts
does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND
at its require line, which is the intended RED.

The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first:
parseMustHavesBlock currently returns [] for every must_haves block on a
CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript
and two /m-anchored \s* patterns can absorb it, inflating a captured indent
by one character and tripping the "not nested under must_haves" guard.
Verified locally against the current (unfixed) compiled module: both the
direct repro and the silent-exit "blank line before must_haves:" variant
return [] today. A parity property test (crlf vs lf must deep-equal for
every block name) matches a pattern this maintainer has required repeatedly
for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held).

The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's
future fix-hint text (pointing at splitLines()) and its self-reference
non-violation (the seam's own correct \r?\n split must never flag itself).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* chore(#3413): src/text-lines.cts owns line-terminator handling

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/
detectEol/joinLines and migrates frontmatter.cts onto it.

parseMustHavesBlock (#3360, confirmed-bug) returned [] for every
must_haves block on a CRLF plan file. Root cause: \r is its own
LineTerminator in ECMAScript, so under /m two \s*-anchored indentation
lookups could match at the position INSIDE a \r\n pair and absorb the
terminator, inflating the captured indent by one character and tripping
the "not nested under must_haves" guard. Two silent exits, one with a
diagnostic and one without (a blank line before must_haves: hits the
silent path). Fixed by converting both lookups from a whole-string /m
match to split-then-scan — splitLines first, then a per-line, non-/m
match — the same structural pattern parseYamlRegion (30 lines away in
the same file) already used safely. Nothing downstream of the two
lookups changed; blockLines is now sliced from the already-split array
instead of re-splitting a substring, but its contents are unchanged for
LF input, and the per-line dash/kv parsing loop is untouched.

A parity property test (CRLF and LF plans parse to identical must_haves
for every block name) matches a pattern this maintainer has required
repeatedly for prior CRLF fixes in this file's neighborhood (Cortex:
7 recorded decisions, verify_intent -> held).

frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion,
isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter)
are rerouted onto splitLines — a literal 1:1 substitution, zero behavior
change, since splitLines IS that same regex plus a type guard.

The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host-
contract, gen-capability-registry, gen-context-index) are deleted and
rerouted onto normalizeEol, which strips a bare unpaired \r exactly like
the deleted copies did (not just \r\n pairs) -- verified against each
script's own --check mode against its real generated output.

local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its
fix-hint message now naming splitLines() instead of the raw regex --
the prohibition finally has a primitive to point at. Detection logic
unchanged in this phase (deliberate scope limit, see design doc Known
limits: the rule doesn't yet recognize safeReadFile/platformReadSync as
a content source, and has no detector for the \s-adjacent-to-anchor
shape that is #3360's actual mechanism -- the CLASS is converged by the
direct fix + regression test regardless).

joinLines/detectEol are NOT wired into frontmatter.cts's own write path
(cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that
platformWriteSync already, unconditionally converts CRLF->LF on every
.md write today as a pre-existing policy owned by a different module,
and ADR-3212's backward-compatibility clause rules out a file-format
change in any phase. Stated explicitly in Known limits rather than left
for a reader to discover.

Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block),
docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md
glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found

Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the
previous commit) immediately surfaced 13 real, pre-existing violations
across 10 files -- undetected until now because the rule never scanned
src/. This is the exact defect class ADR-3212 exists to close, playing
out again one phase after Phase 1 hit the same shape ("the new lint
rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule,
fixed inline rather than deferred or suppressed; there is no
established suppression convention for this rule in src/ and inventing
one now would undermine the point of widening it.

audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts,
phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or
regex character classes widened to \r?\n / [^\r\n], each following the
same pattern already established migrating frontmatter.cts.

phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` --
already semantically CRLF-safe, but the rule's lexical scanner doesn't
recognize \r? guarding a group (only \r? immediately before a literal
\n). Verified the two forms are equivalent across all four EOL/EOF
cases before restructuring, not assumed.

roadmap-upgrade.cts needed two coupled sites, not the one flagged line:
computeMigrationPlan and applyMigration must agree on line
representation for the lines[edit.lineIndex] === edit.from equality
check to hold, and the write-back needed joinLines + detectEol -- a
plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF
wholesale on every migration. This is the first real production
consumer of joinLines/detectEol in this epic (frontmatter.cts's own
write path doesn't use them -- see the previous commit's Known limits).

Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites
the rule doesn't track (.search() and new RegExp(dynamicString) aren't
in its tracked call/construction set). Investigated each empirically --
hand-tracing this exact bug class already produced one wrong conclusion
earlier in this phase (a detectEol design-doc arithmetic error), so
these were verified with real CRLF fixtures rather than reasoned about
on paper:

  - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split
    ('\n')[0]` leaked a trailing \r into a user-visible todo summary on
    CRLF input -- .trim() only strips the string's outer edges, not a
    \r sitting mid-string before the first bare \n. Now splitLines(...)
    [0].
  - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed.
    [^\n]* in targetBulletPattern swallowed a line's trailing \r on
    CRLF input, shifting the computed insert position to land INSIDE
    the \r\n pair; combined with a hardcoded '\n' bullet separator, a
    CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an
    insert. Fixed with two coupled changes (either alone still
    corrupts, verified both ways): [^\r\n]* in the pattern, and the new
    bullet's leading terminator now comes from detectEol(rawContent).
  - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan):
    investigated, genuinely safe, left untouched. The .search(/\n#{2,4}
    .../) boundary-finder and the [^\n]*-based heading match were
    empirically verified on a 3-phase CRLF fixture -- the only stray \r
    ends up at the tail of an intermediate phaseSection string that is
    only ever used for .test()-based idempotency checks, never for an
    exact-match comparison or written back to disk. No corruption on
    round-trip.

Every fix re-verified: npm run build:lib clean, npx eslint
'src/**/*.cts' --no-cache reports 0 problems (was 13), and each
fixed function's existing LF-input tests were spot-checked unchanged.

* fix(#3413): apply orthogonal review findings

Two isolated review engines (correctness + security) ran against the
full diff and found three majors, one real security issue, and several
disclosure-worthy minors. All fixed or explicitly disclosed with
evidence; nothing deferred.

MAJOR — detectEol's tie-break contradicted its own documented contract.
Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc,
CONTEXT.md, the function's own comment) says ties resolve to '\r\n'.
The existing test masked this by reusing the same tie fixture the
buggy code happened to satisfy, rather than a genuine LF-majority
case. Root cause: an Edit attempted earlier in this phase to fix this
exact arithmetic error was blocked by the tier guard, and a later
dispatch was incorrectly told it had already landed. Fixed: condition
is now crlfCount >= bareLfCount; the test fixture corrected to a
genuine 2:1 majority, with a new explicit tie-case test.

MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via
detectEol(rawContent), justified by a comment claiming a hardcoded
'\n' corrupts a CRLF ROADMAP.md. False: this write goes through
platformWriteSync, whose normalizeContent/_normalizeMd unconditionally
converts CRLF->LF for any .md target — the templating was inert dead
code, erased before the file is ever written. Reverted to hardcoded
'\n', comment corrected to state the true reasoning. The separate
[^\n]* -> [^\r\n]* widening one function up (a real splice-position
fix, independent of final EOL) was kept.

MAJOR — roadmap-upgrade.cts's stated rationale for switching onto
splitLines/joinLines was wrong (both functions always agreed on line
representation, before and after — the claimed equality-check risk
never existed), and the change it justified introduced a real
regression: forcing every line onto one dominant terminator silently
rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This
write path uses raw fs.writeFileSync, not platformWriteSync, so unlike
the phase.cts case above the regression is genuinely live.

Fixing this took two attempts. The first attempt (revert to
split('\n')/join('\n') plus a suppression comment) was correctly
blocked by an agent that discovered local/no-crlf-fragile-split is a
PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs —
a hard, out-of-band, ADR-1703-governed guardrail banning any
eslint-disable of this rule anywhere in src/**/*.cts. That agent also
detected and correctly disregarded an injected instruction that
appeared in tool output during a git operation, per this session's
untrusted-content policy. The actual fix: computeMigrationPlan
reverted to roadmapContent.split('\n') (confirmed lint-clean — the
rule's data-flow tracking only follows a variable's initializer, and
this one is declared empty then reassigned in a try block).
applyMigration's write-back now splices edits against the ORIGINAL
content string via indexOf('\n', pos) boundary-walking instead of a
full split/rejoin, so every untouched character — including every
line's own terminator — is copied byte-for-byte. A capture-group split
(/(\r\n|\n)/, preserving terminators inline) was tried first and
empirically confirmed to still trip the rule before this approach was
chosen instead.

MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used
the STRING form of String#replace, so $&, $`, $', $1-$9 inside
must_haves.truths content (author-controlled) were interpreted as
replacement directives, splicing unrelated ROADMAP.md text into the
result. Fixed with the function-replacement form, which is never
pattern-interpreted. Verified before/after with the reviewer's exact
repro.

Also disclosed rather than silently left: test matrix row 31 (four
planned CRLF-materialized regression tests) was never implemented as
separate files — corrected to record the actual verification (a
manual --check run plus incidental existing coverage via each script's
normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF
behavior was claimed byte-for-byte unchanged but the old
yaml.indexOf(blockMatch[0]) substring search could match an unrelated
earlier occurrence of the header text (e.g. inside a quoted value) —
the split-then-scan fix incidentally also closes this, a strict
improvement now recorded in the design doc rather than left implicit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error

Checkpoint 2 came back red with 5 failures on the reviewed sha, both
gaps genuinely undetectable by any local gate.

eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs'
ignores-list entry (ADR-457: generated .cjs artifacts are excluded from
direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits
two lines above it and was the exact precedent read while researching
the six-gate ripple for this module -- missed anyway. Caught by
tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which
only runs on the remote suite.

tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both
`messageId` and `message` on the same RuleTester error assertion --
ESLint's RuleTester rejects that combination outright. This existed
since the test was first authored and was never caught locally: `npx
eslint` only lints the file's syntax, it does not execute RuleTester,
and local `node --test` is hard-blocked in this repo -- the assertion
had never actually RUN before this checkpoint. It was even present in
checkpoint 1's failure list, listed there as one of the "expected RED"
tests; I matched it against my expected-failures list by test NAME
only and never inspected the actual failure detail closely enough to
notice it was failing for the wrong reason (a RuleTester config error,
not the intended message-text mismatch). Fixed by keeping `message`
(the exact-text assertion the test exists to make) and dropping
`messageId`. Verified the crlfFragileSplit message string in
eslint-rules/no-crlf-fragile-split.cjs matches this assertion
character-for-character, and swept every other invalid case in the
file for the same double-specification bug (none found -- all
pre-existing cases use messageId alone).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix

The sole user-visible effect of this phase. No breaking-change label
or Changed fragment needed — ADR-3212's Backward Compatibility section
names the Node floor (Phase 1, already shipped) as the epic's only
breaking change; Phase 2 has none.

* chore(#3413): backfill changeset pr number to 3420

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 20:27:48 -04:00
BeeHiggs
9a305f3b6b fix(#3367): expectedQuickId derives UTC values to match the TZ-pinned subprocess (#3398)
expectedQuickId() in tests/init.test.cjs computed its expected quick_id
using local-time Date getters (getFullYear/getMonth/getDate/getHours/
getMinutes/getSeconds) in the test-runner process, while the CLI
subprocess under test is pinned to TZ=UTC. The two sides only agreed
when the test-runner's own ambient TZ happened to already be UTC.

Swap the six getters to their getUTC* equivalents so the helper is
timezone-invariant. No change to src/init.cts (the CLI side already
relies on its own UTC-pinned subprocess env and is correct) and no
change to the pinned clock instants or expected quick_id strings
asserted by the three affected tests.

Introduced by 80734a969 (#3332).

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-13 18:22:48 -04:00
Tom Boucher
dc3c81e93d chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam

Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts
and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both
suites fail with MODULE_NOT_FOUND, which is the intended RED.

Locks the measured behavior rather than the assumed behavior:
RegExp.escape hex-escapes the leading character of nearly every string
("abc" -> "\x61bc"), so the suite asserts match-equivalence against an
inlined historical oracle (the implementation being deleted) rather
than byte-equivalence of pattern text — 200 seeded fast-check runs plus
a fixed corpus, 0 mismatches. Also locks the latent character-class
range bug this phase fixes as a side effect: a hyphen-bearing value
interpolated into [...] currently forms a real range and matches an
unintended character; post-migration it must not.

* chore(#3412): src/pattern.cts owns runtime-value regex construction

Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam
delegating to the built-in RegExp.escape, deletes every hand-rolled
copy, and raises the Node floor to the Active LTS line.

The census was low, three times over. ADR-3212 counted 10 copies; a
graph query found 12; the new lint rule — once live — found 27 more.
The difference is that the census counted named helper FUNCTIONS while
the rule counts the escape SHAPE, so inline .replace(<class>, '\$&')
copies were never in scope. ADR §1's actual requirement is that no
module outside the seam escapes a value for regex use, so all of them
are, and CLAUDE.md's no-defer rule makes them this change's work.
Fourth consecutive epic here whose copy count was low — the argument
for ADR-3180 Amendment 3's "state N found by the guard" rule.

Also corrected mid-implementation: the survey reported phase-id.cts's
escapeRegex had 0 external importers. It had 8 production importers,
making its removal a public-surface change to an ADR-2121-owned module
and requiring an update to that ADR's locked-surface test. Blast
radius revised Medium-High -> High.

RegExp.escape is match-equivalent but NOT text-equivalent: it
hex-escapes the leading char of nearly every string ("abc" ->
"\x61bc"). Equivalence is proven by a seeded fast-check property test
against the deleted implementation as oracle. It also fixes a latent
bug: a hyphen-bearing value interpolated into a character class
previously formed a real range and matched an unintended character.

Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines,
.nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate
`required-tests` context is unchanged and no job was added or removed,
so branch protection cannot be orphaned by the dropped lanes.

Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with
structural provenance for reviewed pattern-fragment constants rather
than a name heuristic) plus a whole-tree companion guard covering the
directories ESLint's globs miss.

* fix(#3412): close the _SOURCE guard evasion, correct two false claims

Three findings from the orthogonal review pass, all fixed.

1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier-
   name matching with no binding check, so `new RegExp(userInput_SOURCE)`
   — a function parameter — sailed past the guard. That is the same
   rename-evasion class issue #3410 documents, reopened by the very
   fallback meant to complement the structural check. Now bound to the
   identifier's actual binding kind: import, require-derived const, or
   module-scope const; parameters, `let`/`var`, and unresolvable
   bindings fail closed. Four RuleTester cases cover the evasion and
   prove the legitimate cross-module case still passes.

2. src/pattern.cts's own header carried the stale pre-correction counts
   (12 copies / 17 call sites) while CONTEXT.md and the design doc
   carried the corrected ones (~39 / ~44) — a self-contradiction inside
   the PR whose entire purpose is deleting divergent copies. Rewritten,
   preserving the durable lesson: a named-function census cannot see
   inline copies; only a shape-matching guard can.

3. The claim that all deleted copies threw TypeError on non-string was
   false. phase-id.cts's copy — the one with 8 external importers — did
   String(value).replace(...) and never threw. The seam's locked
   signature does not coerce, so this is a real, now-disclosed behavior
   change rather than the pure preservation the tests asserted. Audited
   all 32 invocations across the 8 importers and 6 in-file callers:
   every one is safe by construction (upstream truthy guard or a
   string-producing derivation), verified by runtime probe against the
   compiled modules rather than by TS compilation, which cannot see a
   runtime undefined. Corrected the false claim in both the test comment
   and the design doc, and added it to Known limits.

* docs(#3412): add Changed changeset for the Node 24 floor

The only user-visible break in this phase. The escape-behavior change
is internal and match-equivalent, so it carries no user-facing note.

* fix(#3412): resolve the seam's require graph in script fixtures and packaging

Checkpoint 2 came back red with 90 failures on the node24 lane. Three
distinct defects, all introduced by routing scripts/ through the new
pattern seam, none reproducible by any local gate:

1. ~82 failures — tests/adr-index-gate.test.cjs and
   tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an
   mkdtemp fixture and spawn it there (necessary: those scripts resolve
   their scan root from __dirname/.., so running the real script would
   scan the real repo). Each harness hand-listed the dependencies to
   copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to
   gen-adr-index.cjs made both lists silently incomplete ->
   MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON'
   failures from the same crash.

   Fixed as a class, not an instance: new tests/helpers/copy-script-
   fixture.cjs walks a script's transitive static relative-require graph
   and copies it, so dependencies are derived and never re-declared. It
   throws (naming the unbuilt artifact) instead of letting the child die
   with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming
   scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host-
   contract, sync-runtime-launcher.

2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so
   the new scripts/lint-no-adhoc-regex-escape.cjs would be
   MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from
   the tarball, matching the existing precedent for gen-emitted-
   baseline.cjs, which is excluded for the identical reason, and locked
   with a test modeled on that one. Confirmed against a real npm pack:
   890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs
   present (so the other four scripts' requires are legitimate).

3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped
   source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to
   the retired hand-rolled escaper but NOT text-equivalent: it hex-
   escapes the leading character and all hyphens ('0*\x329',
   '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match
   decisions across all three real interpolation prefixes, zero
   divergence. Those tests now compile each source into the same heading
   regex src/roadmap.cts's searchPhaseInContent builds and assert what
   matches and what does not, including the 'i'-flag canonicalization
   the hex escape has to preserve. Re-pinning the new literals would
   have rebuilt the same brittleness one layer down. Adds a test for the
   property the escape exists for: a dot in '1.2' must not act as a
   wildcard.

Also shares one definition of 'a require' between the packaging guard
and the fixture copier, so the two cannot disagree about what they scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): refuse to copy a fixture dependency outside the fixture root

copyScriptWithDeps resolved each relative require and joined the
repo-relative result onto fixtureRoot. A require resolving OUTSIDE the
repo yields a '../'-prefixed relative path, so path.join climbed out of
the fixture and wrote into the surrounding temp dir (verified:
repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd).

No script in the tree does this today, so this closes an available
escape rather than an active one. Refuses via the existing unresolved-
require path so the failure names the offending specifier. Covered by a
negative proof that the guard fires and that nothing lands outside the
fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract

Applies all findings from the second orthogonal review round, re-run
because real code changed after round 1.

HIGH (security) — extractRequires stripped BLOCK comments before LINE
comments, so a '//' comment containing '/*' opened a phantom block
comment, and a '//' inside a string literal truncated the line. Both
hid real requires: 'const u="http://x"; require("./real.cjs")'
returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were
invisible. Replaced with a real AST parse via espree.

This is ADR-3212's own Decision 4 — tokenizer-first for stateful
grammars — applied to the case it describes; comment/string/regex
nesting is exactly such a grammar, which is why the regex version was
wrong. The function was moved byte-identical out of the #2858 packaging
guard, so the bug PRE-DATES this branch and has been a live blind spot
there: a shipped script could have required an unshipped path
undetected. Fixing it makes that guard strictly stronger than on next.

espree is promoted from a transitive eslint dependency to an explicit
devDependency rather than relying on hoisting. The script parse attempt
sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a
function, making a top-level return legal — scripts/check-coverage-gate
.cjs relies on it, and without the flag the guard throws on a file it
is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js
under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a
real npm pack, so the exact extractor does not newly fail the guard.

MEDIUM (security) — the repo-containment check guarded dependencies but
not the entry path. One escapesContainment predicate now guards both.

LOW (security) — containment was lexical while fs follows symlinks, and
a directory symlink could mint a fresh dedupe key per level. realpath
now resolves both repoRoot and each dependency before the decision, and
the realpath-derived path is the dedupe key. Destination layout still
uses the original repo-relative path, so copied trees are unchanged.

MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests
lost the foreign-prefix contract: every assertion was satisfied by an
impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599
bug class the exact-source prevents. The literal assertions it replaced
were catching this. Now asserts the compiled regex REJECTS a different
prefix with the same number.

MAJOR (standards) — the test hand-duplicated production's heading regex
with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed
the parallel surface instead of policing it: src/roadmap.cts exports
buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports
it. Byte-identical .source and .flags verified for both escaped forms.

MINOR — '..foo' no longer false-flagged as an escape; the inverted
spurious-vs-missing doc claim corrected; the dead allow-test-rule
header removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3412): backfill changeset pr number to 3416

* fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision

Two CI failures on PR #3416, both in code this branch added.

CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a
bracket run be consumed EITHER by the character-class branch OR one
character at a time by the trailing catch-all, so a failing match
explored both parses of every pair. Measured on the real regex:
n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script
scans repo source, so a file with a long bracket run after '.replace(/'
would hang CI outright — a guard against undisciplined pattern
construction was itself the worst pattern in the diff.

Fixed the way ADR-3212 already prescribes: the catch-all branch now
excludes '[' and ']' so a bracket can only be consumed by the class
branch (this is what makes it linear), and every quantifier is bounded
(the locked bounded-quantifiers decision) as a second line of defense.
Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the
constant: a regex literal with a BARE unescaped ']' outside a class is
no longer matched by this backstop. No census shape has that form, and
the AST rule remains the primary detector.

Verified the guard did not go blind doing it: a real census-shape
violation is still reported, and an allow-adhoc-regex-escape
suppression comment is still honored.

Regression test drives the exported findViolations on a
2000-repetition adversarial input and asserts the RESULT. It makes no
wall-clock assertion — elapsed-time tests are forbidden — so a
regression surfaces as a harness timeout, which is the correct signal.

Prompt injection scan — 'must not act as a regex wildcard' in a test
comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if|
my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a
whole test file over one phrase would blunt the scanner permanently,
and the comment has nothing to do with injection.

Neither failure was reachable from the remote runner — CodeQL and the
injection scan are not in that matrix, so the sha it passed was green
and still wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 16:19:57 -04:00
Tom Boucher
622c10b2c6 fix(#3025): refuse cross-runtime skill sync in sync-skills (#3404)
* fix(#3025): refuse cross-runtime skill sync in sync-skills

Skill content and directory layout are runtime-specific — the installer
applies per-runtime converters, adapter headers, brand swaps, and layout
rules at install time, and grok/gemini resolve to ANOTHER runtime's skills
root. A verbatim cross-runtime cp -r therefore produces content the installer
would never have written for the destination, and can damage a runtime the
user never named. #3024 (closed) un-masked this, making the corruption live.

Fix (option b, user decision): add a functional Step 1 guard that refuses
any --to != --from with an actionable installer pointer, before any
resolution or copy. Identity sync (--from == --to) remains a no-op. The
non-functional Step 5 comment is replaced; Arguments/Limitations updated.

Regression: tests/sync-skills-cross-runtime-refuse.test.cjs (source-text-
is-the-product) asserts the guard exits non-zero for cross-runtime, points
at the installer, precedes the cp -r copy, and preserves identity.

* docs(#3025): backfill changeset PR number (#3404)

---------

Co-authored-by: sim <sim@local>
2026-08-13 16:17:10 -04:00
Tom Boucher
6dbc124018 enhance(#3180): the sibling validators share one envelope and one owner — Phase 12 (#3407) 2026-08-13 11:31:16 -04:00
sim
3e70e57e37 fix(#3309): applyRepairs risk-gating test used a fake, non-writable cwd
CI caught what the bench run didn't: "row 12" (every NONE-risk action
must actually apply) called applyRepairs('/fake/cwd', ...) — a literal
path that doesn't exist on disk. This was fine when applyRepairs's
handlers were stubs (pre-migration skeleton), but real handlers now
read/write actual files: createConfig writes config.json,
addNyquistKey/addAiIntegrationPhaseKey read it before patching. Against
a genuinely non-existent path these now correctly fail (ENOENT), and
an earlier fix in this same PR (applied only receives a code on real
success) correctly surfaces that as a failure instead of masking it —
so 3 of 4 codes stopped landing in `applied`, deterministically, on
any environment that actually enforces ENOENT against /fake/cwd.

Uses a real temp project (createTempProject + a valid config.json)
instead. Row 11 (DESTRUCTIVE refusal) and the ADVISE-skip test are
unaffected — both paths return before any handler touches the
filesystem, confirmed by reading applyRepairs's dispatch order.
2026-08-13 07:22:07 -04:00
sim
ce0999bf84 test(#3309): update stale W002/W020 test expectations to match the fixes
Both tests were written against the pre-fix behavior and never updated
once the real fixes landed:

- state-consistency.test.cjs's "KNOWN GAP" test hardcoded the
  expectation that W002 incorrectly fires for a STATE.md phase
  reference whose only home is an archived milestone — that gap is
  now closed (0 diagnostics, confirmed against real buildPlanningSnapshot
  output), so the test is renamed and its expectation flipped.
- worktree-health.test.cjs's two W020 tests asserted the OLD single
  combined "timed out or failed" message/remedy — verified against
  the real pre-migration src/verify.cts:2204-2219 that git_timed_out
  and git_list_failed always had distinct messages; the fix that
  restored this distinction is correct, these tests just never
  caught up to it.
2026-08-13 04:53:08 -04:00
sim
7ddcc19823 fix(#3309): acknowledge health.md growth, fix stale doc-consistency tests
gsd-test found two independent gaps around the newly-generated
health.md tables:

- emitted-attribution.test.cjs requires an acknowledgment for
  health.md's 2271-byte growth (16-code hand-maintained table -> 34-row
  generated table, this phase's explicit acceptance criterion). Adds
  tests/emitted-drift-acks/3309-health-docs-generated.json. Removes
  2573-state-head-freshness.json's now-inert health.md entry (that
  fragment's growth already landed on origin/next, the diff base this
  branch is compared against, so it has nothing left to acknowledge —
  and the guard forbids two fragments naming the same path).

- runtime-converters.test.cjs's health.md content-consistency checks
  asserted stale text from the old hand-written table: a regex that
  false-positived on the new table's own unrelated W020 row (worktree
  scan degradation, a different diagnostic than the W025 isolation
  warning it was meant to detect), and anchors expecting the old
  table's exact last row / footnote wording. Narrowed the regex to
  require the literal use_worktrees config key, and updated the
  anchors to the new table's real shape (I001/I010 as the last rows,
  the new generated-table footnote).
2026-08-13 04:31:10 -04:00
sim
eae2b52e4a fix(#3309): W020 fires on any degraded worktree scan, not just real failures
gsd-test found buildWorktreeHealthField collapsed every
inspectWorktreeHealth failure reason (git_timed_out, git_list_failed,
not_a_git_repo) into one UNREADABLE scope, discarding which one. The
migrated checkW020 then warned unconditionally on any UNREADABLE
scope — but the original (verify.cts:2202-2217) only warned on
git_timed_out/git_list_failed, staying silent on not_a_git_repo (a
.planning/-only fixture with no git repo at all is not a degraded
scan, just the absence of one). This spuriously degraded every test
fixture that isn't a real git repo.

planning-snapshot.cts's worktreeHealth field now carries `reason`
through instead of discarding it; checkW020 branches on it exactly
like the pre-migration code did.
2026-08-13 04:30:57 -04:00
sim
96c7ea9b35 fix(#3309): W023's message drops each colliding directory's status
gsd-test found the migrated W023 dropped a piece of information the
original message included: each colliding phase directory's overall
status (e.g. "Complete"), not just its raw plan/summary/verification
counts. Adds derivePhaseStatusLabel, reconstructing the status label
from already-exposed PhaseSnapshot fields (planCount/summaryCount/
complete/verificationStatus) — no new ambient I/O, no new snapshot
field.

Also fixes a non-conforming test fixture found while verifying:
tests/health-validation.test.cjs's "05-real" fixture used a bare
VERIFICATION.md, which readVerificationStatus never matches (the real
convention, and every other fixture in this repo, use the
*-VERIFICATION.md suffix) — the status was always reading as "missing"
regardless of message formatting. Renamed to 05-real-VERIFICATION.md.
2026-08-13 04:30:48 -04:00
sim
041414c4ad feat(#3309): generate health.md's error-code and repair-action tables
Closes the issue's explicit acceptance criterion: "health.md's tables
are generated rather than hand-maintained, closing the 16-vs-30+
documentation gap structurally." The published roster listed 16 codes
against 30+ actually emitted; W010-W017 and W020-W023 had never been
documented.

Adds description/repairable as static fields on Rule (health-diagnostic-types.cts)
— generation needs a fixed, human-readable summary per code, distinct
from the dynamic per-instance Diagnostic.message a rule's check()
produces. repairable is true only when --repair will actually apply
the remedy: false for ADVISE-only rules AND for DESTRUCTIVE-risk rules
(regenerateState/resetConfig), which are described but never
auto-applied — matches verify.cts's diagnosticToIssueEntry semantics
exactly, after fixing E004/E005's static field to agree with it (both
were wrongly true, an inconsistency caught during this same commit's
own review, not left for later).

New scripts/gen-health-docs.cjs (--write/--check, wired into
lint:generated-sync) regenerates the two tagged table regions in
gsd-core/workflows/health.md from RULES (31 rules) plus the 3
pre-checks that stay outside the rule table by design (E001, E010,
I010) plus a small static Effect/Risk lookup for the 6 real repair
actions — including addAiIntegrationPhaseKey, live in code since an
earlier phase but never documented until now. 34 error-code rows, 6
repair-action rows. The table's old "grep verify.cts for the next free
number" footnote is rewritten to point at the rule table and its lint
guard instead.
2026-08-13 03:13:25 -04:00
sim
4f9e5cf2ed fix(#3309): make the lint guard's W024 exemption explicit, not accidental
W024's committed rule (state-consistency.cts) is a documented permanent
no-op — its real check runs in cmdValidateHealth itself, outside the
rule table, since readStateHeadFreshness needs a git-log shell-out no
Rule.check may perform. The guard's §8.5 fixture-proof check previously
"passed" for W024 only because some test file's title happened to
contain the string "W024" — not because any fixture actually proves it
fires, which it structurally never can. Found by the Spec-axis
orthogonal review.

Adds an explicit PERMANENTLY_INERT_CODES map (currently just W024,
with its reason recorded) that checkFixtureProofInvariant reports
separately from real coverage. The guard's PASS output now says
"30 covered by a real fixture, 1 exempted" instead of implying uniform
proof — a code with no coverage and no exemption entry still fails.
2026-08-13 02:56:47 -04:00
sim
03258a07a2 fix(#3309): applyRepairs must not count a failed repair as applied
applyRepairs pushed a diagnostic's code onto applied unconditionally
after the try/catch around runRepairAction, even when the handler
threw (caught, recorded in details with success:false) or otherwise
failed — making applied mean "attempted" rather than "succeeded," with
no test exercising the failure path. Found by the Spec-axis orthogonal
review.

applied now only receives a code when the repair actually succeeded;
a failed attempt is still fully recorded in details (success:false,
the error message) but no longer misreported as applied. Adds a
regression test forcing addNyquistKey to throw (ENOENT on a config.json
that doesn't exist) and asserts it lands in details, not applied.
2026-08-13 02:56:30 -04:00
sim
1255f960db fix(#3309): restore W027's active-worktree exclusion
The migrated checkW027 (stale worktree) dropped the pre-migration
exclusion of the CLI's own current worktree, since a Rule.check(snapshot)
has no cwd access (§8.1 rule 1 forbids ambient I/O) — flagged as a
disclosed regression during this phase's own design work, then
confirmed as a real, fixable gap by the Spec-axis orthogonal review
rather than an inherent limitation.

Fixes it properly instead of accepting the regression: buildPlanningSnapshot(cwd)
already receives cwd as its own input, so exposing it as snapshot.cwd
is not new ambient I/O, just surfacing an existing parameter — fully
consistent with §8.1 rule 2's "parsed value" allowance. checkW027 now
excludes the entry matching snapshot.cwd before flagging, matching the
original verify.cts:2233-2242 behavior exactly.
2026-08-13 02:56:23 -04:00
sim
42729b21fa fix: scan bin/lib subdirectories in the inventory-manifest generator
gen-inventory-manifest.cjs's cli_modules family did a flat readdirSync
of gsd-core/bin/lib/, invisible to anything shipped in a subdirectory.
Found while registering this phase's 8 health-diagnostic-rules/*.cjs
files in docs/INVENTORY.md (Standards-axis review) — the automated
manifest cross-check couldn't see them even though the manual
INVENTORY.md rows were correct.

Adds collectOneLevelSubdirs (mirrors the existing collectNested's
defensive statOrNull style) and merges flat + one-level-subdirectory
results into cli_modules's single sorted array, using the same
<subdir>/<file>.cjs key format INVENTORY.md's rows already use.

Regenerating the manifest surfaced that three OTHER existing
subdirectories (installer-migrations/, host-integration-adapters/,
observability/ — pre-existing, unrelated to this phase) were equally
invisible and had zero docs/INVENTORY.md rows at all. Added all 15
missing rows rather than leave a gap the fix itself just exposed.

Also fixes 3 pre-existing lint-legacy-dir-name violations in the
installer-migrations rows (legitimate references to the historical
get-shit-done -> gsd-core rename these migrations clean up — marked
with the guard's own gsd-allow-legacy-name exemption) and a stale
health-diagnostic.cjs row that still said "RULES ships empty."
2026-08-13 02:56:12 -04:00
sim
d1760e3c31 refactor(#3309): migrate cmdValidateHealth onto the rule table
Replaces cmdValidateHealth's hand-rolled addIssue/switch accumulation
(961 lines) with buildPlanningSnapshot -> evaluateRules -> map to the
legacy {code, message, fix, repairable} shape, bucketed by severity.
Two pre-checks (home-dir E010/I010, .planning/-root-missing E001) stay
outside the rule table entirely, per ADR-3180 §8.2 rule 4 ("no
precedence system") — building "some rules suppress others" into the
table would itself be the forbidden precedence system.

W024 (STATE.md commit-age freshness) also stays outside the table:
its committed rule is a documented permanent no-op (readStateHeadFreshness's
git-log shell-out is ambient I/O a Rule.check may never perform, and no
PlanningSnapshot field carries a commits-behind count). Migrating onto
the rule table as designed would have silently regressed 7 passing
tests in tests/health-validation.test.cjs — found while wiring this
function, kept as a real check in the wrapper instead (same I/O
license applyRepairs already relies on), fixed inline per this repo's
no-defer policy rather than accepted as a silent loss.

Ports the real repair-handler bodies (createConfig/resetConfig,
regenerateState, addNyquistKey/addAiIntegrationPhaseKey,
backfillMilestones) into health-diagnostic.cts's applyRepairs,
replacing the skeleton's stub. DESTRUCTIVE-risk remedies
(resetConfig/regenerateState) are refused by --repair — a disclosed
breaking change; repairable now means "an automatic repair will
actually run," not merely "a remedy exists to describe," so E004/E005
now report repairable:false. --backfill alone now actually triggers
backfillMilestones, fixing a latent bug where its gate was unreachable
without --repair also being set (verify.cts:2504, confirmed dead code
pre-migration).

Test updates distinguish the two explicitly-authorized behavior
changes (DESTRUCTIVE refusal, backfill-alone fix, W021->W026 split)
from preservation — every changed assertion is commented with why, and
new regression tests were added for both changes plus W021/W026
mutual independence. Drift-guard bookkeeping (bypass-baseline shrunk
to the one disclosed W024 exception, milestone-window and
phase-enumeration exemptions, test-file-count allowlist) updated for
the relocated/new functions this migration introduces.
2026-08-13 02:28:49 -04:00
sim
acc1a7abd6 feat(#3309): add health-diagnostic rule-table lint guard
Enforces ADR-3180 §8.2's 1:1 rule-code invariant (every code unique,
every severity a property of the Rule) and §8.5's fixture-proof
invariant (every code has a describe()/test() block naming it,
verified statically against tests/health-diagnostic-rules/*.test.cjs
and tests/health-diagnostic.test.cjs) for the new RULES table.

Adapted from the design doc's original plan of separate
tests/fixtures/health-diagnostic/<code>.* files: implementation used
inline temp-dir fixtures instead (mirrors tests/planning-snapshot.test.cjs),
so coverage is checked statically against test-file structure, mirroring
lint-fix-has-regression-test.cjs's house style. Wired into lint:ci
adjacent to lint-planning-snapshot-bypass-drift.cjs, its closest sibling.

Passes clean against the real tree: 31 codes, all unique, all covered.
2026-08-13 01:53:20 -04:00
sim
c469aeeacd refactor(#3309): add milestone-archive-hygiene health-diagnostic rules
W018, W019 — MILESTONES.md archive completeness and unrecognized root
.md file checks, migrated onto the frozen rule table per ADR-3180 §8.2.
2026-08-13 01:37:42 -04:00
sim
a56679a4ca refactor(#3309): add worktree-health health-diagnostic rules
W020 (3 conditions), W017, W027 — git worktree list degradation,
orphan and stale worktree checks, migrated onto the frozen rule table
per ADR-3180 §8.2. W027 has a documented fidelity reduction: rules
have no cwd access, so it can no longer exclude the active worktree.
2026-08-13 01:37:42 -04:00
sim
9e74f00ba0 refactor(#3309): add roadmap-disk-consistency health-diagnostic rules
W006, W007 — ROADMAP entries with no matching disk dir, and disk dirs
with no ROADMAP entry, migrated onto the frozen rule table per
ADR-3180 §8.2.
2026-08-13 01:37:42 -04:00
sim
f51f00ef0e refactor(#3309): add agent-install health-diagnostic rules
W010 — agent install completeness across its 4 internal conditions,
migrated onto the frozen rule table per ADR-3180 §8.2.
2026-08-13 01:37:42 -04:00
sim
1b09fb13d3 refactor(#3309): add phase-structure health-diagnostic rules
W005, W023, I001, W009 — phase directory naming, duplicate phase keys,
plan-summary presence, and validation-architecture checks, migrated
onto the frozen rule table per ADR-3180 §8.2.
2026-08-13 01:37:33 -04:00
sim
80484249df refactor(#3309): add config-validation health-diagnostic rules
W003, E005, W004, W008, W016, W012, W013, W014, W015, W022 — config.json
existence, parseability, and field-validity checks, migrated onto the
frozen rule table per ADR-3180 §8.2.
2026-08-13 01:37:33 -04:00
sim
8a7ef78906 refactor(#3309): add state-consistency health-diagnostic rules
W024 (deliberately inert, no snapshot field yet for stale state_head),
W002, W011, W021, W026 — STATE.md cross-checks against ROADMAP/config,
migrated onto the frozen rule table per ADR-3180 §8.2.
2026-08-13 01:37:33 -04:00
sim
7195308384 refactor(#3309): add root-existence health-diagnostic rules
E002, E003, E004, W001 — PROJECT.md/ROADMAP.md/STATE.md existence and
PROJECT.md section-completeness checks, migrated onto the frozen rule
table per ADR-3180 §8.2.
2026-08-13 01:37:24 -04:00
sim
8c9ca3f7c6 refactor(#3309): add allPhaseDirNames field to planning-snapshot
W007 (orphan disk dir with no ROADMAP entry) cannot be sourced from
phaseDirs, which is windowed to ROADMAP-declared phases only — an
orphan dir can never appear in an already-ROADMAP-filtered set. Adds
an unwindowed allPhaseDirNames field so the rule can actually fire.
2026-08-13 01:37:01 -04:00
sim
c5543e533c refactor(#3309): extend planning-snapshot.cts with 8 more parsed fields
Phase 11 of epic #3180 (ADR-3180 §8.1 rule 2). PlanningSnapshot grows from
7 fields to 15: projectSections, statePhaseTokens, stateStatus,
roadmapDeclaredPhases, roadmapPhaseCheckboxes, researchValidationStatus,
milestoneArchiveStatus, planningRootFiles.

Every field is a reused owner (buildRoadmapPhaseVariants/
buildNotStartedPhaseVariants from src/validate.cts, stateFieldValue) or a
small relocation of already-working verify.cts logic (PHASE_NUMBER_TOKEN_SOURCE
scanning, the checkMilestonePrefixMismatches sectionRx walk, W009/W018's
file-existence checks) — never a new algorithm, and never raw document text:
§8.1 rule 2 forbids exposing raw text, not exposing a parsed list or boolean
derived from it once by the snapshot builder.

roadmapPhaseCheckboxes deliberately reads the same ROADMAP checkbox
isPhaseComplete (§7.4, disk-strict) refuses to consult — that owner decides
completion and must not read it; this field only exposes what the checkbox
says, for a diagnostic (W011) whose whole purpose is flagging disagreement.
Not a re-derivation of §7.4, recorded explicitly to prevent that reading.

Adds PROJECT_UNREADABLE to UNUSABLE_REASON (ninth #1879 site), closing a
gap the implementing agent correctly flagged rather than silently leaving
absent-vs-corrupt collapsed for PROJECT.md, matching the STATE_UNREADABLE/
CONFIG_UNREADABLE precedent from this same effort's prior commits.

currentPhaseLabel/statePhaseTokens/stateStatus share one STATE.md read
(buildStateFields) rather than three independent reads.

Additive only — all prior fields and worstScope/buildPhaseSnapshot
unchanged.
2026-08-13 01:10:53 -04:00
sim
ef10bba707 refactor(#3309): add health-diagnostic.cts skeleton (types + evaluator)
Phase 11 of epic #3180 (ADR-3180 §8.2/§8.3/§8.5). New src/health-diagnostic.cts:
SEVERITY/REMEDY_ACTION (7 members: 6 real repair actions + ADVISE)/REMEDY_RISK
(NONE/DESTRUCTIVE) frozen enums, Diagnostic/Remedy/Rule types, an empty RULES
table (rules land in the next commits), evaluateRules (with a duplicate-code
defense-in-depth check ahead of the lint guard), and applyRepairs (the
DESTRUCTIVE-risk-refusal dispatcher — §8.3 rule 3 — with stub handlers; real
repair bodies port in the migration step).

Six-gate .cts ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md +
manifest, CONTEXT.md glossary entry.
2026-08-13 00:51:12 -04:00
sim
6aa378b261 refactor(#3309): extend planning-snapshot.cts with config/agentInstall/worktreeHealth
Phase 11 of epic #3180 (ADR-3180 §8.2/§8.3/§8.5) foundation. Extends the
already-merged Phase-10 PlanningSnapshot additively with three fields the
upcoming health-diagnostic rule table needs and Phase 10 never required:

- config: {value, scope, exists} — parsed .planning/config.json. `exists`
  distinguishes absent (no diagnostic, non-answer) from present-but-invalid
  (CONFIG_UNREADABLE diagnostic, corruption) — both collapse to scope
  UNREADABLE, so a rule needs the extra bit to tell "not configured yet"
  apart from "config.json is broken."
- agentInstall / worktreeHealth — not .planning/-sourced, wrap the existing
  checkAgentsInstalled/inspectWorktreeHealth owners with the same arguments
  cmdValidateHealth already passes them, so a later migration step reads
  these fields instead of calling the owners itself.

Adds CONFIG_UNREADABLE to src/unusable-input.cts's UNUSABLE_REASON (eighth
#1879 site), mirroring STATE_UNREADABLE's exact shape from Phase 10.

Additive only — the four Phase-10 fields and worstScope/buildPhaseSnapshot
are unchanged; existing tests for them are untouched.
2026-08-13 00:41:39 -04:00
Tom Boucher
3c4df10a50 Merge pull request #3402 from open-gsd/refactor/3308-planning-snapshot-parsed-projection 2026-08-13 00:12:48 -04:00
sim
2bd1d12369 fix(#3308): compare guard-produced file fields against POSIX-normalized paths in Windows CI
PR #3402's real Windows CI (windows-latest node22/24) caught what gsd-test's
Linux-only lanes structurally cannot: tests/planning-snapshot-bypass-drift.test.cjs
compared the guard's own POSIX-normalized output (findSnapshotBypassDrift's
`file` field, dedupeViolationsForBaseline/sortEntries entries, a written
baseline read back from disk) against REGISTERED_FILE, which is built via
path.join('src', 'verify.cts') and is therefore backslash-separated on
Windows. The guard always normalizes its OUTPUT to POSIX via toPosixRel
regardless of the input separator form, so the comparison only ever
coincidentally passed on POSIX.

Adds REGISTERED_FILE_POSIX for every assertion against a guard-PRODUCED
value (including baseline fixtures fed into diffAgainstBaseline, which are
matched by exact string key against the guard's normalized output).
REGISTERED_FILE itself is unchanged and still used, correctly, everywhere it
is the relPath INPUT to findSnapshotBypassDrift or a DIAGNOSTIC_RULE_FUNCTIONS
Map-key lookup — both need the platform-native form to match the guard's own
Map key, which is also path.join-constructed.

No behavior change on POSIX (both constants are byte-identical there).
2026-08-12 22:54:30 -04:00
Tom Boucher
2283168612 fix(#3170): anchor milestone one-liner extraction to a summary-shaped heading (#3401)
* test(#3170): extractOneLinerFromBody must anchor to a summary-shaped heading

Regression for #3170: the function matched the first heading's first bold
run, so an incidental first heading (rule list, deviation notes) contributed
its bold text as the milestone accomplishment. Rows 1/2 fail RED on next; rows
3/4 guard Overview recognition and the #2660 Summary-heading form.

* fix(#3170): anchor milestone one-liner extraction to a summary-shaped heading

extractOneLinerFromBody matched the first heading's first bold run regardless
of section, so an incidental first heading (rule list, deviation notes)
contributed its bold text as the milestone accomplishment written into
MILESTONES.md. Iterate headings and extract from the first Summary/Overview/
Accomplishments one with a bold run, falling back to null when none exists.
The #2660 Summary-heading forms and the frontmatter one-liner precedence are
preserved.

* docs(#3170): add changeset

* test(#3170): align extractOneLinerFromBody unit fixtures with summary-heading contract

The core-utils unit fixtures used generic # Title headings encoding the old
'any first heading' contract; the #3170 fix anchors to a Summary/Overview/
Accomplishments heading (the function is summary-specific). Update the
heading text to Summary-shaped; the extraction assertions (bold, frontmatter
strip, colon-label, CRLF, unicode) are unchanged.

* docs(#3170): backfill changeset PR number (3401)

---------

Co-authored-by: sim <sim@local>
2026-08-12 22:21:53 -04:00
sim
21c46ecf52 fix: scope safe.directory ownership bypass to the real-repo-root git ls-files call
Found while running gsd-test for #3308: tests/commit-files-pathspec.test.cjs's
repo-wide `--files` scan runs `git ls-files -z -- *.md` directly against the
checked-out repo root (not a createTempGitProject() fixture, unlike every
other gitOrThrow call in this file). Inside a container-provisioned test
runner the checkout's on-disk owner can legitimately differ from the running
UID, tripping git's CVE-2022-24765 dubious-ownership guard and failing the
scan closed (exitCode 128) rather than reporting a real file-list result —
reproduced on gsd-test's linux-node22 and linux-node24 lanes.

Adds `-c safe.directory=*` to that ONE invocation only, so the bypass is
scoped to this call rather than a global `git config` write that would leak
into every other git call in the process.

No source behavior changed; test-infrastructure resilience only.
2026-08-12 22:05:05 -04:00
sim
d1b659703e test(#3308): failing-first tests for planning-snapshot parsed projection
ADR-3180 epic #3180 Phase 10 (§8.1): tests for the not-yet-existing
src/planning-snapshot.cts (buildPlanningSnapshot, worstScope), the
not-yet-existing scripts/lint-planning-snapshot-bypass-drift.cjs guard,
and the new STATE_UNREADABLE reason on tests/unusable-input.test.cjs's
already-shipped UNUSABLE_REASON enum. RED by construction: the modules
under test do not exist yet.
2026-08-12 21:43:25 -04:00
Tom Boucher
b8cb031ce2 fix(#3163): scope phase.add insertion to the current milestone (#3400)
* test(#3163): phase add must insert in the active milestone, not the trailing archive

Regression for #3163: cmdPhaseAdd/cmdPhaseAddBatch pick the insertion point
via rawContent.lastIndexOf('\n---'), the file's last horizontal rule — which
on a roadmap with shipped/history material after the active phase list sits
deep in archive. Rows 1/2/4 fail RED on next (entry lands after the archive
heading); row 3 guards the no-milestone legacy fallback.

* fix(#3163): scope phase.add insertion to the current milestone window

cmdPhaseAdd and cmdPhaseAddBatch picked the insertion point via
rawContent.lastIndexOf('\n---') — the file's last horizontal rule, which on
a roadmap with shipped/history material after the active phase list sits deep
in archive. Extract phaseEntryInsertOffset(rawContent, cwd): scope the search
to currentMilestoneRawRanges' primary window so the entry lands at the end of
the active phase list. Fall back to the legacy whole-file heuristic when no
current milestone resolves, preserving simple no-milestone roadmaps. Applies
to both cmdPhaseAdd and cmdPhaseAddBatch (identical expression); the decimal
insert path was already header-anchored and is untouched.

* docs(#3163): add changeset

* docs(#3163): backfill changeset PR number (3400)

---------

Co-authored-by: sim <sim@local>
2026-08-12 21:01:07 -04:00
Rezolv
3ff9a7ffcd fix(#2570): parse leading date from last_activity so stale_activity fires with a description suffix (#2571)
* fix(#2570): parse leading date from last_activity so stale_activity fires with a description suffix

templates/state.md prescribes `Last activity: [YYYY-MM-DD] — [What happened]`,
and gsd-core's own STATE.md mirrors that suffix into frontmatter. Date.parse on
the whole string returned NaN, and because staleActivity treats null as "not
stale" (fails open), the only idle/staleness detector never fired on any project
whose last_activity kept its description.

parseActivityTimestamp now reads the leading ISO date/time token when a
whole-string parse fails, validating the calendar date (ADR-227: reject an
impossible date rather than let Date.parse roll it forward) and preferring the
whole-string parse when it succeeds so a trailing zone name is not dropped.

Composes with #3099 (LAST_ACTIVITY_UNPARSEABLE diagnostic), which merged to next
after this branch: both key off parseActivityTimestamp === null, so a value whose
leading date now parses takes the stale path and does NOT emit the diagnostic. A
regression test in tests/smart-entry.unit.test.cjs asserts exactly that (stale
true, emission count 0), guarding against two staleness signals on one field.

Rebased onto next (flattened): resolved the add/add test conflict by keeping both
the #2570 and #3099 describe blocks. Tests: unit + property, 80 pass.

* fix(#2570): fail open when a named zone can't be reconstructed from the token (#2571 B1)

The 2026-08-08 flatten dropped the zone handling earlier rounds built, so the
fallback path -- reached only when a description suffix makes the whole-string
parse fail, the #2570 case -- reconstructed `${date}${time}` WITHOUT any named
zone. ISO_LEADING_RE's offset group captures only Z / +-HH:MM, so " GMT"/" EST"
land in the un-captured suffix; Date.parse then read the reconstruction as LOCAL
time, shifting the instant by the host's UTC offset -- a wrong, host-dependent
value the diff's own comment warned against but guarded only on the other branch.

Fix (the simpler of the two offered in review): when the remainder after the
matched token begins with a letter (a named zone we cannot preserve), return
null -- fail open to not-stale, matching the base's honest behaviour and
ADR-227's "never propagate a wrong instant". The #2570 template suffix
(" -- description") starts with a separator, so it still reconstructs and reads
stale as intended.

Tests (both fail-first, verified RED on the pre-fix head):
- smart-entry.unit: a named-zone + description suffix (54 days old) enters the
  fallback and must read not-stale, not a still-old local instant. Host-
  independent by construction.
- smart-entry.property (f): named-zone + suffix over 1-week..1-year ages and 8
  zones stays total and fails open.

Discloses the removal M2 flagged: TRAILING_ZONE_RE / UTC_ZONE_NAMES /
timeCarriesOffset were dropped by the flatten; this restores the SAFETY (no
wrong instant) via the simpler null contract rather than the allowlist.

* fix(#2570): narrow the stale_activity fallback guard to a zone-designator shape

The round-9 fail-open guard `/^\s*[A-Za-z]/` treated any letter-led remainder as
an unpreservable named zone, so a leading real date followed by a bare
space/tab/colon and an ordinary description (a hand-edited STATE.md that omits the
template em dash) returned null and re-opened #2570 for exactly those shapes.

Narrow the guard to ZONE_DESIGNATOR_RE -- a standalone short all-caps run -- and
consult it ONLY when the leading token captured a time-of-day: a zone qualifies a
clock time, so a bare date carries no zone hazard and always reconstructs to its
UTC midnight. A plain description (including one that opens with a tech acronym
like "CI green") reconstructs; a real named zone on a timed value (GMT/EST/...)
still fails open (ADR-227: never propagate a wrong, host-dependent instant).

Widen the property generator to the non-em-dash separators (space/tab/colon), the
arm that structurally could not reach the fallback before, and add unit cases for
whitespace/tab/colon-separated and bare-date+acronym descriptions. All fail-first
on the prior guard; green across UTC/LA/Tokyo/Kiritimati.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-12 20:46:12 -04:00
0xdhx
6950ae3679 test(#2269): repo-wide regression guard for query commit --files scoping (#2290)
* test(#2269): repo-wide regression guard for query commit --files scoping

regression protection that fix does not carry: a repo-wide scan, edge-case
pins, property tests, and a behavioral test.

- Repo-wide scan across all five directories that carry live invocations
  (gsd-core/workflows, gsd-core/references, agents, commands, skills), so a
  future unscoped `commit` / `query commit` site fails CI wherever it lands.
  Backslash-continued lines are joined first, and a quote-parity walk
  distinguishes a real `--files` flag from one mentioned inside the quoted
  commit message.
- Edge-case pins for the shapes the live content does not exercise: the
  `query`-less spelling, a flag preceding the command (`--cwd`), a prose
  mid-sentence mention, and a `commit_docs` JSON-key false positive.
- Three `fast-check` properties over the quote-parity logic, following the
  tests/adr-parser.property.test.cjs precedent.
- A behavioral test that stands up a real temp git repo, leaves an unstaged
  `.planning/` stray plus an unrelated staged file, runs the real
  `gsd-tools commit --files`, and asserts via git status/diff that only the
  intended artifact landed. The scope is derived from secure-phase.md's own
  commit line, so reverting that line's `--files` fails this test too.

Census over the five roots at this base: 87 invocations, 0 unscoped.

* test(#2269): scan mid-prose argument-bearing invocations too

The repo-wide guard's INVOCATION_RE is line-start-anchored, which keeps
bare backtick mentions out of scope but also blinded the scan to fully
argument-bearing `query commit` invocations embedded mid-sentence in
instructional prose. Three such live invocations sit inside the scan's
own roots today (new-milestone.md, new-project.md, plan-phase.md) — all
scoped, but never entering the candidate set, so trimming their
`--files` clause would reintroduce #2269 behind a green suite.

Add a second tier: MIDLINE_INVOCATION_RE matches the invocation token
anywhere a quoted commit message follows (`commit "` — the executable
shape a bare mention never carries), and invocationCandidates() extracts
the backtick-bounded invocation substring so hasScopedFiles's
quote-parity walk is not skewed by surrounding prose quotes.

Census after widening: 87 anchored + 3 mid-line = 90 invocations,
0 unscoped; the mid-line tier picks up exactly the three cited sites
and nothing else across all five scan roots.

* test(#2269): align the --files value predicate with the runtime's flag filter

The scan's --files value test was /--files\s+\S/, which scores
`--files --amend` as scoped because `-` is \S. routeCommit disagrees:

    args.slice(filesIndex + 1).filter(a => !a.startsWith('--'))

drops every `--`-prefixed token, so `--files --amend` yields files=[] and
lands on the same unscoped default branch as a trailing bare `--files` —
#2269 verbatim. Two live sites sit one token-deletion from the shape:
gsd-core/references/git-planning-commit.md and
gsd-core/workflows/execute-plan.md both run
`... commit "" --files .planning/codebase/*.md --amend`.

The predicate now mirrors the runtime's own rule. Single-dash tokens stay
values, because the runtime filters on '--', not '-'.

Pinned: `--files --amend` and `--files --no-verify` as unscoped, plus the
two live `--files <glob> --amend` shapes and `--files -weird-name.md` as
scoped negative controls, so the fix cannot over-correct into "any
--files near a flag is unscoped" with nothing failing.

Also tightens the fast-check path generator, which is not cosmetic. It
excluded only ["\s], so it could draw a `--`-prefixed token and assert it
scoped: measured 26 hits in 200,000 draws (~1 in 7,700), i.e. ~1 CI run
in 77 at the default 100 runs would have failed as a mystery flake once
this fix landed. The complement is now its own property, and it fails
against the pre-fix predicate (counterexample ["","--#"]).

The behavioral test's --files presence check moves to the same predicate,
and its two failure modes are now separate messages: a genuinely missing
--files (the #2269 regression) versus an unquoted value that only breaks
this test's own scope derivation.

* test(#2269): score each invocation on a line separately, not the whole line

invocationCandidates returned [line] for any line-start match, so scoring
was satisfied by one --files anywhere on the line and a later scoped
invocation vouched for an earlier unscoped one:

    gsd_run query commit "a" && gsd_run query commit "b" --files x.md
    gsd_run query commit "a" ;  gsd-tools query phase-list --files y.md

Same "one hit satisfies the whole candidate" class as the --files value
bug in the previous commit. No live line has the shape today, so this is
pinned rather than left to be rediscovered.

A line-start invocation is now split at shell command separators before
scoring. The separator must be followed by a binary token, so a binary
inside a command substitution — `commit "$(gsd-tools query x)" --files
a.md` — is not treated as a second invocation; that negative control is
asserted, since the obvious "split at every binary token" implementation
turns it into a false offender.

Census on the current tree is unchanged: 89 anchored + 3 mid-line = 92
invocations, 0 unscoped, and the split produces zero extra segments on
live content.

* test(#2269): drop the unnecessary file-level allow-test-rule exemption

no-source-grep only fires on a readFileSync whose path expression carries
BOTH a .cjs/.js/.ts extension and a quoted bin|lib|gsd-core|src literal
(looksLikeSourcePath, eslint-rules/no-source-grep.cjs). This scanner reads
.md only, so the rule never triggered and the exemption bought nothing.

It was not inert, though: the escape is file-level — getAllComments() sees
the header and the rule returns {} for the whole file — so it silently
disabled no-source-grep for the pre-existing #2112/#2523 tests here and
for anything added later.

Verified by removing it and running eslint on the file: clean.

* test(#2269): mirror routeCommit by tokenizing, not by scanning line text

The scan's predicate searched the whole line for `--files` with double-quote
parity, while the runtime does an exact-token argv lookup scoped to a single
command. Those semantics disagreed, and the disagreement was exploitable in
both directions:

  guard=true  runtime=false | ... commit "docs: update ROADMAP.md" && echo done --files unused.md
  guard=true  runtime=false | ... commit 'docs: explain --files usage'
  guard=false runtime=true  | ... commit 'prints a " sometimes' --files .planning/PLAN.md

The first is #2269 verbatim: `--files` is echo's argument and never reaches
gsd-tools argv, so cmdCommit takes the blanket-`.planning/` default while the
guard stays silent. Injecting that line into secure-phase.md produced 0
offenders.

Each earlier round fixed one of these by widening the approximation, which
only moved the disagreement. So stop approximating: tokenize the line the way
a shell would — honouring BOTH quote characters, backslash escapes and unquoted
control operators — then run routeCommit's own predicate over the tokens.
Two previously hand-encoded special cases now fall out for free: `--files=x`
is unscoped (indexOf needs the exact token) and `--files -weird.md` is scoped
(the runtime filters on '--', not '-').

Operators are marked structurally rather than re-identified by comparing a
token's text against an operator set, because `commit '|' --files a.md`
produces a token whose VALUE is `|` and which is ordinary message text.

The property tests are part of this commit, not a follow-up: all three built
their line from a hardcoded double-quoted template, so no number of runs could
generate the single-quoted shape. They pinned one quoting dialect while reading
as though they pinned the predicate, which is what let the above through. The
delimiter is now drawn, and a fourth property asserts directly against a
re-implementation of routeCommit's own predicate.

Also removes a second copy of the old heuristic from the behavioral test, and
widens the mid-prose tier to `'` — it keyed on `commit "` only, so a
single-quoted mid-prose invocation was invisible to the scan entirely.

* test(#2269): scan docs/, which carries live invocations the claim excluded

The comment asserted the five roots were "every directory that carries live
invocations". That was false against the current tree, not hypothetically:
docs/zh-CN/references/ carries 7 live `query commit` invocations — the Chinese
mirrors of three gsd-core/references/ files that ARE scanned.

They are all scoped today, which is exactly why this needed its own assertion:
adding or dropping the root does not move the offender count, so the coverage
loss was silent in both directions. The scan now records which roots actually
contributed invocations and asserts docs/ is among them. Stated as a reach
property rather than a census figure, since a hardcoded count is the drift this
file has already been bitten by.

The locales are the right place to care about: ja-JP/ko-KR/pt-BR have no
references/ subtree at all, so the translations already drift per-locale — the
condition under which an unscoped example is reintroduced in one copy
unnoticed. New files under an existing root were always picked up by the
recursive walk; the gap was only ever at the root level.

* test(#2269): don't flag invocations inside HTML comments; drop a stale census

Two smaller findings from the same review.

The scan treated any `gsd_run … commit "` occurrence as executable, so an HTML
comment documenting the historical defect was reported as an offender:

  <!-- WRONG: gsd_run query commit "docs: message" (missing --files!) -->

Nothing in the scanned roots trips this today, but it made the guard hostile to
documenting the very bug it protects against — a plausible thing to add
precisely BECAUSE this issue exists. Comment spans are now stripped before
scanning, preserving newlines so a multi-line comment cannot fuse the text on
either side of it into one logical line. The strip lives beside the other scan
primitives rather than inside the test, so the assertion and the scan cannot
drift apart.

And the comment claiming "the total match count stays at 86" was stale (89).
No assertion read it. Rather than correct the figure, state the invariant it
was standing in for — dropping `.*` swaps onboard.md out of coverage WITHOUT
changing the offender count, since that site is scoped either way. The number
had already drifted three times; the property will not.

* test(#2269): consume comments, single-&, and redirections before scoring argv

Round-10 Major (davesienkowski): tokenize modelled &&/||/;/| and nothing
else, so `--files >/dev/null 2>&1`, `--files > out.md`, `# --files`, and
`& echo --files` all scored as scoped while routeCommit sees files=[] and
takes the blanket-.planning/ default. The live tail in
execute-phase-requirement-revert.md is one token-deletion from the silent
false negative.

The tokenizer now ends the command at a word-start unquoted `#`, treats a
single `&` as a control operator, and consumes redirections (with glued
IO numbers, `2>&1` included) as redir-marked tokens that hasScopedFiles
excludes from argv. Quoted and mid-word `#` stay literal — ship.md's
`PR #${PR_NUMBER}` message is pinned scoped. `$VAR` tokens deliberately
stay values: statically undecidable, and the three original #2269 fix
sites all pass `--files "${PHASE_DIR}/…"`.

* test(#2269): a foreign command's `commit` no longer reads as a gsd offender

Round-10 Minor (davesienkowski): segmentInvocations' no-hit fallback
returned the whole line, so `gsd_run query state && git commit -m "x"`
and `gsd_run query state | grep commit` — matched by the anchor's
deliberately loose `.*` — were scored as unscoped gsd commits and
false-flagged with the wrong failure message.

The fallback now applies only to a segment that actually carries the gsd
binary followed by a commit token (env-prefixed or wrapped invocations,
where dropping the line would be silent coverage loss); a line whose
`commit` belongs to another command contributes no candidates.

* test(#2269): drop the commands/ exemption from the per-root reach assertion

Round-10 Nit (davesienkowski): the exemption documented a gap that no
longer exists — commands/gsd/review-backlog.md contributes a candidate,
so commands/ is held to the same dead-weight test as every other root.

* test(#2269): put the tokenizer's escape branches inside the tested domain

Round-10 Minor 1 (trek-e): both backslash-escape branches in tokenize
were unreachable by any assertion — every generator stripped every
backslash, and no hand-pinned case carried one.

Double-quoted property messages may now contain `"` and `\`: embedFor
escapes them into the LINE while the property keeps the raw string as
the argv the shell would deliver, so the escape branch sits inside the
differential oracle's domain. Single-quoted messages still exclude
backslashes — the shell has no escape inside '…', so there is no escaped
spelling to generate. Four hand-pinned cases cover both branches in both
directions (escaped quote before a real --files, escaped quote hiding a
message-internal --files, escaped space in the message, escaped space in
the value).

* test(#2269): normalize path separators in the scan's diagnostic strings

Round-10 Minor 2 (trek-e): scanned/offenders entries were built from raw
readdirSync (recursive) names, so a failure message on Windows would
render backslashed paths, against the repo's normalize-unconditionally
convention. Join with the raw entry; report with the normalized one.

* test(#2269): pin the unquoted escape branch in the unscoped direction too

Claim-audit catch on the round's own response draft: the unquoted
backslash-escape branch was pinned only in the scoped direction, and the
unquoted property generator strips backslashes, so the unscoped
direction was outside every assertion. One pin closes it.

* test(#2269): decide an invocation by its command shape, not by its markup

Both scan tiers answered "is this text executable" with a syntactic proxy —
one anchored the command at line start, the other required a quoted commit
message — and both were wrong, in opposite directions. The anchor flagged a
fenced block that deliberately SHOWS the unscoped form. The quoted-message
tier could not see an unquoted invocation at all: `gsd_run query commit
fixup` reaches the identical cmdCommit, entered no candidate set, and
trimming its --files clause reintroduced #2269 with nothing to fail.

Markdown context was the obvious replacement and is refused, on measurement.
Keying on fences means parsing them — fence character, opening run length,
nesting, tilde fences, four-space indented blocks, unclosed markers — and
every bug in that parser is a silent false negative. I built that version
first and it lost four executable shapes before I stopped: `cd "$ROOT" &&
gsd_run query commit …`, `if …; then gsd_run …; fi`, four-space indented
blocks, and a four-backtick fence containing three-backtick runs. Each loss
was invisible; the census stayed at 99.

So the discriminator is the command shape:

    <binary> [ query | -flag [value] ]* <commit-token> <at least one arg>

The middle clause separates a command from a SENTENCE containing the same two
words — `Update STATE.md using gsd-tools.cjs query (or legacy gsd-tools)
commit mutations:` fails it at `(or`, with no markup parsed. The trailing
clause separates an invocation from a mention. Each line is scanned whole and
its inline code spans are scanned too, unioned: the whole-line pass reaches
every executable shape regardless of markup, and the span pass reaches the one
case it cannot, where a backtick glues to the binary token. The union is
additive, so a mis-parsed span can only add a visible false positive, never
hide an invocation.

This closes the unquoted case in BOTH the code-span and bare-prose forms, so
no part of it is left declined. Named residual, pinned as a test rather than
left to be discovered: an undelimited prose mention running straight into its
sentence (`see gsd_run query commit for the scoping rules`) is flagged, because
nothing distinguishes it from an invocation with arguments without guessing at
English. Bounded and measured — the roots carry 93 bare-prose mentions of the
binary and 0 with a commit token, the repo's convention is to backtick a
command reference, and the failure is a visible red.

Two things fell out. An interpreter prefix (`node gsd-tools.cjs commit …`,
live in docs/CLI-TOOLS.md) was invisible to both old tiers and is now in
scope; that pulled in the CLI usage synopses, so a bracketed optional FLAG now
marks synopsis notation — the discrimination is the bracketed flag, never
"contains a bracket", since 24 live invocations carry brackets inside their
quoted message (one as its --files value) and none brackets a flag.

tokenize and hasScopedFiles are untouched — byte-identical. Only the question
"is this text an invocation at all" changed. Census re-derived with the file's
own primitives: 99 candidates across 62 files, 0 unscoped, workflows 67 /
references 11 / agents 12 / commands 1 / skills 1 / docs 7 — identical to the
previous tier logic. Against the pre-fix tree (200daa456^) it still flags
exactly the 3 sites #2269 was filed about.

* test(#2269): let a wrong-example declare itself, in comment position only

A block showing the unscoped form on purpose scored as an offender. That is a
false positive against content correct as written, and the worst kind: the fix
a contributor reaches for is to mangle a teaching example until the linter
stops complaining. The HTML-comment escape only covered examples written as
comments.

Exempting fenced blocks was the prescribed remedy and is refused on
measurement: 96 of the 99 live invocations sit inside fences. Against the
pre-fix tree (200daa456^) the scan flags exactly the 3 sites #2269 was filed
about, and 0 with fences exempted. That does not narrow the guard, it disables
it — and no property of the surrounding markup can stand in for the author's
intent anyway, which is the general form of the same point.

So intent is declared: `# gsd-scan-ignore: <reason>` exempts the line, and a
reason is required because a bare token is not a declaration. Undeclared, the
identical line stays an offender — it is byte-for-byte what a real regression
looks like.

The marker is honoured ONLY in comment position, via the same unquoted-`#`
rule tokenize() already applies, so the two cannot disagree about where the
command ends. Matching the token anywhere on the line would let the commit
MESSAGE carry it — `commit "docs: explain gsd-scan-ignore: semantics"` — and
silently exempt a real offender. A guard that can be talked out of firing by
its own documentation is strictly worse than the false positive the marker
exists to fix, so both directions are pinned.

The marker is a new convention here and I would rather be redirected than
assume: happy to rename it, key it to an HTML comment, or drop it for whatever
shape you prefer.

* test(#2269): pin the escaped-backtick case in both directions

The span pass skips a backslash-escaped backtick, because an escaped backtick
is literal text and pairing it invents a code span the rendered document does
not have — the invented span then reads as an unscoped invocation, a false
offender against prose that merely displays a backtick.

That guard had no test: reverting it broke nothing, and a property with no
test is one the next edit removes for free. Pinned in both directions, so the
assertion covers the escape rather than the absence of span handling.

* test(#2269): close a bypass in the ignore marker's comment-position test

The marker's comment-position check keyed on "preceded by whitespace", which
is not the rule the shell applies and not the rule tokenize() applies. In
`commit docs:\ # gsd-scan-ignore: reason` the backslash escapes the space, so
the shell keeps `docs: #` as ONE word: the `#` is literal, the command runs,
and it runs UNSCOPED. The raw-text check saw a comment and exempted the line.

That is the guard being disarmed by text the author controls, which is worse
than the false positive the marker was added to fix — a guard that can be
talked out of firing is not a guard. Found by an adversarial audit of this
round's own claims, not by the tests, which is why both halves below are now
pinned.

Two independent conditions, deliberately:

- commentPortion tracks WORD START the way tokenize does, rather than looking
  at the preceding character. That closes the escaped-separator case directly.
- A marker that survives tokenization as an ARGUMENT disqualifies the line
  outright. tokenize() drops everything from a real comment onward, so a
  genuine declaration leaves no token carrying the marker; anything that does
  reached argv, which means the runtime executed it.

The second is what makes the closure structural rather than a matter of
getting commentPortion's edges exactly right — and it has edges. A
REDIRECTION swallows `#` and its text into a redir token, so the shell passes
it to the redirect target and never treats it as a comment, while raw-text
reading still sees one. Pinned as its own case, since it is the shape only the
cross-check separates.

Reverting either half fails the marker test independently.

* test(#2269): require a tracking reference on every scan-ignore declaration

The gsd-scan-ignore: marker accepted any non-space reason text, giving the
exemption no expiry and no ledger — the permanent-allow-test-rule shape
RULESET.TESTS.delete-bad-tests names. ADR-456 already settled this for the
sibling allow-test-rule: convention, so the marker now requires the same
#NNN issue reference (or an https:// URL).

scripts/lint-allow-test-rule-refs.cjs walks tests/ only and keys on the
ESLint comment form, so it cannot see a marker living in a .md file. Rather
than teach a second token to a script whose whole contract is that form,
the scan enforces the rule over its own roots — it already runs in CI on
every shard.

An attempted declaration that carries no reference is reported AS one,
asserted before the offender list. It is also an offender (it does not
exempt), and letting the generic assertion win would tell an author who had
already explained the line that their commit is unscoped — sending them to
re-read a flag that was never the problem.

Live declarations in the six scan roots: 0, so nothing is grandfathered.

* test(#2269): name every remedy in the failure, and document the convention

The offender assertion named only --files, but the guard has three distinct
causes and only one of them is the bug. An ordinary English sentence that
runs `gsd_run query commit` straight into its prose is flagged — by design,
since nothing separates it from an invocation with arguments without
guessing at English — and a contributor told only that their commit is
unscoped will mangle the sentence until the guard shuts up. That is exactly
the outcome the declaration marker was invented to prevent.

The message now names all three remedies (scope it / backtick the mention /
declare the wrong-example) and points at CONTRIBUTING.md. This is the
standard the repo already states one section over for its sibling gate:
"The failure output names its own remedy".

The help text is hoisted out of the assertion so it can be pinned. A failure
message is unreachable on the passing path, so nothing would have noticed a
remedy being edited back out of it.

The marker lives in .md files across six roots, so documenting it in this
test's comments reaches nobody who hits it. CONTRIBUTING.md now carries the
convention beside the sibling allow-test-rule: exception, and its example is
a live instance of itself — removing the declaration makes the example an
offender (verified).

CONTRIBUTING.md is outside changeset lint's USER_FACING_PREFIXES, so the
diff still returns ok_no_user_facing_changes (verified, not assumed).

* test(#2269): decide synopsis notation by the first argument, not by the line

SYNOPSIS_TOKEN_RE disqualified the whole segment, so notation anywhere on a
line silently disqualified a real invocation. Wrong in both directions, and
both reachable:

  See [--files](#anchor) then run gsd_run query commit "docs: x"
      ^ an ordinary markdown link, and docs/ is a scan root

  gsd_run query commit "docs: x" [--amend]
                                 ^ a real, executable, unscoped call

Both scored 0 candidates. The second is the dangerous one: `[--amend]` is a
literal word to the shell, so the line runs, reaches routeCommit with
files=[], and sweeps the index — #2269 verbatim, with the guard silent on
exactly the defect it exists to catch.

Positionally there is no ambiguity. A synopsis documents a call it does not
make, so its first argument is a placeholder; a real call's first argument
is its commit message. The test moved to that position.

`<message>` reaches the predicate as a REDIRECTION — `<` is a redirection
character, so tokenize() reads it the way the shell would. That mangling is
the identifying feature rather than an obstacle, and keying on it avoids the
false negative a raw-text match would have introduced: inside a quoted
message `<Widget>` is ordinary text, one token, no redirection. Pinned.

All five live synopsis lines (docs/CLI-TOOLS.md and its four localized
mirrors) remain excluded — verified by running the primitives over each.
Reversion control: restoring the whole-line rule fails 'usage-synopsis
notation documents the CLI and is not a call to it', and only that test.

* test(#2269): reach invocations behind a subshell, a shell -c, and an escaped backtick

Three shapes that were executable and invisible.

(1) `(gsd_run query commit "docs: x")`. Subshell grouping changes no argv,
so `(` was never a tokenizer metacharacter — which left `(gsd_run` as one
token the binary anchor could not match. The anchor now strips a leading
run of `(`.

Backticks are deliberately NOT stripped with it, and the first attempt that
did strip them is why this is stated rather than assumed: to a shell a
backtick is never part of a binary name, so a backticked invocation belongs
to the code-span pass, which extracts the command from inside the
delimiters. Stripping them in the anchor made the whole-line pass find the
same invocation a second time, absorbing the surrounding sentence as
arguments — the census went 99 -> 102 with no verdict changing. Measured,
reverted, and pinned as a comment so the next reader does not re-try it.

(2) `bash -c "gsd_run query commit fixup"`. A shell invoked with -c runs its
next argument AS a command, so the invocation sits inside a quoted token
that no markup rule can reach. The recursion is keyed on the INVOKER, never
on "a quoted token that parses as a command" — the wider rule would flag a
commit MESSAGE that quotes an invocation, a false positive against ordinary
documentation. Pinned in both directions.

(3) codeSpans skipped an entire backtick run on an odd backslash prefix,
which is stricter than the rule its own comment states: an escaped backtick
consumes exactly ONE, and the rest of the run is still a delimiter. Fixed,
and pinned with a case that now yields a span where it previously yielded
none. The reviewer's own example stays at zero spans — correctly, because a
1-run opener does not pair with a 2-run closer — and that is pinned too, so
the distinction is not re-litigated as a regression.

Census re-derived at this head over the six roots with the file's own
primitives: 99 candidates across 65 files, 0 unscoped (workflows 67 /
references 11 / agents 12 / commands 1 / skills 1 / docs 7) — identical to
the previous head, so the widening is coverage-neutral by measurement.

Reversion controls: 3/3 fire against a named test.

* test(#2269): assert scan coverage over the repo, not over each root

The per-root reach assertion was wrong in both directions at once.

Too strong: commands/ contributes exactly one invocation and skills/ one, so
an unrelated PR retiring review-backlog or re-syncing the Chinese mirrors
turned this red with a failure that had nothing to do with #2269. A root is
now allowed to legitimately go to zero.

Too weak, and this is the part worth having: it could only ever re-confirm
the roots already listed. Removing a root from scanRoots was SILENT under it
— the assertion iterates the roots that remain — and a directory that
ACQUIRES invocations without being a root was invisible to it. That is the
gap this file has actually been bitten by twice: agents/ in one round,
docs/zh-CN/ in the next, each found by a reviewer rather than by the suite.

Replaced with the property those checks were approximating: over every
tracked .md in the repo, a file carrying a live invocation must be covered
by a scan root. Dropping any root now fails (verified for docs/ and
agents/), and so does a new directory acquiring one.

This also closes the residual @davesienkowski raised in round 10 —
completeness was asserted only for docs/, never as a general property — and
the one I answered then by saying a git ls-files walk inside the test was
something I would rather propose explicitly than smuggle in. Proposing it:
git ls-files is already used against the repo by seven test files here, one
of which fails closed on a non-zero exit exactly as this does. An
unenumerable file list is an UNKNOWN coverage set, not an empty one.

Measured: 1523 tracked .md, 99 candidates inside the roots, and exactly one
candidate-bearing file outside them — CHANGELOG.md, whose invocation is
scoped. It is excluded with its reason rather than silently: it is
regenerated from .changeset/ fragments, so a marker added to it would not
survive the next release, and it records commands that shipped rather than
instructing anyone to run one.

Suite duration 6.5s -> 9.3s for the wider walk.

* test(#2269): put the recognition half inside the property domain

All four existing properties aim at hasScopedFiles — the half backed by a
runtime oracle, and the half that has been stable for rounds. Every defect
found since lives in the other half: whether a line is an invocation at all.
That asymmetry is the problem, because a miss there is a silent false
negative, where a miss in the scope predicate has an oracle watching it.

So the new properties generate the CONTEXT rather than the arguments. The
recognition rule is "an invocation is found by its command shape, whatever
markup surrounds it", and that is a claim about a domain a generator can
cover: subshells, interpreters, env prefixes, shell keywords, prompts, list
and blockquote markers, indentation, chaining, and `sh -c` wrapping. Each of
those was previously a hand-pinned example, several added only after a
reviewer found the gap.

It paid immediately. Three defects, none of which exists on today's tree:

- A markdown BLOCKQUOTE marker is the one markup form the tokenizer cannot
  ignore, because `>` is also a redirection: `> gsd_run query commit "docs:
  x"` reads as a redirection whose target is the binary. Found on the
  property's first run. 34 blockquoted lines in the six roots invoke this
  binary today; none carries a commit token yet.
- `> sh -c "gsd_run commit a"` — the blockquote strip fed only one of the
  three passes. The passes now apply to every VIEW of the line.
- `(sh -c "gsd_run commit a"` — the subshell strip was applied to the gsd
  binary test but not the shell-invoker test. One helper now, not two copies.

The last two are COMPOSITIONS of shapes that each pass alone, which is the
class hand-written examples are worst at and the specific reason this
finding was worth taking as stated rather than as more examples.

The wrapper axis is drawn only with an unquoted body: a -c payload is itself
quoted, so a quoted message inside it needs a nested-quoting domain, and a
wrong generator domain is this file's most repeated own-goal (two
seed-dependent reds). Constraining the body is what makes the axis safe.

Verified rather than claimed:
- 20,000 generated cases, 0 counterexamples; suite green on three
  independent seeds.
- Reverting each of the four recognition fixes in isolation is now caught by
  the property, not only by the hand-pinned case. Before the wrapper axis it
  caught 2 of 4 — measured, which is why the axis was added.
- Census unchanged: 99 candidates / 65 files / 0 unscoped, same per-root
  split, over 1523 tracked .md.
- Discrimination against the pre-fix tree (200daa456^): 97 candidates,
  exactly 3 offenders — next.md, secure-phase.md, validate-phase.md.

* test(#2269): cover the scan's own assembly, found by three silent controls

Running a reversion control over every fix in this round left three silent,
and all three were real gaps rather than control artefacts.

Two were the same shape: the WIRING between the walkers and the scan's
result lists was covered by nothing. Every test drove documentCandidates /
documentUntrackedDeclarations directly, so replacing the untracked-
declaration source with an empty list — and emptying the uncovered-file
list — both left the suite green. The real corpus cannot catch either: it is
clean, so those assertions can only ever observe an empty result. That is
the structural reason a synthetic corpus is needed and not a nicety.

Factored the per-document classification into scanDocument() and the
uncovered-file walk into uncoveredFiles(), both beside the other primitives
for the reason already stated there — an assertion must not pass against a
private copy while the scan does something else — and drove each with a
synthetic input. Both controls now fire.

The third was worse, because it looked like a check and was not one:
`assert.match(OFFENDER_HELP, /backtick/i)` is satisfied by the word
"backticked" in the clause explaining why a backticked mention is skipped,
so deleting the backtick REMEDY left the assertion green. Matched on the
instruction instead.

Reversion-control matrix over the whole round: 12 controls, 12 fire against
a named test. Suite 29/29, 0 skipped.

* test(#2269): keep this file's own prose out of the exemption lint

Self-found by running scripts/lint-allow-test-rule-refs.cjs, which the round
had a specific reason to run: the lint extracts everything after the token
on a line and requires a #NNN or URL in it, so the comment explaining that
requirement was itself read as a new untracked exemption --

  tests/commit-files-pathspec.test.cjs :: ` reason, and scripts/...

-- and lint-tests would have gone red on a prose line. Reworded so no line
carries the bare token; the lint now reports no novel offenders.

Fitting rather than embarrassing: this is the exact class Major 1 is about,
and the lint caught it on the file arguing for the same discipline.

* test(#2269): fix four defects an adversarial review of this round found

Ran a cross-AI adversarial review over the round's own claims before pushing.
It confirmed 10 of 12 and returned four MISSED findings; all four were real.

1. `shellDashCPayloads` searched the whole segment for a shell name, so
   `echo bash -c "gsd_run query commit fixup"` became a candidate. echo
   PRINTS the string; it does not run it. The invoker must be at the command
   position — only an env assignment, a shell keyword, or a list/prompt
   marker may precede it. `echo` and `printf` are commands and no longer
   qualify; `then` / `$` / `FOO=1` / a subshell still do. Both directions
   pinned.

2. SHELL_INVOKER_RE was a guess at four spellings. `ash`, `csh`, `tcsh`,
   `fish` and `yash` all take -c and all run what follows, and missing them
   is a silent false negative in the one function whose job is reaching a
   command the tokenizer cannot see. All nine spellings pinned.

3. A marker with NO reason at all fell through to the generic offender
   diagnosis, because the loose detector required `\S` after the colon. That
   is the likeliest way to get the marker wrong, so it is the case that most
   needs the specific message. The reason is now EXTRACTED rather than
   matched in one shot, so an empty reason is still an ATTEMPT.

4. The predicate now MIRRORS scripts/lint-allow-test-rule-refs' own
   ISSUE_REF_RE (`/#\d+|https?:\/\//`) instead of approximating it. The
   review refuted my claim of an HTTPS-only rule: the regex accepted
   `http://` and the prose describing it did not. The regex was right and
   the sentence was wrong — so the sentence is gone and the constant is
   cited. A contributor who satisfies one marker and not the other would
   otherwise have been handed two conventions wearing one name.

Also a generator-domain correction, not a narrowing: `node bash -c "..."`
is not an executable line — node takes a script path and `bash` is not one —
so the interpreter context is no longer drawn with a shell wrapper. Leaving
it in had the property demanding recognition of a non-command.

The review's other refutation is NOT actioned, and deliberately: it measured
9 failures, every one a fixture hook dying on `spawnSync /bin/sh EPERM`
inside its own sandbox, with the scanner and property tests passing there.
That is the documented reviewer-sandbox artefact class, not a defect in the
claims. Locally the suite is 29/29, 0 skipped, on three independent seeds.

Census re-derived after the rework: unchanged at 99 candidates / 65 files /
0 unscoped, same per-root split, 1523 tracked .md walked.

* test(#2269): pin the invoker set exactly, and let command modifiers through

Two follow-ons from the review's invoker findings, both verified rather than
reasoned about.

The invoker regex was hand-written and I did not trust it, so I enumerated
it instead of sampling: over every <=2-letter prefix it admits exactly
{sh, ash, bash, csh, dash, fish, ksh, tcsh, yash, zsh} and nothing else.
`ssh` is the near-miss that mattered — admitting it would treat a REMOTE
command as a local shell running the payload — and it is correctly rejected.
Pinned along with cash / josh / publish / wish / rsh.

The mirror of that gap is the prefix side: `time bash -c "..."` really does
run the payload, and the command-position rule introduced one commit ago
stopped at `time`. Command modifiers that pass straight through (time, exec,
nohup, env, command) now skip like the shell keywords do.

Named residual rather than a half-fix: a modifier carrying its OWN flags
(`sudo -u alice bash -c ...`) still stops the search, because skipping
arbitrary flag/value pairs means modelling each modifier's option grammar.
No live instance in the six roots, and the failure direction is a missed
candidate.

Census unchanged: 99 / 65 files / 0 unscoped. Suite 29/29, 0 skipped.

* test(#2269): say http(s):// where the predicate accepts it

An audit of this round's own response comment caught that the fix for the
HTTPS-only overclaim went only into the test's internal comment. The three
strings a contributor actually READS — the offender help, the malformed-
declaration message, and CONTRIBUTING.md — still said `https://` while the
shared predicate is `/#\d+|https?:\/\//` and accepts `http://`.

So the implementation mirrored the sibling lint while the documented
convention stayed narrower than it: exactly the two-conventions-wearing-one-
name problem the mirroring was adopted to avoid, reintroduced in the half a
contributor sees. Both directions were already pinned as tests; only the
prose was wrong.

* test(#2269): reach inside command substitutions, whatever precedes them

`$(gsd_run query commit "docs: x")` scored ZERO candidates in both
directions. `$(` glues to the binary exactly as `(` does, leaving
`$(gsd_run` as one token no command-name test could match — the same
silent false negative the subshell opener produced, on the invocation
idiom the tree uses most.

The strip is widened rather than duplicated, so it covers the shell
invoker test too: one helper, per the rule already stated beside it.

WHAT PRECEDES THE OPENER IS NOT ENUMERATED, and that is the whole of
the design. Keying on an assignment prefix covers most of the live
substitution sites and misses the rest — measured with the file's own
binary predicate rather than a line regex, because two regexes gave two
different totals: 432 tokens whose last `$(` is followed by a real gsd
binary, of which 423 are plain assignments and 9 are not. Those 9 span
three distinct non-assignment prefixes, all live:

  for REVIEW_FLAG in $(gsd_run review-lane flags)              (x3)
  ${PLAN_PRE_HOOKS_JSON:-$(gsd_run loop render-hooks plan:pre)} (x3)
  `ROADMAP=$(gsd_run ...)  -- backtick-glued assignment          (x3)

Ordinary text (`pre$(...)`) and an indexed assignment (`A[0]=$(...)`)
glue just as hard. Every such enumeration is one idiom behind the shell,
and every miss is silent, so the rule is positional: strip to the LAST
substitution opener in the token. Whatever preceded it was, by
construction, not the command.

Arithmetic falls out of the same mechanism rather than a special case:
`$((x))` leaves a `(` in front of the name, which no binary carries, so
it is refused — mirroring the shell, which itself needs `$( (` spaced
before it reads a nested subshell there. Pinned as a zero-candidate
case, as is the array literal `arr=(...)`, which contains no `$(` at all.

Scan census re-derived with the file's own primitives: 99 candidates
across 65 files, 0 unscoped — workflows 67 / references 11 / agents 12 /
commands 1 / skills 1 / docs 7. Unchanged, so the widening is
coverage-neutral by measurement: no live `$( ... )` site carries a
commit token.

Discrimination re-run: a planted unscoped substitution in a scan root is
caught and named; its `--files` twin is not flagged.

* test(#2269): stop the -c pass skipping a substitution-captured invoker

Self-found by sweeping the defect class the round's finding names — a
command context glues to the binary token — across every command-name
test rather than only the one that was reported.

`shellDashCPayloads` treats any `VAR=` token as a skippable prefix, on
the reasoning that `FOO=1 bash -c "…"` prefixes the command. But
`V=$(bash -c "…")` IS the command: skipping it walks the invoker search
past the shell onto `-c`, which is not an invoker, so the payload is
never reached and an unscoped commit inside it stays invisible.

An assignment is now skippable only when it carries no substitution, and
the prefix pattern admits the indexed form (`A[0]=`) for the same reason
the strip above does. `FOO=1 bash -c "…"` and the full invoker matrix
are unchanged and still pass, which is what makes this a narrowing of
the skip rather than a removal of it.

No live instance in the six roots — the failure direction is a missed
candidate, which is the direction this file treats as the one it cannot
afford.

* test(#2269): quoting decides the opener, per character not per token

Found by an adversarial pass over this round's own claims, then fixed
again after that pass refuted the first fix.

THE DEFECT. `printf %s '$(gsd_run' query commit fixup` runs no gsd
command: the opener is single-quoted literal text. But tokenization
removes quotes, so the token's value is `$(gsd_run` and the strip read
it as a binary, fabricating an invocation out of a string argument.

THE FIRST FIX WAS WRONG, IN THE UNAFFORDABLE DIRECTION. Recording "did
this token consume a quote" and refusing the strip on it closes the case
above and opens a worse one: a quote need not cover the opener.
`echo "pre"$(gsd_run query commit fixup)` executes, and so does
`$(gsd_""run query commit fixup)`, and a token-wide flag silences the
guard on both. That trades a visible false positive for a silent miss,
which is the trade this file refuses everywhere else.

So the mask is PER CHARACTER: tokenize carries a '0'/'1' string parallel
to each token's value, marking characters that were quoted or
backslash-escaped, and the command-name helpers strip only an opener
whose own characters were bare. Both directions are pinned — the three
quoted-opener shapes score zero, the two mixed-quoting shapes score one.

This also closes the same shape for the subshell opener (`'(gsd_run'`),
which predates this round, and it retires BINARY_LEAD_MARKUP_RE: the
walk it replaced could not consult the mask, and a regex plus a mask
would have been two descriptions of one rule.

NAMED RESIDUAL, unchanged by this and out of scope. `printf %s 'gsd_run'
query commit fixup` still reads as an invocation, because no markup is
stripped there — the token's value simply IS the binary name. Closing it
needs quote provenance carried through the command-shape test, not just
the strip. The direction is a visible false positive with a declared
remedy, not a silent miss.

Scan census unchanged: 99 candidates / 65 files / 0 unscoped.
2026-08-12 20:45:11 -04:00