Commit Graph

993 Commits

Author SHA1 Message Date
Tom Boucher
0a0905705a fix(#4256): resolve todos from the root via todosDir everywhere (#4479)
* test(#4256): pin todos as root-scoped under workstreams (RED)

* fix(#4256): resolve todos from the root via todosDir everywhere

* chore(#4256): changeset fragment (pr number to backfill)

* chore(#4256): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 08:24:14 -04:00
Tom Boucher
6ebe6372ce fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474)
* fix(#4243): anchor stateReplaceProgressPercent bold form to line start

The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag,
so a bold percent-ish label quoted MID-SENTENCE inside prose — an
Accumulated Context bullet mentioning **Progress:** — captured the
machine-segment rewrite and destroyed the rest of its line, silently, while
the real Progress line stayed stale (and the frontmatter moved on without
it, breaking the #4213 surfaces-agree contract). Every caller
(cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation)
feeds the whole document, so all three were exposed.

Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the
exact idiom #4453 applied to stateReplaceField's bold branch (same-line
confinement per #4010: the leading class is [ \t]*, deliberately not \s*,
which can consume the newlines before the label into the match; $ is
explicit-and-inert and documents end-of-line).

#2177's recorded requirements all stand: frontmatter is stripped before
matching, the suffix-preserving machine-segment swap is untouched, and
bold-beats-plain priority now governs line-start forms, so an earlier
free-text plain Progress: line still cannot capture the rewrite ahead of the
real bold status line. Per the maintainer ruling (2026-09-07), #2177's
incidental bold-anywhere matching was not load-bearing.

* test(#4243): scope the C4 region check with splitLines, not a bare \n split

lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's
content.split(/\n## /)[0] — a bare \n split on readFileSync content is
CRLF-fragile under Windows autocrlf. Same scoping via splitLines()
(src/text-lines.cts), which splits on \r?\n.

* chore(#4243): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 04:27:35 -04:00
Tom Boucher
c4b6dbd486 fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468)
* test(#4247): failing-first regressions for checklist-form update-plan-progress

* fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry

* fix(#4247): single local source for the phase-heading anchor grammar

* docs(#4247): note the missing_phase_details refusal in cli-tools reference

* docs(#4247): backfill pr number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-07 02:20:47 -04:00
Tom Boucher
33e393ba4c fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453)
* test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip

* fix(#4243): anchor stateReplaceField bold form to line start

The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold
label quoted mid-sentence inside prose — the issue's **Status:** inside an
Accumulated Context bullet — captured the rewrite and destroyed the rest of
its line, silently, whenever a whole-body caller fed the function every
section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan
writes). The plain branch was always line-anchored; only the bold branch
lagged.

Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's
same-line confinement idiom for the leading class (deliberately not the
issue's suggested ^\s* — it can consume the newlines before the label into
the match) and #4186's recognition-by-anchoring discipline. Frontmatter
half of the issue (unknown-key drops, invented milestone defaults) is
already fixed on next by #2202/#3216/#4129; pinned here with the issue's
requested regression fixtures.

* test(#4243): pin survival contract, not derived percent, in frontmatter rows

Bench RED run caught two assertion defects in the pin rows: the unknown
progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a
declared derived subkey - omitted under the #3573 no-roadmap withhold,
recomputed when measured (#4129) - so pinning its value over-pins derived
semantics. The rows now pin what the issue demands: unknown/custom keys
survive, stored counters are kept under the withhold, milestone identity is
never reset to invented defaults.

* chore(#4243): changeset for the anchored bold-field fix

* chore(#4243): backfill PR number in changeset
2026-09-07 00:03:15 -04:00
Tom Boucher
8c8eda46b0 fix(#4225): scope the sibling-worktree phase-number horizon to the active workstream (#4450)
* test(#4225): failing-first matrix for phase.add --ws workstream-scoped numbering

Nine rows driven through the real CLI: the issue's verbatim topology
(root roadmap @39 committed, workstream @2, sibling git worktree carrying
the root roadmap), same-workstream sibling boundary, empty-workstream
first phase, coincidental root-maximum, no---ws control (the #3849
global horizon, byte-for-byte), cross-workstream isolation, sibling
lacking the workstream (fail open), add-batch parity, and a
next-decimal control. Rows 1/2/3/6/7/8 are RED on next @38e4ce5f62
(numbering computed from the sibling ROOT roadmaps: 40 instead of 3).

* fix(#4225): scope the #3849 sibling-worktree widening horizon to the active workstream

collectSiblingWorktreePhaseNums scanned each sibling git worktree's ROOT
.planning/ (phases/ dirs + ROADMAP.md headers) unconditionally. Under
--ws (GSD_WORKSTREAM), every local number source flows through
planningDir(cwd) and lands in the workstream scope, but the widening
horizon still merged the siblings' ROOT-roadmap numbers into it — so
phase.add --ws in a workstream at Phase 2 inside a project whose root
roadmap sits at Phase 39 minted Phase 40 (directory 40-<slug>, and a
Depends on: Phase 39 that does not exist in the workstream's numbering
universe).

The horizon now resolves each sibling's planning dir through the SAME
canonical resolver, planningDir(wt, ws), with the env workstream read
once via planningDir's own discriminator: a workstream-scoped allocation
scans the sibling's copy of the SAME workstream (a number taken by that
workstream on another branch is still taken — the #3849 widening
survives, scoped), and never the sibling's root roadmap or another
workstream's. No workstream active: ws is null and the root-scope
horizon is byte-for-byte the #3849 behavior. A sibling lacking the
workstream directory contributes nothing (fail open, unchanged).

phase.add and phase.add-batch share the helper; both scopes of both
verbs are covered by the matrix in the previous commit. Output shape
and the publishStateContract boundary are untouched — only the number
changes.

* fix(#4225): rename siblingPlanning -> siblingPlanningDir (review nit)

* chore(#4225): changeset fragment (pr number to backfill)

* chore(#4225): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 22:24:07 -04:00
Brenden Smerbeck
e54d3aa159 enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key

- Add compact_content: false to the nested workflow object in
  gsd-core/bin/shared/config-defaults.manifest.json
- Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts
  so an absent key resolves to false via config-get --raw
- validKeys entry in config-schema.manifest.json already present

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(#4401): behavioral and boundary tests for workflow.compact_content

- 19 behavioral tests covering config-set/config-get round trip, invalid-shape
  rejection (banana, 42, empty string), the corrected null-unset semantics
  (#2046), absent-key resolution against config-defaults.manifest.json,
  config-new-project wiring, and doc-row shape assertions
- Drops the install-tree fixture-parity block (and its docstring item) that
  asserted gsd-core/references/compact-content-gate.md and
  gsd-core/workflows/compact/map-codebase.md fixture entries — those paths
  belong to #4402 and do not exist on this filtered branch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(#4401): document workflow.compact_content in both config references

- One 4-cell row in docs/CONFIGURATION.md (workflow.* run)
- One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md
- Both cross-reference ADR-4139

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): add changeset

- Added-type fragment, pr: 4401 (issue number; backfill to the real PR number
  is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR-
  FIELD-DRIFT)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(#4401): backfill changeset pr field to #4441

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS

SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false
instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling
entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving
a single-source-of-truth drift risk: a future manifest-only edit to the
default could silently diverge from this literal, only caught later by
the D-03 test if it ever happened to manifest.

Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and
derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority
sibling pattern. Found during maintainer review (review-open-prs) of
this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP

The previous commit added compact_content to CONFIG_DEFAULTS in
src/config-loader.cts but missed the matching entry in
tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat
CONFIG_DEFAULTS keys to their namespaced doc form before checking
gsd-core/references/planning-config.md for a match. Without it, the
test looked for a bare `compact_content` doc reference instead of the
actual `workflow.compact_content` row, and failed:
"CONFIG_DEFAULTS keys missing from planning-config.md: compact_content".

Found by actually running gsd-test against the branch rather than
trusting the plausible-looking fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4401): register compact-content-4139 test in the docs-guard lane

tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md
directly (fs.readFileSync) to assert the workflow.compact_content doc row's
shape, which makes it a doc-reading test file under the #3753 docs-guard
lane. It was never added to scripts/docs-guard-registry.cjs's
DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so
tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed:
"compact-content-4139.test.cjs reads a docs/ path but is not registered in
the docs-guard lane and carries no docs-guard-exempt marker".

Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed
path it reads; gsd-core/references/planning-config.md is outside this
registry's docs/ scope, matching the sibling config-field-docs.test.cjs
entry's existing convention).

Found by actually running gsd-test against the branch — this gap predates
the maintainer's config-loader.cts fix and was already present in the
original PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: sim <sim@local>
2026-09-06 19:52:59 -04:00
Tom Boucher
38e4ce5f62 fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin

Three defects from #4186:

1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the
   free-prose body Status field, so prose merely mentioning a status word
   was silently rewritten to a credible wrong token (a .planning/ path in
   Italian prose -> status: planning; verifica -> verifying; completezza ->
   completed). Recognition is now an ANCHORED whole-field match against a
   declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS,
   state-document.cts) — case/whitespace-tolerant, branch-order artifacts
   preserved (Planning complete -> planning; Phase complete — ready for
   verification -> verifying). The recorded lenient fallback (#3873 row 26)
   stands: unrecognized prose passes through verbatim. Read-side consumers
   (W011, statusline) ride the same function.

2. The progress recount skew (stray *-SUMMARY.md inflating
   completed_plans) is already dead on next via #1988/PR #2016
   (countMatchedSummaries pairs summaries to plans) — verified live and
   pinned with regression rows composed against the #4129/#4359 ratchet.

3. state record-session with no args executed and wrote STATE.md; it now
   errors like state update (stopped-at or resume-file required), handler-
   side so SDK callers are covered too. Four tests pinning the bare-call
   write are updated to the new contract.

* fix(#4186): update status pins to the anchored vocabulary contract

Bench round 1 follow-ups:

- Legacy bare 'Milestone complete' kept as reader-side vocabulary
  (ADR-2207 removed the writers, not recognition of legacy files).
- state.test pins updated: 'Paused at Plan 3' and round-trip
  'Executing Plan 5' were pins of the substring guessing itself —
  the round-trip now uses the real handler form 'Executing Phase 5'.
- record-session no-op/no-fields tests repurposed to the usage-error
  contract (CLI + SDK-level ExitError), byte-unchanged assertions kept.
- statusline tests repinned: vocabulary values collapse to keywords;
  narratives render the documented first-word fallback instead of a
  guessed token. Hook doc comment updated to match.
- docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md.
- docs/CLI-TOOLS.md: record-session signature notes the required flag.

* fix(#4186): repair a dangling sentence in the schema docstring

* test(#4186): bound the completed_plans scan regex (#2128 class)

* chore(#4186): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-06 17:08:24 -04:00
Tom Boucher
f09e7ed08c fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375)
* test(#4137): keg-only Homebrew Cellar path falls back to raw execPath

Regression tests for the Homebrew branch of normalizeNodePath: the rewrite
to <prefix>/bin/node must be existsSync-guarded like the mise/volta
branches, falling through to the raw execPath when the keg-only formula
was never linked into <prefix>/bin. Also makes the existing #3181/#2185
Cellar assertions hermetic by injecting existsSync stubs (granting
existence to exactly the one candidate each asserts) so they no longer
depend on the runner machine's real /usr/local/bin/node.

* fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath

The Homebrew branch of normalizeNodePath returned <prefix>/bin/node
unconditionally — the only one of five runtime branches that never probed
its rewrite candidate. On a keg-only or versioned Homebrew install
(node@24 never brew-linked) that path does not exist, so every managed
hook command baked by resolveNodeRunner/buildBakedNodeToken/
buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh:
<prefix>/bin/node: No such file or directory.

Guard the rewrite with the already-injected existsSync exactly like the
mise and volta branches: when <prefix>/bin/node exists (linked formula)
the rewrite is byte-identical to today; when it does not, fall through to
the raw execPath — a working keg path instead of an immediately broken
one. Also drops two now-unused constants from the regression tests.

* test(#4137): make the #977 non-fnm Cellar assertions hermetic too

The Bug #977 folded block's two 'still maps to stable symlink' assertions
called normalizeNodePath without an existsSync stub, silently depending
on the runner machine's real /usr/local/bin/node (present on the Linux
bench image, absent for /opt/homebrew). With the #4137 guard these become
environment-dependent; grant each exactly the one candidate it asserts.

* chore(#4137): add changeset fragment

* chore(#4137): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-06 16:04:21 -04:00
Michel Moreira
54085516c1 fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371)
* fix(#4211): materialize Kimi's agent tree recursively during surface apply

kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and
install copies that tree recursively (_copyStaged). Surface apply fell through
to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`:
the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was
written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name
that already carries it. A surface change could therefore corrupt Kimi's
installed artifacts while still reporting success.

Three divergences from the install path, all in src/surface.cts:

- _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped
  to exactly what install's _removeGsdEntries owns for this kind (the two root
  files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned
  and preserved.
- applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for
  an unmodified full profile, as it already does for the agents kind and as
  createRuntimeArtifactInstallPlan does for every kind — without it Kimi's
  generated subagents lost their path-prefix rewrites and attribution trailer,
  and an unmodified full profile staged only the skill-referenced subset.
- applySurface runs rewriteStagedSkillBodies for kimi-agents, which the
  install plan routes through it alongside skills.

* chore: add changeset for #4211

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 15:11:47 -04:00
Tom Boucher
acb3cc974b fix(#4197): dedup the update-context fast path against the selected global dir (#4413)
* fix(#4197): dedup the update-context fast path against the selected global candidate

The preferredConfigDir fast path derived scope from a cwd-relative match
alone, so a global install reported LOCAL whenever the shell sat in
$HOME — and run_update then drove the installer through its --local arm
(settings.local.json + the #338 relocation) against a global install.

Extract resolveGlobalCandidate (env candidates first, then $HOME-relative,
first hasInstall hit wins) and use it in BOTH paths: the fast path now
answers LOCAL only for a cwd-relative match that is not the selected
global dir, which is the same dedup the cascade applies at its isLocal
check. A preferred dir that is also the env-directed global now answers
GLOBAL on both paths (the cascade's answer), pinned by a parity test.

The discriminator is the selected global candidate, not the $HOME
pathname: with CLAUDE_CONFIG_DIR directing the global elsewhere,
$HOME/.claude probed from cwd === $HOME is a genuine local install, and
a pathname check would re-break parity (regression-pinned).

* chore(#4197): add changeset

* chore(#4197): backfill PR number in changeset

---------

Co-authored-by: agent-4197 <agent-4197@gsd.local>
2026-09-06 14:09:32 -04:00
Tom Boucher
b7917882bb fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure

The 240-char pending-todo bullet cap must be deterministic w.r.t. where the
repo is checked out. Deterministic long-base-path fixtures (a single 110-char
segment, no real macOS dependency) reproduce next's own macos shard 3/3
failure (run 34038716700) on every OS: with an absolute link the bullet
exceeds the cap and the documented needs-first truncation drops the
'Needs <solution>' clause. Rows cover the determinism property (byte-identical
bullets under short and long bases), the CLI surface, relative-path stability,
legacy no-projectRoot behavior, drop-order preservation, and adversarial
edges (outside-root, path===root, non-string path).

* fix(#4384): render the pending-todo bullet link repo-relative

renderPendingTodosMarkdown gains an optional projectRoot; when given and the
todo's path is absolute, the bullet's [todo file](…) target becomes
toPosixPath(path.relative(projectRoot, path)) — the idiom already used for
project_exists. cmdInitTodos passes cwd.

The JSON todos[].path field stays absolute (#2376). Only the rendered display
link changes: embedding the machine-variable absolute base let macOS's
/private/var/folders/… temp paths consume the 240-char budget and drop the
'Needs' clause on long-path machines only — next's own macos-latest shard 3/3
went red on exactly this (run 34038716700), Linux's short /tmp passed. The
240-char whole-bullet cap and the needs→title→area drop order are unchanged;
this matches PR #4384's own canonical example, docs, and unit tests, which all
show repo-relative links. Docs updated at all three surfaces that describe the
bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was
still pre-#4384 'count and reference' prose).

Fixes the macOS regression introduced by #4384; next is red on its own CI.

* test(#4384): fix substring false positive in the outside-root regression row

The ../-form relative link legitimately contains the absolute path as a
substring, so !line.includes(absolutePath) fired on correct output (caught by
the first remote verify run, linux-node24 44018/44019). Assert the property
itself instead: extract the link target and require it to be non-absolute and
not equal to the absolute path.

* chore(#4398): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 13:23:26 -04:00
Tom Boucher
708d9a0b82 fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file (#4388)
* test(#4187): bare VERIFICATION.md regression matrix for the status surface

Both query verbs must agree on every row: bare file, suffixed variants,
missing file, other-dir placement, and the staleness seam. Row 1 is the
failing-first regression from the issue repro.

* fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file

readVerificationStatus and its internal staleness check
(findStaleVerificationSummary) called the shared resolver without
allowBare, so a phase whose only report was a bare VERIFICATION.md read
as missing and was told to re-run execute-phase while
verification.resolve-file, determinePhaseStatus, and both init
verification_path projectors all resolved the same file. Both call
sites now pass allowBare: true, matching the other five; tier order
(dashed > bare) is unchanged, so only bare-only directories change
behavior.

* fix(#4187): correct call-site counts in allowBare docblocks

Adversarial review caught the comments claiming five of six call sites
opted in; the current tree has six call sites with four previously
passing allowBare — the two module-internal status-path sites were both
holdouts, not one.

* chore(#4187): changeset for the bare VERIFICATION.md status fix

* chore(#4187): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 11:54:03 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
b7406b293f enhance(#2618): render pending todos as one bounded bullet per todo (#4384) 2026-09-06 08:06:39 -04:00
Tom Boucher
66e4034fe4 fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380)
* test(#4138): failing-first regression — begin-phase without --phase must fail closed

* fix(#4138): begin-phase without --phase exits non-zero and writes nothing

* chore(#4138): changeset fragment for begin-phase arg validation

* chore(#4138): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-06 07:03:39 -04:00
Tom Boucher
03738824de enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) 2026-09-06 05:46:49 -04:00
Tom Boucher
0aa4202f6a fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376)
* test(#4135): regression rows for pristine regen coverage collapse

RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a
null-returning stub (wired into verifyFile after the #4145 orphan tier,
behavior-neutral) so the git-history rows fail behaviorally, not at require
time. Failing-first rows: baseline_covered aggregate on a 1-of-13
multi-version fixture, coverageHeadline typed renderer, the opt-in
--min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass
and malformed-value boundaries), git-history baseline recovery (dropped-line
catch + surviving-line verify + older-commit hop), findPristineInGit unit,
Step 5a workflow headline contract, and the installer-side
describeBaselineCoverage honest N-of-M summary with the collapse disk-state
pinned. Negative-space rows pin today: non-git ok_no_baseline posture,
no-match-no-adoption, #3657 drift never rescued, canonical precedence, and
no git tier without --pristine-dir.

* fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening

The #3407 promotion rule regenerates gsd-pristine/ baselines from the
INCOMING release source and keeps only candidates byte-identical with the
OUTGOING recorded hash — correct in isolation, but on a multi-version jump
the surviving set is precisely the files upstream did NOT change. The
verifier then reports ok_no_baseline (advisory, exit 0) for everything
else, and no surface distinguishes a 12-of-13-unverified green run from a
fully-verified one: the human summary printed Checked/Failures only, the
JSON had no coverage aggregate, and the installer's update output gave
per-bucket counts without N-of-M framing.

All three issue directions, none exclusive:

- Report coverage prominently: --json gains an additive baseline_covered
  aggregate; the human summary leads with 'Baseline coverage: N of M
  file(s)...' on every run plus an advisory section naming each skipped
  file and reason; the installer prints an honest covered-of-modified line
  via the exported describeBaselineCoverage helper (typed return, exact
  contract); workflow Step 5a computes and prints the headline before any
  pass/fail framing.
- Fail louder on low coverage: opt-in --min-baseline-coverage <0..1>
  exits with new documented code 3 when coverage falls below the
  threshold (>= semantics; empty run vacuously passes; content failure
  exit 1 outranks it; malformed values are usage errors, exit 2).
  Default posture unchanged — no_baseline stays advisory per #934.
- Widen the promotion rule (its only trustworthy form): when no baseline
  resolves under gsd-pristine/ and a hash is recorded, the verifier now
  recovers the baseline from the config dir's own git history — the
  workflow's documented Option A — anchored by the same authority every
  tier trusts, exact pristine_hashes sha-256 equality. Read-only
  (git log/git show, windowsHide per #685), bounded (100 commits/file,
  10s/subprocess), null-on-any-failure so ok_no_baseline remains the
  universal fallback. Tier order: canonical join -> #4145 orphan scan ->
  git history -> OK_NO_BASELINE; #3657 drift and canonical precedence
  untouched.

Hash validation in saveLocalPatches is NOT relaxed — the collapse is
legitimate conservatism; hiding it was the bug. Measured on the issue's
shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report
baseline_covered 1/13 with the headline and can gate at exit 3; a
git-managed config dir with the outgoing bytes in history verifies 13/13.

Review fixes folded in: workflow headline derives the unverified count
from checked - baseline_covered (not the drift+no_baseline sum), and the
new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on
the marker line.

Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it.

* chore(#4135): backfill PR number into changeset fragment

---------

Co-authored-by: agent-4135 <agent-4135@gsd.local>
2026-09-06 05:26:05 -04:00
Tom Boucher
7bb366e836 fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening

Block A (flag): check decision-coverage-plan --context <path> must route
identically to the positional form; flag wins over positional context;
valueless --context falls through to the #2770 fail-closed caller error;
verify keeps its positional surface (flag is plan-only). RED on base:
the flag token lands in the args[2] phase slot (false uncovered) or the
args[3] context slot (silent CONTEXT.md-missing skip).

Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper
(?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the
no-adjacent-overlap property; a differential property compares the module
against a frozen copy of the pre-hardening grammars (reference validated
against the base build: 60k generated lines, 0 mismatches); 40k cliff
shapes assert correct outcomes with no wall-time asserts (repo rule).

A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags.

* fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions

(A) check decision-coverage-plan --context <path> — sibling convention
(check predicate, #2008): --flag value pairs parsed by the new shared
partitionPredicateArgs (parsePredicateFlags reimplemented as its flags
half — one parser, cannot diverge), the flag winning over a same-purpose
positional, positionals kept (no sibling deprecates them; the plan-phase
workflow caller passes positionals), valueless --context falls through
to the #2770 fail-closed caller error. Repair of the routing accident
where --context landed in the args[2] phase slot (false uncovered) or
the literal token in the args[3] context slot (silent green skip).

(B) parseDecisions regex seam hardened, byte-identical on all legal
inputs: the three bullet grammars consume the ID atomically via the
(?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split,
~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to
[^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group
indices unchanged (handlers untouched). Pinned by regex-lattice tests,
a differential fast-check property vs the frozen pre-hardening grammars,
and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo
rule — no deterministic engine step counter exists in Node).

* docs+test(#4130): document --context invocation; harden lattice test tooling

- docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan
  gate directly' block documenting both the positional and --context
  forms, flag precedence, and the valueless-flag fail-closed semantics
  (same place the gate's behavior is documented; sibling check predicate
  documents its flags the same way).
- Two changeset fragments per the maintainer brief (Added: flag; Fixed:
  hardening), PR numbers to be backfilled.
- tests/decisions.test.cjs review fixes: readRegExpTemplate template
  escaping (bare ')' SyntaxError), range-aware lattice checker with
  backreference skip and template unescape, honest A1 contract, lint
  escape warning.

* fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context

Suite-caught fixes from the first verify run:
- cmdDecisionCoveragePlan now refuses a flag-shaped token as the
  positional context path: a bare valueless --context stays a positional
  (sibling parser semantics, unchanged) but reading it as a PATH would
  turn a caller mistake into a silent 'CONTEXT.md missing' green skip —
  exactly what #2770's fail-closed law forbids. Now falls through to
  the missing-context-argument error, as documented.
- A8 test compares decoy-positional+flag against flag-with-phase (phase
  held constant) so the row isolates WHICH context was read; the old
  form compared against a no-phase invocation that could never match.

* chore(#4130): backfill PR number in changeset fragments (PR #4374)

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:17 -04:00
Tom Boucher
6adf3098ac fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines

RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a
null-returning stub so the new rows fail behaviorally, not at require time.
Failing-first rows: verifier resolution (no_baseline must drop to 0 when an
exact-hash orphan exists), findPristineByHash unit row, and the two
saveLocalPatches relocation rows. Negative-space rows pin today's behavior:
missing baselines still report ok_no_baseline, mismatching orphans are never
adopted or deleted, canonical precedence and the #3657 drift posture are
untouched.

* fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans

Both pristine readers joined the manifest-keyed path strictly, so a snapshot
stored without the gsd-core/ prefix (an earlier release's writer) was reported
as ok_no_baseline by the verifier and pushed into regeneration by
saveLocalPatches — where incoming-release candidates can never satisfy the
recorded outgoing hash, leaving the correct baseline permanently unconsumed.

- src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash —
  deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the
  recorded pristine_hashes entry (the same authority the #3657 drift guard
  trusts), symlink-skipping, canonical path excluded via skipRel.
- verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded
  hash, adopt byte-identical content found anywhere under gsd-pristine/ before
  reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and
  the frozen REASON/report shapes are untouched; the verifier stays read-only.
- install.js saveLocalPatches(): preserve-check rescue — relocate a
  hash-matching orphan to the canonical path (copy, hash-verify, then remove
  the orphan) so the state self-heals on the next update instead of repeating
  forever. Honest accounting: new non-overlapping rescued counter.
- Workflow doc: one-sentence note on hash-based snapshot resolution.
- Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore
  entries for the compiled artifact, seedFixture mkdir fix in the new rows.

Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145)

* fix(#4145): review follow-up — orphan scan never consumes a canonical path

Adversarial review finding: with two modified files sharing byte-identical
outgoing content, recoverOrphanedPristine could adopt the OTHER file's
canonical pristine as its rescue source — relocating it (copy + delete at
its home path) and ping-ponging the single baseline between the two files
across updates. findPristineByHash's skip parameter now accepts a Set, and
saveLocalPatches passes the normalized manifest keys so every canonical
path is excluded; only genuine non-canonical orphans are eligible for
removal (no strict-join reader ever consults those). Adds the
canonical-theft regression row, a Set-skip unit assertion, and tightens the
workflow doc sentence the same pass flagged as overstated.

* fix(#4145): INVENTORY roster row + symlink-fixture correction

Two leftovers from the ab17b7a1e5 bench run, both root-caused:
- docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs
  (#3762 gate: every manifest entry carries a row).
- The findPristineByHash symlink unit fixture placed its symlink target
  INSIDE the scanned root, so the walk legitimately matched the real target
  file. The implementation skips the symlink itself; the fixture now keeps
  the target outside the scanned tree so the assertion tests what it claims.

* changeset(#4145): fixed fragment for pristine baseline hash resolution

---------

Co-authored-by: gsd-agent <agent@gsd.local>
2026-09-06 02:04:18 -04:00
Tom Boucher
c3e2da153b fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names

A first-milestone ROADMAP.md H1 that puts the version after the name
(# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's
own version token, which the ADR-3180 §7.2 pinned name rule returns as a
COMPLETE-scope milestone name. Failing-first coverage:

- getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only
  fallback) must yield TRUNCATED {version, name: null}, never ')'
- the refusal is level-agnostic (H2/H3)
- punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only)
- listMilestoneHeadings enumerates the heading with name: null
- init manager CLI reports milestone_name: null and no lone ')' anywhere
- property (seed 20260905, 300 runs): a word-char remainder is always a
  name, a punctuation-only remainder never is
- negative space: canonical delimiter forms, parenthetical names (#3171),
  trailing markers, digit-only names, CRLF headings, version-last-no-parens
  control

* fix(#4134): refuse punctuation-only milestone heading names

extractMilestoneHeadingName returns everything after the heading's own
version token as the name (ADR-3180 §7.2 pinned rule), which assumes
version-then-name. A name-then-version heading — the H1 a first-ever
ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves
exactly ')' after the token, and that fragment was returned as a
COMPLETE-scope milestone name, propagating into init.* JSON output and
buildStateFrontmatter's STATE.md writes.

A remainder with no letter or digit anywhere (any script) is heading
structure, not a curated name: refuse it as name: null so callers report
the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names
that merely contain punctuation are unaffected — '(' stays an ordinary
name character (#3171) — and digit-only names qualify.

Also closes the template gap that lets the shape occur: the roadmapper
agent's output_formats now templates the version-free canonical H1
('# Roadmap: [Project Name]', per templates/roadmap.md) instead of
leaving a first milestone's title line to invention. The new section
shifts the file's existing bare-gsd-tools prose mention from line 647
to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it.

Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134)

* chore(#4134): add changeset

* chore(#4134): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 23:31:35 -04:00
Tom Boucher
e6d047decc fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359)
* test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment

* fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet

Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/):

P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at
the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the
one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch
that owns the denominator. A completed phase whose verification routes stale
(#2348 clean-commit-time drift) or is missing no longer under-counts forever.

P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on
a measured scan: totals derived both directions (#2440), completed counters
up-only (#2969 — the schema-declared progress-ratchet, now enforced on the
write path like the read path always has), percent recomputed from the merged
counters. The #3756 unmeasured guard and the #3242 explicit-progress contract
are unchanged.

P3 — phase complete's atomic 3-file commit passes the post-completion
ROADMAP-derived counters through the #2736 authoritativeFm seam (new object
direction for the progress key; completedOnlyRaise at the post-preservation
re-assert), because the transaction's disk scan reads the pre-completion
ROADMAP and failed to increment on the completing phase's own write.

* fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites

The pre-preservation merge could lower a correctly-higher disk-derived
counter (a verification-passed phase whose ROADMAP table row drifted behind
the disk signal). completedOnlyRaise now governs both application sites: the
intent and the derivation agree on direction (up), never on subtraction.

* fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal

The re-parsed derived block carries string scalars ("2") while the curated
snapshot carries numbers (2); substituting the curated spelling over an
equal derived one was a no-op in substance but a shape churn the ADR-3473
§8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived
warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater
curated counter replaces the derived value now; percent gets the same
verbatim rule.

* changeset(#4129): backfill PR 4359

---------

Co-authored-by: sim <sim@local>
2026-09-05 23:02:07 -04:00
Tom Boucher
06eba5fdb0 fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs

Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags,
discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and
parity properties (well-formed digit-prefixed ids parse to their exact id;
a non-digit injected into the prefix fails loud). Update the #2347
non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and
graduate the representative d5-prefix corpus fixture from could-not-parse
to parsed-but-uncovered.

All new rows are RED against origin/next; they go green with the parser
fix in the next commit.

* fix(#4130): parse phase-prefixed decision IDs (D4-01)

The three declaration grammars, the parse-miss guard, the #3939 join
regexes, and the token evidence all anchored on the literal 'D-' (or
'**D-'), so an ID carrying a digit-run phase prefix between the leading
letter and the hyphen matched nothing — while the #2347 shape detector
correctly called those bullets decision-shaped, collapsing the whole
CONTEXT.md to could-not-parse with 0 extracted instead of a coverage
verdict.

Derive the extractor ID grammar from one shared DECISION_ID_SOURCE
('D[0-9]*-' + the existing alnum tail, full id captured), widen the
guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial
prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose
like 'Deferred-until' stays none-present), and align the bare-token
evidence. Both gates and the gap-checker share the parser, so all three
surfaces read phase-prefixed decisions now; the gate messages name the
accepted forms including the phase-prefixed one.

* docs(#4130): document the phase-prefixed decision identifier form

The canonical CONTEXT.md reference said decisions carry 'a sequential
D-NN identifier' with no mention of the optional phase-number prefix the
parser now accepts (D4-01) or the alphanumeric tail it always accepted
(D-INFRA-01). Name both in the Decision identifier format section, EN
and ja-JP.

* chore(#4130): changeset

* chore(#4130): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 21:05:19 -04:00
Tom Boucher
0be5bf865a enhance(#3783): audit-uat summary segments current-milestone vs archived debt (#4336)
* test(#3783): add failing coverage for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3783): segment audit-uat summary into current_milestone and archived buckets

Additive: current_milestone/archived are new; total_items, total_files, parse_gap_files, by_phase, and by_category are unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): add changeset fragment for audit-uat summary segmentation

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#3783): allowlist the new audit-uat-summary-segmentation test file

lint-test-file-count.cjs baselines the "audit" module (keyed off bin/lib/audit.cjs)
at 6 pre-existing files; this adds the new dedicated suite as a 7th, matching the
module's existing one-file-per-feature-slice precedent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3783): fix phase/file number mismatch in the mixed-milestone fixture

The active phase fixture used dir "02-current" with file "01-UAT.md" — a
cross-phase stray per phase-id.cts's isPhaseArtifact/scopeToPhase (#3511),
so the file was silently excluded from the scan and current_milestone read
{files:0, items:0} instead of {files:1, items:1}. Confirmed by direct CLI
run against a hand-built fixture before recommitting. Renamed the file to
02-UAT.md to match its directory's phase number, matching every other
fixture in this suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3783): backfill changeset PR number to 4336

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 19:03:14 -04:00
Tom Boucher
c20675cc4d fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode

The pre-commit protected-branch assertion in the executor agent (#2924)
only fired inside a Claude Code worktree and matched a hardcoded
five-name branch list. It never ran in an ordinary checkout and never
covered this repo's own default branch ("next"), so gsd-executor could
commit planning-repo documents directly onto a shared checkout's
default branch with no PR ever created.

Widen the guard to run in every isolation mode, and resolve the
protected branch via the repository's actual default branch (with the
existing five-name list retained as a fallback when the resolver
itself cannot be invoked) plus any configured git.protected_branches.
Add a git.allow_default_branch_commits escape hatch for projects that
intentionally execute on their default branch. Also point the
separate <final_commit> commit helper back at the same guard, so it
cannot be sidestepped by that path.

Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3819): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:42 -04:00
Zy Deng
4c60879b5d fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources

* chore(#4132): record PR number in changeset

* test(#4132): cover rejected commands source alias

* fix(#4132): reject aliased package fallback

* test(#4132): cover rejected agents source alias

* test(#4132): cover partially aliased marker provider

* fix(#4132): reject partially aliased source providers

* test(#4132): cover routed source identity probes

* fix(#4132): route installed source identity probes

* refactor(#4132): tighten installer source metadata

* test(#4132): cover corpus trust boundary attacks

* fix(#4132): close installed corpus trust gaps

* refactor(#4132): keep installer authority private

* fix(#4132): preserve private installer fallback

* test(#4132): preserve fixture source authority

* fix(#4132): reject overlapping source fallback

* fix(#4132): avoid redundant installed corpus reads

* refactor(#4132): simplify provider resolution

* test(#4132): sync install tree fixtures after rebase

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 15:32:52 -04:00
Michel Moreira
86b745b48b fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 14:20:26 -04:00
Tom Boucher
7ff196c505 fix(#4096): honor --dry-run in todo complete and write completion keys inside the frontmatter fence (#4325)
* fix(#4096): honor --dry-run in todo complete and upsert completion keys inside the frontmatter fence

* review(#4096): tighten todo complete flag rejection to any dash-prefixed token

* chore(#4096): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 13:49:58 -04:00
Tom Boucher
3d03ae65e6 fix(#4094): withhold all four STATE.md progress counters under the milestone-unbounded guard (#4322)
* test(#4094): failing-first matrix for withholding all four progress counters

* fix(#4094): withhold all four progress counters under the milestone-unbounded guard

completed_phases/total_plans/completed_plans are accumulated from the same
phaseDirs walk as total_phases, so the #3354/#3573 withhold condition makes
them equally untrustworthy — yet only total_phases was withheld, and every
resyncing state.* write silently clobbered the three stored siblings with the
under-scoped disk numbers. Extend the withhold-then-fall-back-to-stored
pattern to all three siblings: null sentinels in the disk-scan cache value,
three new stored-counter readers threaded through all three
buildStateFrontmatter call sites, and the same cached-else-stored consumer
fallback. Milestone-bounded projects are untouched (gate-conditional).

* fix(#4094): scope-requires for the new test block, keep the (#3573) warning token, and update two #3578 rows to the withheld-counter contract

- the #4094 describe sat after the closing brace of the section that owned
  the module-level beforeEach destructure, so it needs its own local requires
  (mirroring the #3642 block);
- the #3573 warning keeps its literal '(#3573)' tag (asserted by an existing
  test) with '#4094' appended as a separate token;
- two #3578 status-guard rows in tests/state.test.cjs asserted the pre-#4094
  unconditional disk-scan assignment of completed_phases under the
  roadmap-absent withhold — exactly the silent clobber #4094 removes; the
  status-guard conclusion (must not fire) is unchanged, the counter-value
  assertions now pin the withheld contract.

* test(#4094): lint conformance — splitLines for the persisted-progress parser, local seeder, scoped rmSync disable

* changeset(#4094)

* changeset(#4094): backfill PR number

---------

Co-authored-by: sim <sim@local>
2026-09-05 13:01:28 -04:00
Tom Boucher
2e1ede6d99 fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline

* fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline

* refactor(#4093): collapse IIFE to a plain block (review finding)

* docs(#4093): document the advance-plan recovery decline + changeset

* chore(#4093): backfill PR number in changeset

* fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe

---------

Co-authored-by: sim <sim@local>
2026-09-05 10:46:37 -04:00
Atirna
70f22e4643 fix(#4213): keep STATE.md progress surfaces synchronized (#4231)
* fix(#4213): keep STATE.md progress surfaces synchronized

* fix(#4213): clamp the shared progress bar and keep bold-first priority, changeset + property tests

- formatProgressMachineSegment clamps through clampPercentFromFraction
  (ADR-3180 Decision 7 kernel) with a 0 floor, so a hand-edited
  out-of-range persisted percent renders a clamped bar instead of
  throwing RangeError on repeat() inside the write seam
- stateReplaceProgressPercent restores the #2177 bold-first priority:
  **Progress:** anywhere in the body wins; a plain ^Progress: line is
  the fallback, so free text starting with Progress: cannot capture
  the rewrite ahead of the real status line
- cross-reference comment names the three consumers and the
  cmdStateSync sanctioned exception (ADR-3408 §8.3)
- CONTEXT.md: applyPostSyncPreservation reconciliation documented in
  the STATE.md Transition Module entry
- property tests (never-throws/well-formed, idempotency, round-trip,
  bold-first) + two regression rows through the CLI

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 07:56:06 -04:00
aaka3207
294ec29857 fix(#4053): quote decimal-shaped frontmatter scalars for spec YAML readers (#4165)
* fix(frontmatter): quote decimal-shaped scalars so a spec YAML reader preserves them

A decimal phase identifier written to STATE.md frontmatter (e.g.
`current_phase: 22.10`) was emitted BARE, because `scalarNeedsDoubleQuoting`
only asks whether a value can OPEN a plain scalar — which `22.10` can. A
YAML-spec reader (js-yaml, the statusline, any external tool) then reloads bare
`22.10` as the float 22.1, colliding with `22.1` and dropping the trailing zero.
gsd's own tolerant line-scanner (`extractFrontmatter`) round-trips the raw text
and so hid the defect; a spec reader does not.

Fix: `reconstructFrontmatter`'s general scalar path now also quotes numeric-
looking strings that are not plain all-digit integers (decimals, exponents,
sexagesimal, hex/oct/bin) via `generalScalarNeedsNumericQuoting`, reusing the
existing `YAML_NUMERIC_RE`. Every all-digit string — integer counts, phase
numbers, and leading-zero fixtures like `02` — stays bare, so the state-rebuild
idempotency baseline and the rest of the state corpus are unchanged. This also
quotes `gsd_state_version: 1.0` on write, which matches the authoritative
STATE.md template (`src/state.cts` already emits it quoted).

Regression test drives the real write path and asserts, via js-yaml, that
`22.1` and `22.10` no longer collide and read back string-typed; guards that
integers and free-text stay unquoted.

Fixes #4053

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC

* chore(changeset): add Fixed fragment for #4053

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC

* docs(frontmatter): trim the generalScalarNeedsNumericQuoting comment

Cut the over-long doc block down to the essential why and drop the inline
comment that repeated it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC

* docs(test): drop the #4053 explanatory comments from the touched tests

The assertions speak for themselves; remove the added narrative comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC

* fix(#4053): correct the trade-off comment, changeset PR number, and cover every claimed numeric form

Review follow-ups (trek-e):
- The doc comment claimed a plain integer round-trips harmlessly. That is
  false for leading-zero values (`02` -> 2, `017` -> 17 under js-yaml). Rewrite
  it to state the real, deliberate trade-off: all-digit strings stay bare
  because zero-padded ids (`plan: 01`, `phase: 02`) are the pervasive GSD
  convention and quoting them all is the blanket quoting #4053 asked to avoid;
  the loss is padding not identity (`02` and `2` normalize to the same phase,
  `22.1` and `22.10` do not).
- Changeset carried the auto-closed draft's number (4151); correct to 4165.
- Test exponent, hex, octal, binary and sexagesimal forms through js-yaml, and
  pin the leading-zero trade-off so the documented behaviour is asserted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qc7VN4zTpTSDTS9JXM2cFB

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 11:17:41 +00:00
Dennis Alexis Valin Dittrich
5869febb16 enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change

readVerificationStatus() now recomputes a deterministic sha256 fingerprint
over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY,
requirements, implementation files in the verified change set) and returns
stale on any mismatch, fail-closed when a covered file is missing,
unreadable, or escapes the project root. Legacy reports with no fingerprint
metadata keep the prior SUMMARY-mtime staleness check unchanged.

The verifier computes covered_digest via the new verification.fingerprint
CLI command rather than by hand, since a digest is deterministic math, not
an LLM-estimated value.

* chore(#4155): backfill fork PR number in changeset

* fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap

* fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior

Partial fingerprint metadata (one of covered_files/covered_digest present,
the other missing or malformed) now fails closed to stale instead of
silently downgrading to the legacy mtime-only check. computeCoveredDigest
also canonicalizes with realpathSync before re-confining, so an in-root
symlink whose target escapes the project root can no longer produce a
matching digest. gsd-verifier.md restores the completeness requirement and
checklist item trimmed by the earlier size-budget fix, within the LARGE
tier byte cap.

* chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap

* fix(#4155): address gemini adversarial review findings

computeCoveredDigest now threads the caller-supplied opts.fs seam through
its confinement and read paths instead of always using raw node:fs — a
caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP
2, #2790 follow-up) was silently bypassed for covered-input reads. The
project-root anchor itself still canonicalizes through real fs (it is a
trusted value the caller derived, not attacker-influenced covered-input
data); only per-file candidate reads go through the injected seam.

Covered-file paths are now canonicalized (./ prefixes, redundant slashes,
internal .. segments) before becoming dedup/sort/hash keys or confinement
subjects — closes both a spurious-stale false positive (two spellings of
the same file hashing differently) and a confinement gap (an internal ..
segment that doesn't start the string).

gsd-verifier.md now states covered-file paths are project-root-relative,
not phaseDir-relative, closing an ambiguity that would have made a real
verifier agent's first fingerprint invocation fail closed.

defaultFsImpl's methods now late-bind through fs.<method> rather than
capturing function references at module load — the earlier direct-capture
form was invisible to existing tests' t.mock.method(fs, 'statSync', ...)
seams, a real regression caught by the full suite (not the reviewer).

* fix(#4155): catch a plan/summary added to the phase dir after verification but never declared

The content digest only recomputes hashes for paths the verifier actually
declared in covered_files — it had no way to notice a plan or summary
added to the phase directory after verification if that new file was
never declared, silently regressing behind the legacy mtime check it
replaces (which scans the live directory, not a declared list).

findUncoveredCurrentArtifact re-scans the live phase directory for every
current *-PLAN.md/*-SUMMARY.md and requires each to be represented in
covered_files, closing that gap; a directory scan failure fails closed to
stale rather than silently skipping the check.

CONTEXT.md's Verification Module entry corrected to describe the
fingerprint path's stricter fail-closed FS-error contract (routes to
stale) instead of the module's original degrade-to-safe one (missing /
not-stale), which only the legacy path still keeps.

* refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test

computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/
sorted covered_files independently — one shared helper now backs both
(gemini review's ponytail-lens finding).

Adds one CLI-to-readVerificationStatus test against a genuine
.planning/phases/NN-x/ project with an implementation file outside
.planning/ entirely, closing the review finding that prior #4155 unit
fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot
falls back to phaseDir itself there) and never exercised real multi-level
path resolution.

* fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/

Two independent review rounds (opus critical-reviewer + opus ponytail +
agy, run twice) found two instances of the same fail-open class:

- computeCoveredDigest's per-file reads routed through the caller's
  injected fsImpl. planning-inspect.cts passes a `.planning/`-confined
  containment fs into readVerificationStatus's opts.fs, so any covered
  implementation file outside `.planning/` (mandatory per the issue)
  made the confinement wrapper throw, which was caught and turned into
  a stale digest -- reporting every fingerprinted phase permanently
  stale via `planning.inspect`, regardless of actual drift. Per-file
  reads now always use real node:fs, matching the pre-existing
  treatment of root canonicalization; the realRel-vs-realRoot check is
  the real confinement boundary for this data and needs no seam.

- allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans
  reports readdir failures via a `scope` field, it never throws), so
  an unreadable nested plans/ dir was silently treated as "zero
  artifacts, all covered" instead of failing closed. Now branches on
  scope !== SCOPE.COMPLETE.

Also, per ponytail's second-round findings: reverted an unwarranted
FINGERPRINT_VERSION bump and digest length-prefix from the first fix
(no v1 digest has ever existed -- the feature is unreleased -- and the
prefix closed a collision that grants no capability beyond what a
writer of covered_files already has more cheaply); removed a
verifier-facing escape-hatch instruction whose own example was a case
that should trigger staleness, not bypass it; corrected CONTEXT.md
references to the renamed allCurrentArtifactsCovered and a stale
"unconditional" rescan claim; simplified the isStale derivation,
removed dead FsLike members, and tightened test coverage.

Regression tests for both fail-open bugs are included and were each
confirmed to fail against the pre-fix code before the fix landed.

full test suite: 2558/2560 pass, 2 skipped, 0 fail

* fix(#4155): trim gsd-verifier.md under the LARGE size cap

Fork CI caught what my local runs missed: the superseded/nested-plans
instruction added earlier pushed gsd-verifier.md to 49299 bytes,
147 over the LARGE tier's 49152-byte hard cap
(tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's
wording and dropped a redundant inline comment tag; no content lost.

* chore(#4155): point changeset at the upstream PR number

pr: 19 was the fork PR opened for internal review-lane CI; now that
open-gsd/gsd-core#4290 exists, the changeset field must match it per
CONTRIBUTING.md's release-notes convention.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:42:52 -04:00
Behruz Nassre Esfahani
5ad9a36f35 fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by
spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id
was a hardcoded literal, so `--host` chose only the argv RENDERING while the
LEVEL always came from the installed plan-checker's frontmatter — `low` under
every shipped model profile. Every prompt-fed lane therefore ran at a fast
structural verifier's effort, and because the rendered argument is a CLI config
override it silently beat the effort the operator had configured for that CLI.
At `low` a large source-grounded prompt makes a model end its turn with no final
message, so the lane came back empty and its stub read as a crash.

Effort is a property of the review, so the lane declares it. Two new fields on
ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` —
carried through each capability manifest and the generated registry, set on the
three lanes with an argv effort channel and null on the other nine. A new pure
`resolveLaneEffort()` resolves config key -> lane default -> nothing, where
"nothing" emits no effort argument at all and the reviewer CLI's own
configuration decides; `inherit` selects that path explicitly and an
unrecognized level falls back to the lane default rather than being forwarded to
a CLI that would reject it. The host's negotiated effortSurface still gates the
rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving
in-process also removes up to twelve subprocess spawns per review.

The empty-output stub now names the effort the lane ran at and distinguishes a
clean exit from a timeout kill, a non-zero exit, and a process that never ran —
`status` is null for both a timeout and a signal, so those were indistinguishable
before. The hint is hedged: a clean empty exit is most often a model stopping
short, but it is also consistent with a CLI writing its output elsewhere.

Also: the capability validator now knows both fields, rejects a malformed key or
an out-of-vocabulary default, and rejects a default declared without a config
key (a level the operator could never override). An existing end-to-end row in
tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures
the lane's own key and pins the decoupling in the same real spawn, with the
agent execution tier set to a level that must not appear.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive.

Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:25:44 -04:00
Dennis Alexis Valin Dittrich
925a363879 enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract

Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation.

* feat(4032): apply configured agent tool grants during staging

Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion.

* test(4032): cover host grant and quoted MCP contracts

Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants.

* feat(4032): apply configured agent tool grants across runtimes

Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy.

* fix(4032): register agent tool grants in configuration

Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper.

* fix(4032): translate configured MCP grants for Kilo

Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies.

* fix(4032): decode YAML-escaped tool grants

* fix(4032): emit valid inline agent tool grants

* fix(4032): reject invalid trailing-colon grants

* test(#4032): cover cross-review remediation gaps

* fix(#4032): close cross-runtime grant gaps

* test(#4032): expose Kimi global project context

* fix(#4032): preserve Kimi project config context

* chore(#4032): add release note

* test(#4032): expose fork review regressions

* fix(#4032): address fork review findings

* test(#4032): make byte-stability assertion portable

Compare repeat installs at one root so platform-specific path rendering cannot
masquerade as an agent_tools behavior change.

* chore(#4032): bind changeset to upstream PR 4238

* fix(#4032): address trek-e review findings (2,3,4,5,6,7,8)

Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper,
a comment-only `tools:` header mis-parse that silently dropped
configured grants, and a naive comma-split that could tear a quoted
scalar containing a literal comma. Documents Kilo's inherent
`{server}_{tool}` MCP-permission-key collision (external, fixed
format — not ours to widen) and locks the existing first-seen-wins
resolution in with a regression test.

Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite
step: routing Kimi through that pipeline (needed so project-scoped
agent_tools selectors reach it) was short-circuiting Kimi's own
neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core
text rather than a pre-rewritten Kimi path.

Extends the fast-check token pool and per-runtime install coverage
with the missing comment/comma/broad-runtime cases the prior review
flagged as untested.

* docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step

Documents readGsdEffectiveAgentTools (Install Model Override Resolver
Module) and the appendAgentTools pre-converter pipeline step (Runtime
Artifact Conversion Module), per contributor-standards.md's
new-seam glossary requirement (finding 1).

* fix(#4032): address agy adversarial review findings

An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix
commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps,
plus a genuine new regression and two CONTEXT.md inaccuracies:

- ZCode's comment-only `tools: # note` header matched the inline-value
  branch instead of falling through to the block-list scan, so a following
  mcp__* item leaked through unstripped — the exact defect finding 4 fixed
  in appendAgentTools, unfixed in this sibling function.
- Reverted capabilities/kimi-code/capability.json's noPathRewrite: true.
  kimi-code uses the standard 'agents' kind with converter: null (not
  kimi-agents — confirmed by reading the descriptor, not its prose
  description), so it never went through the pipeline change finding 5
  fixed, and disabling its path rewrite broke every ~/.claude/ embed in
  its shipped agents instead.
- decodeToolScalar never stripped a trailing ` # comment` from a bare
  (unquoted) scalar, so a comment after a block-list item, or after an
  appended grant on an inline line, became part of the "tool name" —
  fixed at the source (one call site fixes every consumer).
- appendAgentTools's comment-index scan wasn't quote-aware, so a `#`
  inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment
  start and corrupted the quote.
- parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of
  appendAgentTools's own output) had the same naive comma-split and
  comment-only-header gaps as findings 4 and 6, unpatched.
- The all-runtime smoke test's presence assertion was built on a guessed
  omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__
  grant recognizable, replaced with a verified allowlist.
- CONTEXT.md claimed a `project:<agent>` selector prefix that does not
  exist (project override is a same-key merge across two config files)
  and mislabeled stageAgentsForRuntimeWithConverter's module.

* fix(#4032): address full-PR review (Opus critical/ponytail + agy)

A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a
second agy full-source adversarial pass) surfaced defects the earlier
finding-scoped passes couldn't reach:

- appendAgentTools corrupted a `tools:` line whose ENTIRE value is a
  leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`,
  invalid YAML) — there is no safe line-surgical rewrite here, so it now
  refuses to touch that shape instead of emitting broken frontmatter.
- decodeToolScalar's malformed-trailing-quote check ran BEFORE comment
  stripping, so a bare tool name with a quote inside its own trailing
  comment (`Bash # note: "internal"`) was wrongly rejected. Reordered.
- findUnquotedCommentIndex (added in the prior remediation commit) was
  built on a wrong model of YAML: a `#` after whitespace starts a real
  comment in a plain scalar regardless of nearby quote characters —
  verified against the actual parser. The one case that DOES need
  protection (a leading quoted scalar) is now refused outright above, so
  the quote-tracking scan was dead weight solving a problem that no
  longer reaches it. Removed; reverted to the plain `[ \t]#` scan.
- Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter,
  distinct from the buildKiloAgentPermissionBlock fixed earlier) with the
  same comment-only-header and naive-comma-split gaps as findings 4 and 6
  — unfixed in both its src/ and bin/install.js copies. Fixed in both,
  exporting splitToolScalars for bin/install.js to reuse rather than
  reimplementing it.
- Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5
  steps, omitting appendAgentTools (now step 3 of 6).
- docs/CONFIGURATION.md didn't state that a --global install still
  discovers agent_tools from the cwd's .planning/config.json (confirmed
  intentional and already covered by a dedicated test, not a bug).
- Removed install-engine.cts's deps.cwd injection seam: zero callers or
  tests ever populated it.

Two claims from this round were verified and rejected, not fixed:
prototype pollution via a `__proto__` selector key (empirically confirmed
`Object.prototype` is never touched — only reassigns the resolver's own
local object's prototype, with no observable effect), and a `*` grant
value crashing YAML parsing as an alias reference (empirically confirmed
it parses as plain scalar text, no crash). A pre-existing, unrelated
defect (extractFrontmatterField returns null for block-list `tools:` on
Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents
today) was filed as a follow-up rather than fixed here — it predates
#4032 and isn't caused or worsened by this PR.

* fix(#4032): update stale slug-derivation-drift-guard fixture line

normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a
side effect of this PR's edits to runtime-artifact-conversion.cts; the
MAJOR-1 fixture's hardcoded realEndLine had gone stale.

* fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools

bin/install.js's installAgentsKindStandalone call site omitted the projectDir
argument the function already supports, so a global install through this
legacy branch silently fell back to the runtime config dir instead of
process.cwd() when resolving project-scoped agent_tools grants — inconsistent
with the sibling installOpencodeFamilyArtifacts call site, which already
threads it correctly.

appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow
sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the
in-sequence commas and appended past its closing bracket, producing invalid
frontmatter. Extended the bailout regex to also refuse a value starting with
`[`, matching the same "whole node, nothing may follow" reasoning already
applied to quoted scalars.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:52:45 -04:00
Dennis Alexis Valin Dittrich
e8800287d5 enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target

* fix(#4153): fail closed unresolved update target

* test(#4153): require a concrete recovery installer

* fix(#4153): use concrete unresolved recovery command

* chore(#4153): bind changeset to fork PR

* test(#4153): cover portable update diagnostics

* fix(#4153): keep update diagnostics portable

* fix(#4153): harden update version diagnostics

* test(#4153): reject jq in update version checks

* test(#4153): expose step-local parser gap

* fix(#4153): keep JSON parsing step-local

* docs(#4153): align update target guidance

* test(#4153): expose workflow runtime fallback

* test(#4153): expose resolver runtime fallback

* fix(#4153): leave unknown workflow runtime empty

* fix(#4153): stop inferring Claude for unknown targets

* test(#4153): preserve Claude workflow targeting

* test(#4153): preserve known runtime directory identity

* fix(#4153): recognize Claude workflow paths

* fix(#4153): reuse known runtime directory identities

* chore(#4153): acknowledge emitted workflow growth

The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes.

Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation

* test(#4153): expose missing Windsurf workflow contract

* docs(#4153): document Windsurf update targets

* chore(#4153): bind changeset to upstream PR

* fix(#4153): gate unresolved-target exit before the VERSION-missing fallback

The VERSION-missing bullet in get_installed_version sat before the
UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version
0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy
"proceed to install" without ever reaching the fail-closed exit this
PR adds, reopening the ill-defined mutating path #4153 closes. Reorder
so the unresolved-target gate runs first and scope the VERSION-missing
bullet to require an already-resolved target.

Also drop two vacuous mutationSpies entries: they checked '--sync'/
'--reapply' (commands/gsd/update.md content) against `step`, a slice of
workflows/update.md — always -1 regardless of correctness. Those routes
bypass get_installed_version entirely and are already covered by
install.test.cjs, reapply-patches.test.cjs, and
skill-frontmatter-contract.test.cjs.

* chore(#4153): point changeset pr field at fork PR #10 for fork CI

* test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent

Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS
(src/update-context.cts) are two independently maintained copies of the
same runtime->dir mapping with no parity check; add one so a future
edit to either surface without the other fails loudly instead of
silently drifting.

Nit 2: call out in the changeset that a custom --config-dir matching no
known runtime, marker file, or env var now resolves unresolved instead
of silently defaulting to claude -- this narrowing is intentional, it's
the fail-closed behavior #4153 asks for.

* fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path

agy (gemini-3.8-flash-high) adversarial review of the full PR:

1. check_latest_version's uc_field() copy-pasted get_installed_version's
   `${2:-$UC}` fallback, but every call site here passes $2 explicitly and
   $UC does not exist in this step's scope -- dead, misleading reference.
   Use $2 directly.
2. No unit test covered resolveUpdateContext's preferredConfigDir fast path
   returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS
   suffix, marker file, or env var (the exact fail-closed case #4153 adds).
   Added.

A third finding (update.md:90 using /gsd:update vs docs using /gsd-update)
was investigated and rejected: /gsd:update is the actual registered
Claude Code command name (commands/gsd/update.md name: gsd:update) and is
locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update
is a separate, pre-existing, intentional prose convention used in
audience-facing docs (README/INVENTORY/FEATURES). Not a defect.

* chore(#4153): backfill changeset pr field to upstream PR #4237

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:17:21 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
Behruz Nassre Esfahani
c6efe2905c fix(#4087): stage the hook helpers the Codex bundle's hooks require (#4117)
* fix(#4087): stage the hook helpers the Codex bundle's hooks require

CODEX_HOOKS_TO_COPY is a flat, hand-maintained filename allowlist that never
recursed, and Codex is excluded from installSharedHooksBundle() — the path that
stages hooks/lib/ for full-bundle runtimes — by an !isCodex gate. Excluding
hooks/lib/ was a correct scoped decision for #3579 until #3911 (2ea5efc15) gave
gsd-context-monitor.js a real require('./lib/hook-exit.js'). From then on every
fresh --codex install staged the hook without its helper and the hook died with
MODULE_NOT_FOUND at module load, before its own try/catch, on every event Codex
registers it for. The install still exited 0, so nothing surfaced it.

Reproduced before changing anything, in a sandboxed CODEX_HOME: four hooks
staged, no lib/, and the installed hook exiting 1 on "Cannot find module
'./lib/hook-exit.js'".

Rather than hand-add today's three helpers — which re-breaks the next time a
Codex-bundled hook grows a lib dependency, exactly how this regressed — the
transitive-require walk already written for Cursor in 704859e9c is extracted out
of writeCursorHooksJson into an exported stageTransitiveHookLibs(), Cursor is
rewired onto it, and the Codex copy loop calls it. bin/install.js already
required that module, so this adds no new seam. Cursor's staged set is
byte-identical to base, compared file by file.

Extraction surfaced a latent defect in that walker, fixed here: its regex read
`./X` and `./lib/X` identically, but from a hook SCRIPT a bare `./X` is a
sibling in hooks/ — gsd-check-update-worker.js requires
`./managed-hooks-registry.cjs`, which is not a lib — so it demanded
hooks/lib/managed-hooks-registry.cjs and the fail-loud guard threw. Seeds now
match only `./lib/X`; lib files still match both, which is the
sibling-within-lib case 704859e9c exists for. Cursor never exposed it because
none of its scripts carries a bare sibling require.

Three further grammar gaps closed after review, each in the fail-closed
direction: an extensionless `require('./lib/x')` is valid CommonJS and was
resolved literally, failing the install on a legitimate require — now resolved
through .js/.cjs and written under its resolved name; a NESTED `./lib/sub/x.js`
could not be expressed by the character class and was a SILENT miss, the one
failure mode this function exists to remove — now refused loudly; and a capture
carrying no alphanumeric character is prose, not a module name — hooks/lib/
injection-patterns.js documents this very mechanism with the literal string
require('./lib/...'), which captured `...` and sent the resolver hunting for
hooks/lib/... . The scan is still not comment-aware, which is disclosed at the
call site rather than papered over.

Seeded from the entries THIS invocation staged rather than probing the
destination, so a file left by an earlier install whose source is no longer
allowlisted cannot contribute helpers for a hook that is no longer shipped.

The #3579 boundary holds: three of ten helpers ship, gsd-graphify-rebuild.sh
among those correctly absent. Seven rows — three driving a real install into a
sandboxed config dir (with HOME sandboxed for the child, since Codex's skills
kind resolves from os.homedir() and the #3712 guard rightly refuses otherwise)
and four pinning the discovery grammar directly. All proven fail-first; the
set-equality row also reds on over-staging, which the count-based version it
replaced did not catch.

Fixes #4087
Fixes #4098

Emitted-Drift-Ack-Hash: hooks/lib/hook-exit.js — newly emitted for codex because the installer now stages the helpers its hooks require; the helper's own content is unchanged
Emitted-Drift-Ack-Hash: hooks/lib/cli-exit.js — newly emitted for codex as hook-exit.js's transitive require; the helper's own content is unchanged
Emitted-Drift-Ack-Hash: hooks/lib/exit-code-registry.js — newly emitted for codex as cli-exit.js's transitive require; the helper's own content is unchanged
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* chore(#4087): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* fix(#4087): stage the hooks/lib helpers the Windsurf guards require

Review of #4117, verified as asked and reproduced against a real install.

Windsurf sets hostBehaviors.skipSharedHooksInstall, so like Cursor it never
reaches installSharedHooksBundle -- the only other stager of hooks/lib --
and writeWindsurfHooksJson staged its two Cascade guards without the
helpers both require at module load: gsd-windsurf-pre-write.js requires
./lib/hook-exit.js and ./lib/git-probe.js, gsd-windsurf-pre-command.js
requires ./lib/hook-exit.js. stageTransitiveHookLibs had one call site,
Cursor's.

Measured on a fresh `--windsurf --global` install into a sandboxed HOME:
the installer exited 0, hooks/ held only the two scripts and package.json,
and executing either installed guard exited 1 with "Cannot find module
'./lib/hook-exit.js'" -- so every pre_write_code and pre_run_command event
failed at load while the install reported success. The same command with
`--cursor` staged four helpers and its hook ran, which is the control.

Pre-existing rather than introduced here: at merge-base 05092ff36 the same
three require lines exist and writeWindsurfHooksJson already staged no
lib/, and this PR's diff carried no reference to Windsurf. Fixed here
anyway because the helper this PR extracted is the right tool and a second
runtime is a few lines onto it.

writeWindsurfHooksJson now calls stageTransitiveHookLibs after staging its
scripts, with the same gsd: -> gsd- transform the scripts receive, so a
helper is rewritten the same way as its caller. The install-tree fixture
regenerates with exactly hook-exit.js, git-probe.js, cli-exit.js and
exit-code-registry.js added and no other fixture moved. Two rows execute
the INSTALLED guards, beside the Codex rows they mirror; the existing
windsurf-hooks-bridge rows run the guards from source and test behaviour,
a different question, and are left as they are. Both new rows fail-first
against the unfixed compiled artifact -- gsd-core/bin/lib, which is what
bin/install.js loads -- on the MODULE_NOT_FOUND assertion.

Emitted-Drift-Ack-Hash: hooks/lib/git-probe.js — first staged for Windsurf, whose pre-write guard requires it; the file itself is unchanged
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:49:37 +00:00
Andreas Brauchli
0ea012c519 fix(#3939): parse decision bullets with a wrapped bold lead-in (#3953)
* fix(#3939): parse decision bullets with a wrapped bold lead-in

parseDecisionLines matched every PHYSICAL line against the three decision-bullet
grammars, and all three require the closing `**` in the same string as the
`- **D-` anchor. A declaration whose bold lead-in wraps across a line break —
the shape discuss-phase itself writes whenever a decision title runs past the
wrap column — matched none of them and fell to the #1365 parse-miss guard, which
forces `could-not-parse` and hard-blocks check.decision-coverage-plan on a
well-formed CONTEXT.md.

Fold physical lines into logical bullets before matching: a declaration whose
bold lead-in is still open at end-of-line absorbs following lines until that run
closes. The three grammars are untouched, so every single-line form parses
exactly as before.

Joining is bounded and preserves the fail-loud contract. A blank or
whitespace-only line, any block-level construct (a list marker of any family,
an ATX heading, a blockquote, a table row), or the end of the block stops it,
and a lead-in that never closes is emitted unchanged — so a genuinely malformed
bullet still reaches the parse-miss guard and still fails loud (#1365), and
cannot be "closed" by an inline `**` belonging to the block below it. The joined
line keeps the first physical line's indent, so the nested cross-reference
signal (#3169) is unchanged. Absorbed lines are scanned once each rather than
re-searching the accumulated candidate, keeping a pathological unterminated run
linear on the plan gate's hot path.

Regression coverage lands in tests/decisions.test.cjs (the owning module's file,
per the regression-test placement policy): all three grammars wrapped, a
three-line wrap, tags/category/continuation preservation, one-line parity
(including inline bold and emphasis inside a wrapped title), CRLF, the
markdown-header path, plus negative proof that every join terminator still
yields could-not-parse and that the FIX-B and #3169 fixtures are unchanged.

Fixes #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3939): add changeset fragment for PR #3953

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3939): property-test the wrap-position invariant

Review follow-up: RULESET.TESTS.property-based-testing requires a parsing /
transformation contract to carry at least one fast-check property asserting a
domain invariant, and the join added by the fix is exactly such a
transformation. The example-based tests pinned four hand-picked wrap points;
these generalize over the whole dimension.

Three properties, on the shared tests/helpers/fast-check-setup.cjs config
(numRuns 200, seeded):

- round-trip: for every grammar (colon-immediate, titled-colon, em-dash), every
  id shape, every tag, with and without a category heading, wrapping the bold
  lead-in at ANY interior space is deepStrictEqual to not wrapping it — where a
  line happens to break carries no information;
- domain invariant: a well-formed wrapped declaration never reaches the
  parse-miss guard (outcome `parsed`) and keeps its declared id;
- fail-loud preservation: an unterminated bold run followed by 0-12 prose lines
  still yields `could-not-parse` with no decision manufactured, however many
  lines the join would have to absorb before giving up.

The corpus is deliberately free of markdown metacharacters: `:` and `*` select a
different grammar (#1639's `[^:*]*` discipline) and a block-construct token
legitimately terminates the join. Both are separate behaviours, example-tested
above; these properties isolate the wrap-position dimension.

Rebuilding the module from `next` with these in place fails 14 (was 12); the two
new failures are the round-trip and never-a-parse-miss properties. The fail-loud
property passes before and after, which is the point of it.

Refs #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3939): fail loud when a wrap splices a decision tag token

Addresses review rounds 2 and 3 on PR #3953.

Folding a soft line break to a single space is markdown's own rule and is
invisible everywhere in a decision bullet except inside the id-adjacent
`[tags]` bracket, which the three grammars turn into `tags` and therefore
into `trackable`. There a spliced space splits one tag token into two
(`[defer` + `red]` -> `defer red`), which does not fail: it parses to a
DIFFERENT tag, silently flipping whether check.decision-coverage-plan
demands coverage for that decision.

The join now stops at such a splice, so the bullet reaches the #1365
parse-miss guard and fails loud instead of guessing. The check is
delimiter-aware, so wraps that land next to `[`, `,` or `]` still join and
still parse identically to the one-line bullet -- a comma-separated tag
list may wrap at any of its separators, across any number of lines. A
bracket further along the title is ordinary text and does not restrict the
join.

Also in this round:

- blockConstructRe's doc comment claimed parity with the sectionizer seam's
  `iterateBullets`, which recognises only the `N. ` ordered form while this
  set also stops at `N) `. The widening is deliberate and one-directional
  (a terminator set may recognise more block openers than a bullet iterator;
  a spare terminator can only make a malformed bullet fail loud, never
  manufacture a decision). Comment corrected to say so, both marker forms
  now tested, and a drift guard asserts the seam still does not yield `N)`
  so the divergence cannot widen silently.
- Documented that the table-row alternative deliberately has no trailing
  whitespace requirement (CommonMark tables may open flush), and that
  over-termination on prose opening `10.` or `|` is accepted fail-loud
  behaviour -- now pinned by a test.
- Coverage the review asked for: a WRAPPED bold lead-in nested under an
  already-open decision (#3169, the existing guard used a single-line nested
  bullet), and title/body whitespace fidelity across every wrap position
  around a double space.
- A fourth fast-check property: wherever a wrap lands inside a `[tags]`
  bracket, the parse either matches the one-line bullet exactly or fails
  loud with nothing extracted -- never a decision whose tags differ.
- Property helpers render through `renderBullet`, which asserts the form
  exists instead of letting an unchecked map lookup yield undefined.

Fail-first: tests/decisions.test.cjs run against origin/next's decisions.cts
fails 18 of 131; against the previous PR head it fails the 2 new tag-splice
guards. All 131 pass with this change. Real-world CONTEXT.md from the report
is unchanged at 37/44 parsed.

Refs #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3939): arm the tag-splice guard on any wrapped line, not just the first

The #3953 round-3 guard read the id-adjacent `[tags]` bracket only from a
bullet's FIRST physical line, via a regex anchored to the bullet start. A
lead-in that wraps twice can open that bracket on a LATER absorbed segment,
where the guard was never armed and `wouldSpliceTagToken` became a no-op:

    - **D-01
      [inform
      ational]: A title.** body text here.

folded to the tag `inform ational` and `trackable: true`, where the one-line
form gives `informational` and `trackable: false` — a silently wrong answer to
the coverage gate, with no thrown error and no parse-miss to signal it. Exactly
the re-classification the round-3 guard exists to prevent, for the case it did
not cover.

`tagBracketOpenAtEolRe` becomes `tagRegionRe`, which asks whether the
id-adjacent bracket REGION is still unsettled rather than whether it opened on
one specific line: group 1 present means the bracket is open, group 1 absent
means the id is read but a `[` may still follow. `joinWrappedBoldLeadIns` keeps
the assembled text in `tagRegion` only while the bracket has yet to open, so a
bracket opening on any segment arms `tagTail`; once armed, the pre-existing
O(1) tail update takes over and `tagRegion` is dropped. A non-empty segment
that is not a bracket-open settles the region immediately, so this bounds the
string to a single extra join and leaves the 5000-line unterminated run linear.

The id class widens to admit an empty id, so a bare `- **D-` still counts as
unsettled. This regex only answers "may an id-adjacent bracket still open
here?", where matching MORE shapes is the conservative direction: an over-broad
match can only make a malformed bullet fail loud, a missed one re-classifies
silently.

The existing property test wraps at exactly one point, and only at spaces —
which round-trip exactly, since the join re-inserts the space it replaced — so
neither the bracket-opens-later state nor an observable splice was reachable
from it. `wrapBoldLeadInMulti` breaks at two or more arbitrary positions after
the id and asserts the same disjunction: parse identically to the one-line
bullet, or fail loud with nothing extracted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCPSU591zVS9vLPd3gKnqn

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 01:18:36 -04:00
Tom Boucher
eb336e9f77 fix(#4081): decode git C-quoted paths in codebase-drift --name-status parser (#4307)
* test(#4081): failing-first regression for quotepath C-quoted paths in codebase-drift

* fix(#4081): decode git C-quoted paths in codebase-drift --name-status parse

* test(#4081): set drift_threshold 1 so decoded-path test triggers action_required

* chore(#4081): add changeset fragment

* chore(#4081): fix changeset fragment formatting

* chore(#4081): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 00:50:35 -04:00
Tom Boucher
02ad0b91f3 fix(#4078): phase.complete next-phase cascade reads dash-grammar checkbox rows (#4301)
* test(#4078): phase.complete mixed-grammar roadmap picks lowest outstanding phase, not positional-last

* fix(#4078): accept dash-grammar checkbox rows in phase.complete next-phase cascade

Stage 2 (roadmap identity scan) and stage 3 (#2028 lowest-outstanding
override) required a colon separator after the phase number, while the
canonical phase lookup has accepted the bullet-house dash grammar
(- [ ] **Phase N — Name**, #2199) for years. On a mixed-grammar roadmap
the only parseable row above N was a later phase.add-ingested colon-form
phase - positionally last - and it won the numeric-minimum vote it should
never have been alone in: completing Phase 1 of 18 selected Phase 18 and
skipped phases 2-17 (#4078).

The checkbox branches now accept the #2199 separator class (em/en-dash,
hyphen, colon); heading branches stay colon-only, mirroring
findRoadmapPhaseInContent exactly.

* test(#4078): align regression fixtures with slug name + checked-box semantics

* fix(#4078): drop unnecessary type assertion flagged by eslint

* chore(#4078): add changeset fragment

* chore(#4078): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 22:03:43 -04:00
Tom Boucher
04ac8723b9 fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure (#4288)
* test(#4024): pin quantitative-criteria trap shapes for verify plan-structure

Rows 1-3 and 20 of the #4024 test matrix reproduce the issue's shapes
(exact grep -c counts, bulk all-N observed-failing claims) and are
expected to FAIL against unmodified next: nothing judges these shapes
today. Corrected-arm rows pin that each rule is silent on its own fix.

* fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure

Add scanQuantitativeCriteria, the third plan-discipline scanner in the
cmdVerifyPlanStructure family (#429, #968). It judges criteria text in
<acceptance_criteria>/<automated>/<verify> blocks against a six-rule ban
list of shapes proven to be traps at HEAD: exact grep -c counts (R1),
bulk all-N observed-failing claims (R2), unquoted $VAR in command
position (R3), fallible git swallowed by a non-final pipeline stage (R4,
warn), wc output compared by string equality (R5), and relative HEAD~N
git anchors (R6; bare git diff warns). Legitimate exit:
<!-- plan-criteria-allow: R# - reason -->. Pure text scan, fail open.

* test(#4024): bind node:test before hook locally below the fold-point

* fix(#4024): R3 command-position anchor tolerates list bullets and inline-code backticks

* test(#4024): bind VERIFY_CJS locally in the unit block instead of relying on fold scope

* fix(#4024): R6 argument span ends at inline-code backtick or redirection

* fix(#4024): satisfy no-adhoc-markdown lint on the R4 stage-boundary regex

* chore(#4024): add changeset fragment

* chore(#4024): backfill PR number in changeset fragment

* fix(#4024): escape backticks in regex literals so drift-lint tokenizers keep function attribution

---------

Co-authored-by: sim <sim@local>
2026-09-04 21:04:49 -04:00
Tom Boucher
2e056488d9 fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED)

Five-case matrix: decline on unsummarized plans (regression), fire on
fully-summarized phase, fail-open on unresolvable phase dir, idempotent
decline, normal advance untouched.

* fix(#4067): derive advance-plan phase-complete from disk, not the plan counter

The phase-complete branch of state.advance-plan was decided purely by
STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale
counter carried into a newly planned phase, or a counter raced by
wave-parallel executors, let 'Phase complete — ready for verification'
land while sibling plans were still executing.

cmdStateAdvancePlan now re-decides that branch from disk before the
write: every plan in the Current Position phase's directory must have a
SUMMARY.md (scanPhasePlans single owner, the same source
state.update-progress recalculates from). Outstanding plans decline the
entire write byte-identically (idempotent, concurrency-safe, counter
stays display-only); an unavailable disk answer fails open to the
counter-derived decision.

* fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs

#3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed
(current-milestone) lookup first so an archived milestone's stale dir
cannot shadow the live one; unscoped retry when the window cannot
answer. Also restore the transform's undefined-data error semantics and
extract scanOutstanding.

* chore(#4067): add changeset fragment

* chore(#4067): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-04 17:49:01 -04:00
Tom Boucher
8249ebcf6e fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate

RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do
not exist yet; every row fails on require. Per #3770 only an intentional
target-test failure may authorize GREEN; zero-test discovery, fixture crashes,
unrelated failures, and unexpected green are INVALID_RED.

* fix(3770): require intentional RED evidence before GREEN

Only an intentional failure of the TARGET test (distinctly named, TAP-reported
assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test
discovery, fixture/load crashes (file-named failures), nonzero exits without a
failing test, unrelated failures, unexpected greens, and malformed/missing
records are INVALID_RED and block GREEN.

- src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses
  the prohibition-enforcement TAP primitives; fail-closed, never throws)
- check tdd-red-evidence <record.json>: validates the persisted record
  (command, exit code, failing test, expected, actual)
- gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now
  requires the evidence record + gate verdict, not a nonzero exit or a RED: tag

* chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs

* fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib

- gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B <
  49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid)
- tests: the row-6 fixture used String.replace (first-occurrence), so the
  `not ok` line still named the target test and the classifier was right to
  accept it; replaceAll makes the failure genuinely unrelated
- eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs
  (lint the src/*.cts source, per ADR-457 migration rule)

Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line

* chore(3770): add changeset

* chore(3770): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:44:38 -04:00
Tom Boucher
580059251a fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing

Red: init.progress/init.resume/init.new-project payloads carry no
partial-init discriminator, and progress.md/resume-project.md/
new-project.md route an interrupted bootstrap (.planning/PROJECT.md +
config.json only) to Route F / STATE reconstruction / a hard error.

* fix(#4040): route partially-created .planning to initialization recovery

A bootstrap interrupted after .planning/PROJECT.md (but before
REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways:
progress.md read it as between-milestones (Route F) or 'no planning
structure', resume-project.md offered STATE.md reconstruction, and
new-project.md errored 'already initialized' — a routing loop with no
recovery exit.

Add a shared buildInitCompletenessFields discriminator
(planning_exists / requirements_exists / milestones_exists /
init_incomplete) to the init.progress, init.resume and init.new-project
payloads, and branch on init_incomplete in progress.md, resume-project.md
and new-project.md BEFORE the legacy branches. MILESTONES.md presence
excludes the archival between-milestones state, so Route F and the
STATE-reconstruction path keep working.

Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes.
Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch.
Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring.

* chore(#4040): add changeset fragment

* chore(#4040): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:03:42 -04:00
Tom Boucher
75ee7b0214 enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate

One query verb computes TDD-applicability for a plan (CLI flag, plan
type: tdd frontmatter, a task's tdd="true" attribute, or the
workflow.tdd_mode config default), mirroring phase.mvp-mode's
precedence-cascade shape. Foundation for epic #4272 Phase 2, which
wires both dispatch backends to consume it instead of restating the
predicate independently.

Also fixes workflow.tdd_mode, workflow.research, and
workflow.nyquist_validation, which never reached
cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone
because loadConfig() never populates config.workflow — a dead
accessor found while wiring this verb's own config read, fixed inline
per the no-defer rule rather than left alongside it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4273): document phase.tdd-applicable's FEATURES.md entry

Add a docs/features/ fragment for the new phase.tdd-applicable query
verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left
untouched: it documents /gsd-* slash commands only, and the sibling
verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI
reference entry anywhere in docs/ either -- only inline prose mentions
in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md
precedent to extend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests

Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): stop whitelisting capability-owned config keys centrally

workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are
each already owned by their own first-party capability's federated config
schema (the tdd/research/nyquist capabilities declare them under their own
capability.json `config`), resolved via isCapabilityConfigKey. Adding them
to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as
the prior commit in this branch did (mirroring workflow.mvp_mode, which
genuinely is central-only), declares the same key in two places at once.
That collision breaks capability-loader.cts's loadRegistry composition:
gsd-test caught this as 84-85 unrelated failures across
capability-cli/capability-command-dispatch/capability-lifecycle test files,
every one showing "unknown capability: <id>" for a freshly-installed
third-party capability that should have resolved fine.

Verified directly (not asserted): reverting only this file, keeping the
config-loader.cts tdd_mode/research/nyquist_validation flattening and the
init.cts call-site fixes from the prior commit, and re-running the exact
capability install + capability set repro from
tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the
failure with the whitelist entries present and clears it without them.
loadConfig() still surfaces all three flattened values correctly with no
central whitelist entry (confirmed directly against the compiled module) —
the whitelist additions were never required for the #4273 fix to work; they
were an incorrect over-application of the mvp_mode precedent to keys that
aren't central.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file

Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4273): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 14:14:56 -04:00
Adnan
f4bf449296 fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat

cmdAuditUat admits `human_needed` OR `gaps_found`, but
parseVerificationItems had a body only for the first and returned an
empty array for the second — standing on a comment deferring to
`plan-phase --gaps`, a different command audit-uat never reaches. Since
cmdAuditUat pushes a file into `results` only when `items.length > 0`, a
`gaps_found` report did not under-report: it vanished, taking its
phase's `by_phase` row with it, so a clean-looking total gave the reader
no cue anything was skipped.

Eligibility now has one owner (the caller) and parseVerificationItems
reports what the file says.

The closed-entry filter could not be built on extractFrontmatter: its
array-item parser keeps only each `- ` entry's FIRST line and has no
notion of nested key/value objects, so an entry's `status:`/
`resolution:` siblings never reach its output and a closed entry is
indistinguishable from an open one downstream. Rather than grow a
competing object-list parser — or change extractFrontmatter, whose blast
radius is every frontmatter consumer in the repo — this reads the raw
segment BEFORE the flattening, via the existing anchored
sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps`
machinery that already parses exactly this `- `-opened, indentation-
continued shape.

The human_needed path is byte-for-byte unchanged: same reader, same
display names, same numbering, no resolved-entry filtering — pinned by
a test and verified by identical CLI output on base and head.
parseGapsItems keeps its narrower `status: resolved` rule so no
*-UAT.md behaviour moves.

Closes #3850

* chore(#3850): backfill changeset pr number for #3879

* fix(#3850): one parse per entry, one fence parser, one resolved-entry rule

Adversarial review on #3879: B1, B2, M3, m5, m8 and n9.

B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence
regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell
5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found`
report vanished from the audit exactly as it did before this fix — this issue's
own symptom, on a platform the repo already has a named defect class for.
`extractFrontmatter`'s BOM+fence logic is now factored out as
`frontmatterRegion` and shared. One fence parser, not two.

B2 — the resolved-entry skip paired two DIFFERENT parsers by array index:
`parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A
block sequence written at its key's indent — ordinary, legal YAML — makes them
disagree about entry count, and from the first disagreement every index names a
different entry, so an OPEN entry inherits a CLOSED one's resolution and is
silently dropped. That is the defect this PR exists to fix, reintroduced inside
the fix. Display name and sibling fields now come from ONE parse of the raw
slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as
`parseYamlRegion` does, so the string is byte-identical to what
`extractFrontmatter` produced. The flattened array remains the #2286 GATE, but
is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the
LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment.

M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both
readers use, rather than two copies differing only in `result`.

m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an
acceptance criterion #3850 does not contain: the issue has no AC section, and
its suggested fix (2) states the skip unconditionally, naming a file with 14 of
16 entries resolved. That file is `human_needed`, so the asymmetry left the
reporter's own scenario over-reporting by 14.

m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and
says the column-0 boundary rule is now a cross-module contract.

n9 — the vestigial bare block is gone and its body de-indented.

Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF
fixture (M4 — it survived by accident, now pinned) and the unified skip rule.
Fail-first verified by running the new tests against the pre-fix build: the BOM,
nested-sequence and unified-skip cases are red there.

* fix(#3850): read the entries as objects, not as re-parsed display text

Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1
(#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml:
`parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry
now flattens to `test: A, resolution: R` rather than to its first line.

The original mechanism existed ONLY to work around that lossy first-line
flattening — it sliced the raw frontmatter segment and re-parsed each entry by
hand so a `resolution:` sibling was visible at all. With a real parser upstream
that workaround is obsolete, so it is deleted rather than repaired:
`sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the
`splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are
all gone.

`frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` —
the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence,
same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping
one step before the display flattening. `flattenObjectListItem` is exposed
alongside it so a caller deriving a display name produces the byte-identical
string `extractFrontmatter` would have.

That collapses the review's blockers into properties of the parse rather than
things this fix has to get right:

- B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI.
- B2 (index pairing) — there is no second reader. Display name and sibling
  fields come from one object.
- M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers.
- M4 (CRLF) — js-yaml's, not ours; verified through the CLI.

Also confirmed on the rebased base, per review: #3850 still reproduces on `next`
after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found`
fixture), so this PR is still doing work #3707 did not do. Nothing was dropped
as redundant.

One behaviour note: `entryField` returns a present value verbatim and treats
only whitespace-only as absent. Trimming would rewrite an author's `truth:` on
its way to becoming the display name.

* fix(#3850): keep every frontmatter list entry at its own row

Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result
to objects, and filtering COMPACTS: `parseHumanVerificationItems` then
numbered the survivors by their position in the compacted array. On a list
mixing object and non-object entries the non-object rows disappeared outright
and the rest were renumbered — #3850's own vanishing-row defect, reached
through entry SHAPE instead of file STATUS. Base never had it: it walked the
display array, so every row surfaced at its own position.

Renamed to `frontmatterListEntries` and it no longer filters (the name now
matches what it returns). Deciding what a non-object entry MEANS is a
caller's judgement; dropping it is nobody's.

Both readers now walk the DISPLAY array — one element per row, the array
#2286 already gates on — and consult the parsed array only for "does this
entry carry a closure field?". `parsedEntriesFor` owns that pairing and
checks the two lengths agree before trusting an index; all-null is the
correct degradation, since over-reporting a closed row is recoverable and
closing the wrong one is not. Names stay byte-identical to base for every
entry shape, including a nested sequence (`[nested]`, not `["nested"]`).

Same class closed in the gaps reader: a non-object `gaps:` entry surfaced
nothing at all and now surfaces as `unknown`, which is this module's
documented fail-safe direction (`parseGapsItems`) on a false-negative bug.

Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase
dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule
inlined twice; `extractFrontmatter` now routes through it, so "one fence
parser" is enforced rather than asserted in a comment.

Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and
its "shared by both readers" comment corrected — it has one call site, and
the two readers differ deliberately, each mirroring its own established
sibling (`parseGapsItems` vs #2286). Documented at the divergence.

Tests: `B2` asserted a name substring, so it passed while the row was
mis-numbered and would have passed through outright loss; it now asserts
positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim,
B2c the survivors' file positions across skipped rows, B2d the gaps reader.
All four fail-first against the reviewed head; 332/332 green with the fix.

* fix(#3850): make status authoritative, and let the two gaps readers agree

Round 4 review, all five findings.

Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as
closure regardless of `status:`, so `status: failed` + `resolution:
"attempted retry, still failing"` vanished from the report — the
silently-vanishing-item defect #3850 exists to close, reached by field
combination instead of file status.

Closure is now per key, because the two keys have different conventions
and one rule cannot serve both:

  `gaps:`               `status: resolved` only, byte-identical to the
                        rule `parseGapsItems` applies to a `## Gaps`
                        markdown section, so one authored entry cannot
                        read closed in one reader and open in the other.
  `human_verification:` a bare `resolution:` still closes, since that is
                        how verifier-written entries record it — but a
                        readable `status:` that contradicts it wins.

A single unified rule was the first draft and is wrong: it closes a
frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which
`parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring
claims it mirrors that reader's fail-safe status handling.

The contradiction guard is not a judgment call about YAML. It is the rule
this codebase already applies to the same field pair: `validateResolution`
(probe-core.cts) rejects a populated `resolution:` on a non-resolved status
outright — "a populated payload is an authoring mistake ... Reject it so
the mistake surfaces." A reporter cannot throw, so it surfaces the item.

Minor 1. Direct unit tests for `frontmatterListEntries` and
`flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that
historically co-changes with `frontmatter.cts`. They were reachable only
through `uat.cts`' readers before.

Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted
directly. Verified unreachable through content rather than assumed: both
readers enter through `frontmatterRegion`, `extractFrontmatter`'s only
extra argument gates a warning, and `normalizeParsedValue`'s `value.map`
is 1:1. It is a drift alarm for a future edit to either parser, so the
helper is exported for tests rather than left as the one unpinned branch.

Minor 3. The vestigial `const skipResolved = true` and its dead
conditional are gone.

Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:`
entry has no `test:` in its vocabulary — the template's entries carry
truth/status/reason/artifacts/missing — so it was speculative support for
a field the shape does not have, and it collided with the 1..N row numbers
`parseHumanVerificationItems` assigns by array position. Not reading it
makes the collision impossible; an offset would have rewritten an authored
value, against `entryField`'s verbatim contract.

Docs, changeset and the dispatcher docstring all stated the unconditional
rule and are corrected — three prior rounds here were comment/code drift.

Fail-first proven: restoring the universal rule reddens all three new unit
tests and both rewritten properties.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-04 17:47:01 +00:00
Rezolv
590edec7a7 fix(#3956): require positive evidence for verify artifacts/key-links pass (#4004)
* fix(#3956): require positive evidence for verify artifacts/key-links pass

An all-string or path-less must_haves.artifacts / key_links block is
item-by-item skipped, leaving zero checked results, yet the pass verdict
was computed as `passed === results.length` (0 === 0), so all_passed /
all_verified read true with status valid and exit 0: a silent false GREEN
over zero acceptance evidence.

Add a positive-evidence floor (results.length > 0) to both verdicts,
mirroring the no-vacuous-pass rule at src/uat-predicate.cts. A well-formed
block, the fully-empty-block error, the parser's string tolerance, and
key-links pending (#1202) semantics are all unchanged.

Governing: ADR-3473 section 8 / 37C (absence, emptiness and failure must
not encode as success) and Decision 3 (failure is a value).

* chore(#3956): add changeset for verify vacuous-pass fix

* test(#3956): add mixed-block coverage and correct the key-links vacuous-pass comment

Addresses review on #4004:
- Correct the cmdVerifyKeyLinks positive-evidence-floor comment: only bare-string
  items are continue-skipped; a from:-less object is NOT skipped (it falls through
  to a verified:false hard failure), so it was never part of the vacuous-pass
  surface. The prior comment overclaimed symmetry with the artifacts side.
- Add a mixed-block regression test per verb (one bare-string prose bullet + one
  well-formed entry): the string is skipped, results.length === 1 > 0, and the
  verdict follows the single real entry — pinning that the floor does not
  over-reject a partial block.
- Tighten the changeset wording to match (all-bare-string, not "no path:/from: key").

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-03 21:46:50 +00:00
Rezolv
a788afb120 fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives (#4021)
* fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives

stateReplaceField's bold and plain patterns used `\s*` for the label-to-value
gap, which matches the newline after an empty field; `(.*)` then captured the
following line and the rebuild discarded it -- silent STATE.md data loss on any
`state update` against an empty body field (Status:, Stopped at:, Paused at:),
with exit 0 and no warning.

Confine the gap to same-line whitespace (`[ \t]*`), mirroring the already-correct
read side (stateExtractField, src/state-document.cts:404/:409), and pin the
label-value separator to a single space when the label line had none, so an empty
field yields `**Status:** value` rather than a glued `**Status:**value`. Non-empty
and pipe-table replacements are byte-identical to prior behaviour.

ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; this aligns
the writer to it. Regression test fails before / passes after and covers bold and
plain shapes, LF and CRLF, the non-empty byte-identity guard, and an end-to-end
transitionCore characterization at the consumer (ADR-3180 Decision 4(c)).

* chore(#4010): add changeset for the stateReplaceField empty-field fix

* test(#4010): add boundary and property coverage; scope the changeset's unchanged claim

Addresses review on #4021:
- Add boundary tests for the shapes the example tests missed: an empty field at
  end-of-document (no following line, bold + plain), two consecutive empty fields
  (only the target is filled, the other empty field's line survives), and an empty
  new value on an empty field (joinFieldReplacement synthesizes no dangling
  separator and the following line is preserved).
- Add a fast-check property over the bold/plain branches and joinFieldReplacement:
  for any field name, any values (empty fields included), and any new value,
  replacing one field changes only its own line and never the total line count —
  the invariant #4010 violated, now guarded directly.
- Scope the changeset's "unchanged" claim to ordinary space/tab separators (an
  exotic vertical-tab/form-feed separator, which no GSD template emits, now
  normalises to a single space).

* test(#4010): pin glued-separator non-empty field, scope joinFieldReplacement JSDoc

Round-3 review carried forward a Minor finding: joinFieldReplacement's JSDoc
still claimed non-empty replacements are unconditionally "byte-identical to
prior behaviour", but a non-empty field written with no label-to-value
separator (**Status:**value) gains a single inserted space under the narrowed
[ \t]* gap. Round 2 scoped only the changeset prose; the source JSDoc was left
making the false unconditional claim.

- Scope the JSDoc's byte-identity claim to ordinary space/tab separators and
  name the no-separator normalization as the one intentional exception.
- Add a test pinning the glued-separator case (**Status:**Planning): exactly
  one space inserted, following line survives, not byte-identical.

Emitted .cjs is gitignored (class-1), so no emitted-drift-ack applies.
build:lib clean; 74/74 state-document tests pass.

Claude-Session: https://claude.ai/code/session_01Mzmut6aeqZ1APfUBAkBZTR

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-03 17:14:12 -04:00
Tom Boucher
515191f07d feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only)

* test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED)

Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/
merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and
confirms the prior research pass's Open Question 1: a coordinator crash
between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge)
leaves BATCH.json at "pending" with no STATE.md row yet (only written in
Step 9), so --resume's eligibility re-derivation would dispatch a second
executor into a new worktree for the same item, orphaning the first.

This test asserts worktree-dispatch.md's Step 6 excludes an item whose
SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's
existing PLAN.md-existence check one layer earlier. Fails against the
current worktree-dispatch.md, which has no such guard.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the full trace and fix-location rationale.

* fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN)

worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round
via the same quick-batch resume call resume-mode.md uses, but had no check
for "did this item already finish executing" the way planner-wave.md
already checks "did this item already get planned" (PLAN.md existence)
before re-planning. A coordinator crash between Step 6 (executor commits,
SUMMARY.md written) and Step 7 (merge) left the item eligible for a second
dispatch on --resume, orphaning the first worktree's real, already-
committed work and silently losing it once the second executor's SUMMARY.md
write clobbered the first at the same item_dir path.

Adds a SUMMARY.md-existence exclusion before spawn-plan is computed,
symmetric to planner-wave.md's PLAN.md check. The excluded item is not
lost: merge-wave.md's own mergeable-wave criterion (status=pending,
SUMMARY.md on disk, not yet merged) already picks it up independently of
this eligible/spawn list.

Workflow-prose-only fix — touches no already-merged/reviewed .cts module.
See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the fix-location rationale (why not resumeBatch itself).

* test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules

Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677,
epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership
attempts", "scope drift", "submodules"):

- Arbitrary-worktree ownership tampering: a manifest entry naming a
  non-agent branch is silently dropped at normalization before any git
  subprocess runs; a manifest entry naming a plausible agent-branch that
  was never actually created by this repo's own worktree.create (a
  genuinely foreign repo/branch) is blocked via base_mismatch. Both leave
  the foreign location and repoRoot's HEAD provably untouched.

- Advisory scope drift: a committed path outside declared files_modified
  still merges successfully (advisory, never blocking) while surfacing a
  scope_out_of_declared warning naming the drifted path; an exact
  declared-scope match produces zero warnings (boundary case).

- Real .gitmodules submodule integration: a repo containing a real local
  git submodule merges cleanly through executeWorktreeWaveCleanupPlan for
  an unrelated plan; a real gitlink pointer bump (declared) merges cleanly
  with the superproject tree reflecting the new pinned commit; an
  undeclared bump is advisory-only and surfaces a scope warning naming
  vendor/sub, same as any other undeclared modification.

No src/*.cts changes — all three gaps were coverage-only; the underlying
primitives already behaved correctly (independently verified against real
git subprocess output before writing each assertion).

* docs(#3677): document how to diagnose a preserved quick-batch worktree

Extends the one-sentence "worktree is preserved (never deleted)" mention
into a concrete diagnosis procedure: where the preserved directory is, how
to read the executor's real commits/diff against the plan's declared
files_modified, how to read the item's own SUMMARY.md independent of merge
outcome, how to manually merge-and-clean-up or discard, and how to re-run
--resume afterward. Also documents that a SUMMARY.md-written-but-still-
pending item (the crash-window case fixed in this same PR) needs no manual
intervention — --resume routes it straight to the merge step.

* chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only)

* fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable

Orthogonal review (Spec finding): the crash-window regression test added
earlier this phase only asserted readStep('worktree-dispatch.md') + regex
matches against the markdown prose — proving the DOCUMENTATION says the
right thing, never that the runtime condition (pending status + on-disk
SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own
"Alternatives considered" explicitly rejects "document recovery without
fault injection" for exactly this reason.

Extracts the filtering decision into a pure, independently testable
function, filterAlreadyExecuted(eligibleIds, executedIds) in
src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed`
CLI verb (src/quick-batch-command-router.cts) — the same pure-decision-
then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already
establish. worktree-dispatch.md now calls this verb explicitly instead of
only describing the decision in prose. A genuine fixture-based test in
tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch),
writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the
REAL resumeBatch, and proves both that resumeBatch alone still reports the
item eligible AND that filterAlreadyExecuted (fed a real filesystem check)
correctly excludes it. The prior prose-assertion tests are kept — they now
prove the workflow markdown is correctly WIRED to the verb — but are no
longer the only proof.

Self-discovered defect while building that fixture (fixed inline, not
deferred): tracing merge-wave.md against /gsd:quick's own prior art
(QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed
$QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed
coordinator correctly does not re-dispatch an already-executed item (this
fix), but nothing durably recorded that item's worktree_path/branch/base
either — Step 7 in the resumed process would have had no data to build its
cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/
dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT
a reuse of the pre-existing `worktree` field, whose loadBatch validation
requires the path to exist on disk (verified empirically: reusing it made
the batch permanently unloadable the moment a legitimately-merged worktree
was removed). worktree-dispatch.md persists the triple once a worktree is
created; merge-wave.md falls back to it when the ephemeral manifest lacks
an entry, clears it after a successful merge, and fails closed rather than
guessing if no record exists anywhere.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1
and §9.3 for the full trace, empirical verification notes, and rejected
alternatives (reusing `worktree` directly).

* test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees

Orthogonal review (Security finding): the two existing ownership-tampering
tests didn't test ownership — one was trivially rejected by
WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch-
NAME filtering, not ownership), the other pointed at a wholly separate,
never-linked foreign repo, so merge-base failed immediately because the
branch didn't exist as a ref at all. Neither exercised the real scenario:
a manifest entry whose worktree_path/branch are swapped to point at a
DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot,
with a branch name passing the shape check and a base in allowed_bases.

Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts)
directly: this is NOT a reachable gap. Git enforces branch-per-worktree
uniqueness, so a swapped-in entry.branch can only match worktree_path's
ACTUAL checked-out branch if it names that sibling's own real, uniquely-
generated branch name — which manifest tampering confined to one batch's
own record has no way to know (branch names are
agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is
collision-checked GLOBALLY across every existing quick task and batch, not
merely within one batch).

Adds a stronger test that empirically proves this: two REAL, concurrently-
alive sibling worktrees of the same repo (both via real `git worktree add`,
both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with
worktree_path/branch swapped between them in both directions. Both attempts
are blocked via branch_mismatch; both real worktrees, their branches, and
one sibling's real uncommitted-to-main commit survive completely untouched.
Supplements (does not replace) the original two tests, which still prove
distinct, real boundaries.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2
for the full trace, including the one explicitly-documented (not fixed)
trust boundary this investigation surfaced: the primitive defends against
fabricated data, not a caller bug that misattributes a real-but-wrong
item's own triple to a different item.

* chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only)

* docs(#3677): add changeset for PR 4240

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 09:47:22 -04:00
Tom Boucher
d5f8191f66 fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical

* fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick

* fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter

Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run)
Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction

* chore(#3730): backfill changeset pr number

* fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch

---------

Co-authored-by: sim <sim@local>
2026-09-03 01:32:41 -04:00