* fix(#2176): ground the Antigravity reviewer in the repo under review
- capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass
the repo root on both invocation arms
- anchor _AGY_PROMPT to the absolute repo root; mandate a
REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable
- stamp a [reviewed-without-repo-access] marker on self-reported or
scratch-anchored output; Consensus Summary down-weights marked reviews
- apply the same absolute-root anchor to the cursor-agent prompt (AC5)
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* docs(#2176): changeset fragment for PR #2184
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells
- regenerate tests/workflow-size-baseline.json for review.md's growth
- cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the
wrong root from a repo subdirectory)
- blind-review tells anchored: self-report to the first lines of output,
scratch tell to a workspace-declaration phrasing — a grounded review
quoting either string is no longer mis-stamped
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* test: regenerate golden-install-parity fixtures for the review.md change
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* test(#2176): pass the transcript path to bash with forward slashes
The behavioral detection test substitutes a mkdtemp path into the bash
compound; on Windows runners that path contains backslashes, which bash
strips, so the transcript is never found and the first assertion fails
(windows-latest/24 lane). Git Bash accepts D:/-style paths.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* fix(#2176): use /gsd:review namespace syntax in workflow comment
The slash-command namespace invariant (#3443) bans retired /gsd-<cmd>
references in Claude-facing sources; a cursor-anchor comment used
/gsd-review. Size baseline + golden fixtures regenerated for the byte
change.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* test(#2176): derive the POSIX path via path.sep, not a hardcoded separator
Review finding: out.replaceAll('\\', '/') hardcodes both separators;
use the separator-safe out.split(path.sep).join(path.posix.sep) idiom.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* test(#2176): use the merged toPosixPath seam for the bash path
Per maintainer note: #2247's shell-command-projection now centralizes
running-OS → POSIX path conversion; import it instead of the inline
split/join idiom.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:
- toPosixPath(p) — this machine's native path → POSIX (running-OS relative;
for local filesystem paths).
- toNativePath(p) — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
to a POSIX/bash TARGET (which may differ from the running
OS) and for parsing mixed-separator input.
core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.
- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
runtime-artifact-install-plan, drift, init, worktree-safety,
installer-migrations, installer-migration-authoring, install-engine, surface,
verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.
Closes#2246
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The windows-robustness guard (bug #685) requires every external-binary
spawn to set windowsHide:true so no console window flashes on Windows.
Golden fixtures regenerated for the hook byte change.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling
The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to
auto-approve a gate="blocking-human" checkpoint and escalates it so a human can
vet the package. execute-phase's checkpoint_handling step then dispatched purely
on checkpoint *type* and never read gate -- so under --auto/--chain it
auto-approved the checkpoint the executor had just refused to auto-approve.
Net effect: the slopsquatting defence was inert in exactly the unattended mode
where it matters. An [ASSUMED]/[SUS] package reached install with no human ever
seeing the prompt.
- gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the
package-legitimacy what-built markers) ahead of every auto-mode branch.
- gsd-core/references/checkpoints.md: document the gate attribute and its two
values. blocking-human previously appeared nowhere outside gsd-executor.md,
so no planner had a documented way to author a non-auto-approvable checkpoint.
- tests/package-legitimacy-gate.test.cjs: the existing regression test asserted
the executor half only, which is why it stayed green while the gate was open.
Now asserts the orchestrator half too.
* chore(changeset): link to issue #2107
* chore(changeset): backfill PR number 2113
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa
* test(#2107): refresh golden-install-parity hashes for edited gsd-core files
The golden fixtures pin content hashes for gsd-core/references/checkpoints.md
and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated
via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa
* fix(#2107): keep the carve-out inside the ADR-857 host-loop budget
The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so
optional-feature logic keeps migrating out of the host loop. The carve-out
first landed 623 bytes over that ceiling.
Move the two-layer rationale (why gsd-executor escalates these checkpoints)
into references/checkpoints.md, where the gate is now documented, and reduce
the workflow to the operative rule. execute-phase.md is 93589 bytes, under
the ceiling; the gate token and both <what-built> marker strings are kept
because the orchestrator matches on them.
Refresh the two baselines the edit invalidates: golden-install-parity
fixtures (only the checkpoints.md and execute-phase.md hashes move) and
workflow-size-baseline.json (one line). The ADR-857 ceiling itself is
untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa
* fix(#2107): executor honors blocking-human on the decision branch + gate transport
Review found the fix incomplete one layer down. Two executor-layer gaps:
1. Blocker — agents/gsd-executor.md auto-mode dispatch gated
checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision
branch below auto-selected the first option with no gate check. The executor
resolves a decision itself (auto-selects and continues) without returning it,
so the orchestrator carve-out never runs for it. A planner following the new
checkpoints.md rule 6 ("gate a decision whose default would be wrong to
assume") would have it silently auto-selected under --auto/--chain — the exact
#2107 harm, one checkpoint type over. The decision branch now STOPs and
returns for an explicit human decision when gate="blocking-human".
2. Major (transport) — checkpoint_return_format carried no field conveying the
gate to the freshly-spawned orchestrator, so recognition of the proactive
pre-install checkpoint rested on freeform prose. Added a **Gate:** field to
the return format and re-pointed the execute-phase carve-out at it
("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md
drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests
- New: 'auto mode does not auto-select a blocking-human decision checkpoint'
asserts the executor decision branch STOPs on blocking-human. Verified red on
the pre-fix executor (2 fail), green with the fix (27 pass).
- New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:**
field carries blocking-human across the executor->orchestrator boundary.
- New: 'auto-select rule for decision is conditional' — orchestrator-side mirror
of the human-verify conditional test, for the execute-phase decision branch.
- Fix vacuous test: both conditional tests now assert the anchor matched
(length > 0) before iterating, so anchor drift can no longer pass with zero
assertions.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2107): refresh golden + size baselines for executor + execute-phase edits
Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the
gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the
runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is
unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973,
execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
- readGitStatus calls execFileSync via the child_process namespace so tests
can inject spawn failures through the shared module object
- two deterministic tests: ERR_CHILD_PROCESS_STDOUT_MAXBUFFER-shaped and
ETIMEDOUT-shaped throws both degrade to null (segment absent), proving
the fail-soft paths the PR previously only asserted
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
- changeset reformatted to the bold-lead + (#2163) house convention
- explicit 8 MiB maxBuffer on the git spawn (default 1 MiB could overflow on
huge dirty repos; overflow still degrades to segment-absent)
- child_process require hoisted to module level per file style
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
New statusline.show_git config (default false). When enabled, a git
segment renders after the directory: current branch plus compact
work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓
when clean and in sync), e.g. " │ main+2~1?3".
One git status --porcelain=v2 --branch spawn per render via execFileSync
with a fixed argument array (no shell), a 1.5s timeout, and the
workspace dir passed with -C. Fails silently — segment absent outside a
repo, without git, or on timeout. Default output is unchanged when the
flag is absent.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* feat(#2161): opt-in absolute token count on the statusline context meter
New statusline.show_context_tokens config (default false). When enabled,
the context meter shows the absolute token total after the percentage,
e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read,
and output tokens from context_window.current_usage (matching /context).
Default output is byte-for-byte unchanged when the flag is absent or
false. The .planning config is now read once per render and shared with
the last-command/position block instead of being re-read.
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* docs(#2161): changeset fragment for PR #2174
* fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format
- formatTokens promotes to the M branch when k-rounding reaches 1000
(999,500-999,999 rendered "1000k" instead of "1.0M")
- boundary tests at 999499/999500/999999/1000000/1000001
- Number() guards on the four usage fields (silent string-concat gap)
- changeset body ends with the (#2161) citation per house convention
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style
- config-set accept/reject tests for statusline.show_context_tokens
(mirrors the post-planning-gaps precedent the issue scope names)
- changeset + docs no longer claim parity with /context: the suffix sums
four fields while the meter %% derives from used_percentage (three), so
the figures can diverge slightly
- module.exports one entry per line
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
* test: regenerate golden-install-parity fixtures for the statusline hook change
Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#2206): strip trailing slashes in isGitIgnored to avoid CRLF check-ignore quirk
isGitIgnored was called with a trailing slash (`.planning/`) in config-loader.
git check-ignore has a longstanding quirk: a CRLF .gitignore with blank lines
falsely reports any path WITH a trailing slash as ignored. This silently set
commit_docs=false on Windows repos (where CRLF .gitignore is the norm), skipping
all planning-doc commits. Normalize trailing slashes inside isGitIgnored so every
call site is protected.
Closes#2206
* docs(#2206): add changeset fragment
* docs(#2206): backfill PR number
* fix(#2203): traceability parser matches REQ-IDs in any column, not just the first
The traceability table-row parser required the REQ-ID in the first column
(`^| REQ-ID |`). A table that leads with a status column (e.g. `| ☐ | REQ-01 |`) matched zero rows, so phase complete warned every body REQ-ID was missing.
Match REQ-IDs in any pipe-delimited cell (drop the ^ anchor).
Closes#2203
* docs(#2203): add changeset fragment
* docs(#2203): backfill PR number
* fix(#2202): preserve unknown frontmatter keys in syncStateFrontmatter
syncStateFrontmatter rebuilds frontmatter from a fixed schema, dropping any
custom/unknown key on every mutating verb. Before reconstruction, merge any
existing frontmatter key the schema does not own. Schema keys still win.
Closes#2202
* docs(#2202): add changeset fragment
* docs(#2202): backfill PR number
* fix(#2202): add regression test + remove redundant type assertion
- tests/state.test.cjs: behavioral regression test asserting custom/unknown
STATE.md frontmatter keys survive a mutating verb (they were silently dropped
before the syncStateFrontmatter carry-forward).
- src/state.cts: drop the unnecessary `as Record<string, unknown>` assertion
that tripped @typescript-eslint/no-unnecessary-type-assertion (the lint-tests
gate failure).
Refs #2202
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The <execution_context> reference presented @~/.claude/gsd-core/... as canonical
and described the paths as files "the executor reads before starting". Both are
misleading: the prefix is install- and runtime-relative (Claude global vs Cursor
.cursor/gsd-core vs an absolute --local path), so a committed plan is not
clone-portable, and /gsd-execute-phase loads the workflow from its own installed
copy rather than gating on the committed block.
- docs/reference/plan-md.md: describe the install-relative, non-clone-portable
nature of the block and contrast it with repository-relative <context>.
- .out-of-scope/plan-md-execution-context-portability.md: record the #2238
wontfix decision (PLAN.md is a machine artifact; #2158 precedent) with a
revisit-if condition.
Refs #2238. Closes#2240.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(2119): single SECURITY.md writer — auditor is return-only
The gsd-security-auditor held Write/Edit and was instructed to write
SECURITY.md (no <N>- prefix, no template frontmatter), while the
orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md
from templates/SECURITY.md. Two writers, two naming conventions, two
shapes — the auditor's unprefixed file was invisible to the workflow's
*-SECURITY.md glob detector and unparseable for the threats_open gate.
Fix (option 1 from the issue): make the auditor return-only.
- Remove Write/Edit from auditor's tools
- Rewrite all 'Write SECURITY.md' instructions to 'Return structured
verdict' with threats_open count
- Add explicit constraint in workflow Step 5 spawn prompt
- Update existing test (was asserting Write in tools — now asserts absence)
- Add new regression test for single-writer contract
- Update docs/AGENTS.md stale Tools/Produces rows
- Regenerate golden fixtures + agent size baseline
* docs(changeset): backfill PR number (#2154)
* chore(#2119): regenerate pi/qwen golden fixtures after next merge
The single-writer change edits gsd-core/workflows/secure-phase.md and
agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge
straggler) were the only runtime fixtures still holding pre-change hashes for
those files. All other runtimes already reflect the change. Regenerated via
the sanctioned gen-golden-install-parity script.
* merge origin/next — regenerate goldens + baseline for merged state
* fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix)
* fix(#2200): scope phase-complete roadmap writes to the current milestone
runPhaseCompleteTransaction's roadmap mutators ran unanchored / un-milestone-
scoped / first-match over the whole ROADMAP: the phase-checkbox flip could check
a bullet inside a backticked prose literal or an earlier Backlog entry instead of
the closing phase's bullet (a transposition), and the Plans-count writer could
bind to a same-numbered phase in a shipped milestone.
- Add currentMilestoneRawRanges(content, cwd) to roadmap-parser.cts: raw
[start,end) offsets of the active milestone's region(s) (primary section + the
optional Phase Details section), mirroring extractCurrentMilestone's selection.
- Line-anchor the phase checkbox pattern (^ + m flag) so an inline / backticked
prose literal cannot match.
- Apply the phase-checkbox flip, the Plans-count write, and the per-plan checkbox
flips ONLY within the current milestone window(s) via a mutateMilestonePhase
helper (splice later windows first so offsets stay stable). Fall back to whole-
content mutation when there is no versioned active milestone (prior behaviour).
The Progress-table writer stays as-is (already scoped to ## Progress, #2012).
Closes#2200
* docs(#2200): add changeset fragment
* docs(#2200): backfill PR number in changeset fragment
* fix(2112): scope commit to --files pathspec, not entire index
cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files
named in --files but then ran a bare 'git commit' with no pathspec,
absorbing anything else in the index into a commit whose message
described only the named files (#2112).
Fix: append '-- ...stagedPaths' to the commit args when the caller
declared a scope. Three guards are load-bearing:
- stagedPaths (not filesToStage) excludes skipped missing files (#2014)
- explicitFiles gate keeps the default .planning/ path byte-identical
- MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge
- --amend is left without pathspec (different operation)
cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the
full rename is captured atomically.
Also fixes workflow markdown in spec-phase.md and add-tests.md.
All-files-missing now short-circuits to nothing_to_commit instead of
absorbing the entire index under a message describing files that
were not committed.
* docs(changeset): backfill PR number (#2148)
* test: update golden-install-parity fixtures for workflow markdown changes (#2112)
* test: update golden fixtures + workflow baselines for #2112 changes
- claude-local.json golden fixture (now generated via gen script)
- workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes)
- Extended gen-golden-install-parity-zcode.cjs to also regenerate the
claude local-layout fixture
* fix(2118): honor --dry-run in milestone complete with zero-mutation preview
milestone complete treated --dry-run as a no-op: the flag was neither
parsed nor rejected, so a caller who expected a preview instead
triggered the full destructive mutation (archive phases → move audit
artifacts → rewrite STATE.md) with no way to back out.
Fix (option 2 from the issue): add dryRun to MilestoneCompleteOptions,
parse --dry-run in the dispatcher, and return a JSON preview plan
(would_archive, would_update) after the read-only stats gathering but
before any mutations. Also gated platformEnsureDir on !dryRun so the
archive directory is not created during preview.
3 regression tests: no-mutation happy path, --no-archive-phases combo,
and --force bypass combo.
* docs(changeset): backfill PR number (#2155)
* chore(#2118): regenerate pi/qwen golden fixtures after next merge
The milestone --dry-run fix changes gsd-core/bin/gsd-tools.cjs; qwen.json
(missed at authoring) and pi.json (added on next, never carried the fix)
were the only two runtime fixtures still holding the pre-fix hash. All
other runtimes already reflect the change. Regenerated via the sanctioned
gen-golden-install-parity script.
* fix(#2118): surface accomplishments in dry-run preview; fix --dry-run --raw
Orthogonal review findings on the --dry-run preview:
- The preview omitted the already-computed accomplishments (the primary
MILESTONES.md content a real run writes); surface it as a top-level field,
mirroring the real-run result.
- `--dry-run --raw` discarded the structured payload and printed the literal
string "dry-run"; drop the raw-value arg so --raw emits the full preview
JSON, matching the real-run output() call.
Adds tests: --dry-run --raw is parseable JSON, preview includes accomplishments,
and --dry-run --force is proven zero-mutation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2201): accept --phase N flag in the phase verb family (complete, list-plans)
The phase family router treated the first positional (args[2]) as the phase
number, so `phase complete --phase 12` passed the literal '--phase' as the phase
→ 'Phase --phase not found'. The state family already accepted --phase N. Now
complete and list-plans accept --phase N (and --phase=N) as well as the bare
positional; unrecognized flags yield a usage error naming the accepted form.
Closes#2201
* docs(#2201): add changeset fragment
* docs(#2201): backfill PR number
* test(#2199): cover bullet/em-dash ROADMAP phase resolution + milestone count
Adds the bullet-only ROADMAP fixture the suite lacked: an all-bullet em-dash
ROADMAP resolves each phase (no Phase null), colon/en-dash/hyphen bullet
separators all resolve, mixed heading + bullet forms coexist, and the milestone
phase-count counts bullet-form phases instead of collapsing to zero.
* fix(#2199): accept bullet/em-dash phase entries in roadmap lookup + milestone filter
Roadmap phase lookup (findRoadmapPhaseInContent) matched only ATX headings
against a colon-required pattern, so a bullet/checkbox entry like
`- [ ] **Phase N — name**` — which the bundled roadmapper emits in bullet-house-
style ROADMAPs — resolved found:false and `Phase null` was written into STATE.md.
The milestone phase-filter built its phase set from headings only, so a bullet-
only ROADMAP collapsed to a zero-count pass-all filter and progress denominators
broke.
- Add a shared bullet-phase-line pattern (separator: em-dash/en-dash/hyphen/colon).
- findRoadmapPhaseInContent: on a heading-match miss, fall back to the bullet line
for the requested phase; return found:true with the captured name.
- getMilestonePhaseFilter: also scan bullet lines into the milestone phase set.
Closes#2199
* docs(#2199): add changeset fragment
* fix(#2199): bullet phase lookup as last resort + consolidate test (review)
Two corrections to the initial fix:
1. Regression — the bullet fallback inside findRoadmapPhaseInContent was too
eager: it returned a bullet match from the scoped (current-milestone) content
before the caller tried the full-content heading path, so a phase whose
Requirements live in a Phase Details heading (after the active-milestone
section) got a bullet-line section with no Requirements → phase_req_ids null
(broke 3 init tests). Restructure: findRoadmapPhaseInContent is heading-only
again; a separate findRoadmapBulletPhaseInContent runs in getRoadmapPhaseInternal
ONLY after scoped + full heading lookup fails, so a heading with a Requirements
section always wins.
2. lint-test-file-count — the standalone fix-2199-roadmap-bullet-phase.test.cjs
collided with the over-cap 'roadmap' module (FAIL_NOVEL_FILES). Consolidate
the regression into the existing tests/roadmap-parser.test.cjs (its natural
home, under the 2-file cap).
* test(#2199): assert heading-in-full beats bullet-in-scoped (review L3)
The exact first-attempt regression: a phase has a bullet in the active-milestone
scope but its heading (carrying Requirements) lives in a Phase Details section
outside that scope. Pin that the heading section wins over the bullet line so
req_ids resolve and the eager-bullet bug cannot return.
* docs(#2199): backfill PR number in changeset fragment
The Gemini, Claude, and Codex reviewer blocks invoked the CLIs with no explicit
timeout, so each inherited the host default (~2 min on Claude Code). A source-
grounded review of a large plan set takes ~570s (Codex xhigh) / ~525s (headless
Claude) — both exceed that window, so the lane is killed mid-review, its output
is empty, and the cross-AI review silently proceeds with fewer lanes. CodeRabbit
and OpenCode already documented a timeout; the four main lanes did not.
Add a shared timeout-guidance note directing a high Bash timeout (>= 900000;
1200000 for Codex xhigh / headless Claude), referencing BASH_MAX_TIMEOUT_MS for
the Claude Code host cap, and framing a slow-lane empty output as a timeout kill
(not the 0xc0000142 crash it gets misdiagnosed as) so operators re-run with more
time instead of diagnosing a CLI failure.
Closes#2194
Recaptures the 18 golden-install-parity fixtures + workflow-size baseline (only
the review.md entry changed in each; review is LARGE-tier, 47168 < 61440).
Source-text contract guard: review.md IS the product the runtime loads. Assert
its reviewer-invocation section carries Bash timeout guidance (>= 900000ms) for
the Gemini/Claude/Codex lanes and frames a slow-lane empty output as a timeout
rather than a crash.
CI lint (local/no-crlf-fragile-split + no-regex-spaces) flagged the regex
assertions: a bare \n on readFileSync content is CRLF-fragile on Windows
git-autocrlf, and literal double-spaces are hard to count. Switch to .includes()
strings — same validation intent, lint-clean on all platforms.
cmdStateUpdateProgress ran its Progress: patterns against the raw STATE.md
content (frontmatter included). With /i+/m the first match was the YAML
frontmatter `progress:` key, not the body line — so the frontmatter block was
mangled (\s* crossed the newline) while the body line stayed stale, and the
next STATE.md write re-derived percent from that stale line, silently reverting
the update. \2: the .* also discarded any agent-authored descriptive suffix.
- Match against the frontmatter-stripped body only (reusing stripFrontmatter);
reconstruct content as fmPrefix + modified body so the frontmatter is untouched.
- Swap only the machine segment ([bar] NN% or bare NN%), preserving any suffix.
- updated stays false when the body has no Progress: line (no false success from
a frontmatter progress: key).
Closes#2177
Adds the frontmatter-bearing STATE.md fixture the suite lacked. Asserts the
body Progress: line (not the YAML progress: key) is the update target, the
descriptive suffix after [bar] NN% is preserved, the frontmatter block is not
mangled, and a body with no Progress: line reports updated:false even when the
frontmatter has a progress: key (no false success).
normalizeNodePath had hardcoded branches for macOS Intel (/usr/local/Cellar) and
Apple Silicon (/opt/homebrew/Cellar) but none for Linuxbrew
(/home/linuxbrew/.linuxbrew/Cellar). After `brew upgrade node` on Linux, the
version-pinned Cellar path baked into managed hook commands 404'd, and re-running
the installer did not repair it — normalizeNodePath returned the pinned path
unchanged, so the 'already normalized' skip-rewrite check no-op'd.
Generalize to one branch: match <prefix>/Cellar/node(<@ver>)?/<ver>/bin/node and
rewrite to <prefix>/bin/node, deriving <prefix> from the path itself. This covers
macOS Intel, Apple Silicon, Linuxbrew, and any custom HOMEBREW_PREFIX — the same
failure class as #977 (fnm) and #1619 (mise), now closed for Homebrew on every
platform.
Closes#2185
Adds the Linuxbrew Cellar layout (/home/linuxbrew/.linuxbrew/Cellar/...), a
post-version-bump path, a node@version formula, and a custom HOMEBREW_PREFIX —
all mapping to the stable <prefix>/bin/node symlink. Mirrored in both
consolidated describe blocks.
readGsdEffectiveModelOverrides resolved ~/.gsd/defaults.json via os.homedir()
with no seam, so a test asserting project-only overrides could not isolate the
global file. Add an optional { homedir } option (defaults to os.homedir()) to
readGsdGlobalModelOverrides and readGsdEffectiveModelOverrides — the same
dependency-injection shape the sibling warnIfStaleBake already uses. Backward-
compatible: existing callers pass no option and behave identically.
Closes#2152
The subtest asserted project-only model_overrides but called
readGsdEffectiveModelOverrides without isolating HOME, so the real
~/.gsd/defaults.json global overrides bled into the deepEqual on any dev box or
non-hermetic runner that has one. Pass a sandboxed homedir (mirroring the sibling
warnIfStaleBake subtests that already inject homedir).
Review found the hint value (yes|no) matched prefixes ('nope'/'not' read as
'no') and the hint regex was unanchored, so a mid-line prose mention like 'see
**UI hint**: no above' was treated as the authoritative metadata line. Line-
anchor (^ + m flag) so only a real hint line counts; word-boundary on the value
so nope/not do not mean no. Adds coverage for hint:yes over a pure-backend body
and the nope fall-through.
checkUiPresence ran UI_TOKENS (which includes the bare token 'UI') over the raw
phase-section text, so GSD's own '**UI hint**: no' metadata line matched the 'UI'
token and reported hasUI=true — blocking non-frontend phases that explicitly
declare themselves non-frontend via the documented convention (the plan-phase
UI-SPEC gate fired under the default workflow.ui_safety_gate=true).
- An explicit '**UI hint**: yes|no' line is now authoritative (mirrors how
progress.md / new-project.md already parse it via 'UI hint.*yes'). hint:no ->
hasUI=false; hint:yes -> hasUI=true.
- Any '**UI hint**:' line is stripped before token-sniffing, so a hint without a
recognised yes/no cannot false-positive on the bare 'UI' token.
- With no hint line, behaviour is unchanged (token-sniffing on the rest).
Closes#2150
Adds the divergent input the green suite lacked: the documented `**UI hint**:
no` metadata line. Asserts hint:no is authoritative non-frontend (no false
positive), hint:yes is authoritative frontend, a malformed hint is stripped so
its bare UI token does not fire, and hint:no overrides genuine UI language.
hasRow keyed on a bare '| ID |' which could match the ID as the first cell of a
non-traceability table elsewhere in REQUIREMENTS.md, suppressing a real
table_unmatched signal. Require a second cell ('| ID | <phase> |') so only a
traceability-row shape counts as a row.