next
5966 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ba25e3989d |
fix: pin transitive hono dependency to >=4.13.5 (moderate advisory)
A moderate-severity advisory chain (GHSA-gqvv-2mrq-wpjv,
GHSA-g6gw-c38x-mqfc, GHSA-crvj-82cr-hjcx) is published against
hono <4.13.5, pulled in transitively via @anthropic-ai/claude-agent-sdk
-> @modelcontextprotocol/sdk. This has been blocking
tests/npm-integrity-gate.test.cjs identically across every issue in
this session's bug-fixer sweep -- fixed here, in #4460's own PR, per
explicit direction, rather than waiting on a separate tracking issue.
Adds "hono": ">=4.13.5" to package.json's existing overrides block
(same pattern already used for qs, body-parser, @hono/node-server).
npm audit --omit=dev now reports 0 vulnerabilities.
Note: an equivalent fix (commit
|
||
|
|
10ad91dafb |
fix(#4460): drop the superfluous allow-test-rule marker
local/no-source-grep's looksLikeSourcePath only matches readFileSync targets ending in .cjs/.cts/.js/.mjs/.mts/.ts -- WORKFLOW_PATH here points at code-review.md, so the rule can never fire regardless of the marker. Confirmed by reading eslint-rules/no-source-grep.cjs directly before removing it, not assumed. Caught by an independent code-review pass on the sibling #4466 fix, which copied this same now-unnecessary marker pattern -- fixed there too. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7fe3fd80c2 |
fix(#4460): stop sourcing Tier 1's fence -- confirmed non-portable path check
CI diagnostics from the round-6 push (captured via the new stderr-wired harness) gave the actual root cause after three prior guesses failed identically: on Windows CI, `git rev-parse --show-toplevel` returns a mixed-format path (C:/Users/..., drive letter + forward slashes) while GNU realpath (also bundled with Git for Windows) returns a genuine POSIX path for the identical location (/c/Users/...). Tier 1's REPO_ROOT-prefix containment check can never match between these two formats, so every --files entry is misclassified as "outside the repository" on every Windows run -- deterministically, not flakily, and unrelated to the `-m` flag or 8.3 short names (both already tried and both ineffective). This is a real, structural, pre-existing Tier 1 defect, not something this test can fix without expanding #4460's scope (same out-of-scope bucket as #4461, code-review.md's fences not being cross-platform- robust -- see cr-2 in the review notes). The correct fix is the same treatment already applied to Tier 2: stop sourcing Tier 1's fence, and seed REVIEW_FILES directly with the value a working Tier 1 would have produced. This isolates the test to Tier 3's own gate -- the actual subject of #4460 -- from Tier 1's unrelated defect, on every platform. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
779fff03a4 |
fix(#4460): drop unused execFileSync import (lint-tests finding)
Left over from round 5's switch to spawnSync for stderr capture -- CI's lint-tests job (eslint --max-warnings 0) caught the now-unused import. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
eb4b8ba8c9 |
fix(#4460): fix two self-inflicted bugs in round 5's diagnostic harness
Round 5's own gsd-test run failed on Linux (not just Windows), proving
the instrumentation itself was broken, not Tier 1/3:
1. Attaching `__diagnostics` directly onto the returned files array made
assert.deepEqual fail even when the file list was exactly right --
Node's deepEqual compares an array's own properties too, so a decorated
array never structurally equals a same-valued plain array literal.
Switched runTiers() to return {files, diagnostics} instead.
2. The new "[diag] REPO_ROOT=$REPO_ROOT" probe referenced REPO_ROOT
unconditionally, but Tier 1 only sets it inside its own `if [ -n
"$FILES_OVERRIDE" ]` body -- under `set -u`, the "without --files" case
(FILES_OVERRIDE empty) hit an unbound-variable exit before Tier 3 ever
ran. Default-expanded to ${REPO_ROOT:-<unset...>}.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
f6d610788e |
test(#4460): capture Tier 1/3 diagnostics instead of guessing a 4th cause
Round 4's fs.realpathSync.native() fix did not resolve the Windows CI failure either -- the identical widened-to-5-files symptom recurred a third time on PR #4552, proving that diagnosis was also incomplete or wrong. Rather than guess a fourth root cause blind, switch the harness from execFileSync (which discards stderr) to spawnSync capturing it, redirect the tiers' own diagnostic echoes (previously discarded via `> /dev/null`) to stderr instead, and add explicit "[diag] REPO_ROOT=" / "[diag] REVIEW_FILES(post-tier1)+=" probes right after Tier 1 runs. A future failure now carries what Tier 1 actually computed instead of requiring another round of speculation. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
14e84e1fc5 |
fix(#4460): use fs.realpathSync.native for the Windows-CI tmp fixture path
Round 3's `.replace(/\brealpath -m\b/, 'realpath')` workaround did not actually fix the "test (windows-latest, 24, shard 1/3)" failure -- the same widened-to-5-files symptom recurred identically on PR #4552's next push, proving the `-m` flag was never the real cause. Root-caused via tests/helpers.cjs's own documented Windows caveat (tmpRootCandidates(), ~line 369): GitHub's Windows runners report os.tmpdir() in the 8.3 SHORT form (C:\Users\RUNNER~1\...), and fs.realpathSync() -- what this test used -- does not reliably expand that; only fs.realpathSync.native() does. The un-expanded short-form tmpDir path this test's Node side used for cwd/file construction can diverge from what bash's own `git rev-parse --show-toplevel` / `realpath` independently resolve inside Tier 1's containment check, which is exactly the failure mode observed: --files gets classified as "outside the repository", REVIEW_FILES stays empty, and control falls through to the full-diff path instead of exercising the gate under test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4e5e097544 |
fix(#4460): strip realpath -m from Tier 1 in the test (Windows CI)
CI's test (windows-latest, shard 1/3) failed: --files=src/alpha.js
widened to all 5 files instead of staying at 1. Root-caused (not
assumed): this is the SAME pre-existing Tier-1 `realpath -m`
portability gap already documented as out-of-scope for this fix
(confirmed on macOS during manual verification) -- also real on
Windows CI. `-m` only changes behavior for a path that doesn't (yet)
exist; on a platform where it errors or behaves differently, every
--files entry gets misclassified as "outside the repository",
REVIEW_FILES stays empty, and the test ends up exercising the OUTER
`if [ ${#REVIEW_FILES[@]} -eq 0 ]` full-diff fallback instead of ever
reaching the elif this fix's own gate lives on.
Fixed in the TEST only (code-review.md's Tier 1 is untouched -- this
gap is real, pre-existing, and out of #4460's scope per cr-2). Strip
`-m` from the extracted Tier 1 fence before running it: every path in
these fixtures already exists, so `-m` is a behavioral no-op here, and
this makes the test exercise Tier 3's gate (the actual subject of this
fix) on every platform gsd-test runs on. Manually re-verified both
cases locally before re-pushing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
bbbcf43631 |
docs(#4460): backfill changeset PR number
pr: 0 -> pr: 4552 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d3e3a8f535 |
fix(#4460): restore compact-file reachability + fix test stdout capture
Two more gsd-test-surfaced findings:
1. The previous execute-plan.md trim removed the literal
`summary.compact.md` filename mention, breaking
tests/compact-content-variant-guard.test.cjs's reachability check
(ADR-4139 Phase 6): every registered .compact.md variant must be
named by at least one workflow "spine" file, and execute-plan.md was
apparently the only spine naming this one. Restored the bare
filename (kept the shortened surrounding wording) -- read
tests/helpers/compact-content-variant.cjs's checkReachability/
isUnprefixedMatch directly to confirm the fix rather than guessing.
40926 bytes, still 34 under the size cap.
2. The redesigned test (previous commit) still failed: both tiers'
diagnostic `echo`/`printf "Warning: ..."` lines were mixing into the
captured stdout the assertions parse as the file list, so
"--files=src/alpha.js" appeared to produce 2 lines instead of 1.
Wrapped both tier fences in a `{ ...; } > /dev/null` brace group
(not a subshell -- REVIEW_FILES still persists to the enclosing
shell) so only the final printf reaches stdout. Manually re-verified
both cases against a real git fixture before re-running the suite.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
db7a349a8c |
fix(#4460): trim execute-plan.md under its size budget (unrelated regression)
gsd-test surfaced a SEPARATE, unrelated failure while re-verifying this branch: tests/workflow-size-budget.test.cjs found execute-plan.md at 40981 bytes, 21 over the 40960 DEFAULT hard cap. Root-caused (not assumed): already-merged PR #4540 (enhance(#4139), unrelated to #4460/#4459/#4461) added two near-identical explanatory parentheticals about .compact.md template variants across two nearby steps (user_setup, create_summary), pushing the file over. Confirmed directly against origin/next independent of any merge with this branch -- `next` itself already carries this. This branch's fork point predated PR #4540's merge, so gsd-test's merge-testing against the current next only now surfaced it (merged origin/next into this branch in a separate commit first -- 0 conflicts, after discovering and fixing that this worktree's git clone was SHALLOW, via `git fetch --unshallow`, which is what made a plain `git merge origin/next` fail with "refusing to merge unrelated histories"). Fixed by trimming the SECOND (of two near-identical) parentheticals in the create_summary step to a short back-reference to the first -- same information, no duplication, no cap raised (the test explicitly warns against raising it). 40931 bytes, 29 under the cap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5946926b94 |
fix(#4460): rework test to not depend on Tier 2's broken bash (#4461)
A fresh code-review pass found the test's original approach (concatenate and execute Tier 1 + Tier 2 + Tier 3 verbatim, matching the issue's own reproduction) cannot run: Tier 2's own fence -- untouched by this diff -- is not currently parseable bash. Two unescaped `"` inside its embedded `node -e "..."` regex literal (`raw.replace(/^['"]|['"]$/g, '')`) terminate the outer double-quoted string early, which breaks bash's PARSE of the whole concatenated script even though Tier 2's body never executes under --files. Independently confirmed via manual extraction and execution before accepting the finding. This is a real, separately-filed, already-queued sibling issue (#4461, filed by #4460's own reporter specifically to avoid folding it in here) -- not fixed in this PR. Instead reworked the test to run only Tier 1 + Tier 3 verbatim, seeding the Tier-2-equivalent REVIEW_FILES state directly for the "without --files" case (documented in the module docblock, explaining why Tier 2 isn't sourced and pointing at #4461). Also fixed a nit from the same review pass: a code comment overstated Tier 2's guard as "immediately above" when it's ~150 lines away. Manually re-verified both test cases against a real git fixture with a GNU-realpath-compatible `realpath` (matching gsd-test's Linux bench -- this Mac's BSD realpath lacks the `-m` flag Tier 1 uses, a SEPARATE pre-existing portability gap surfaced during this check, masked on Linux CI, not touched by this fix) before re-running the full suite. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b1c78f0d2e |
fix(#4460): gate Tier 3's #2666 cross-check on FILES_OVERRIDE
code-review.md states (line 144) "Skip SUMMARY/git scoping entirely when --files is provided." Tier 2 honors this via `if [ -z "$FILES_OVERRIDE" ]`, but Tier 3's #2666 SUMMARY/diff cross-check had no FILES_OVERRIDE reference at all -- reached via `elif [ -n "$DIFF_BASE" ]` whenever REVIEW_FILES was already non-empty (true under --files, since Tier 1 fills it), so it silently appended the whole phase's changed files onto an explicit user-supplied file list. --files is documented as the highest-precedence scoping tier (D-08) and is the flag Tier 3's own fail-closed path recommends when no reliable diff base is found; a user narrowing a review to two files silently got the whole phase instead, and the reviewer agent spent its budget on files nobody asked about. Gated the elif on the same condition Tier 2 already uses: elif [ -z "$FILES_OVERRIDE" ] && [ -n "$DIFF_BASE" ]; then The issue's own narrowest suggested form, reasoned through against two alternatives (wrapping the whole Tier-3 fence, or changing the stated invariant instead) -- both explicitly rejected there for good reasons concurred with after reading the surrounding code. Added tests/code-review-tier3-files-override-scoping.test.cjs, mirroring the issue's own verified reproduction methodology: extracts the Tier 1/2/3 fences VERBATIM from code-review.md (never reimplemented) and runs them against a real constructed git fixture matching the issue's own scenario exactly (5 files, a SUMMARY listing only 1). Confirms --files stays scoped to exactly the requested file, and separately confirms the #2666 cross-check still widens a genuinely partial SUMMARY scope when --files is absent (proving this is a gate, not a blanket disable). Emitted-Drift-Ack-Growth: code-review.md — #4460 gates the Tier-3 #2666 cross-check on FILES_OVERRIDE, matching Tier 2's own guard, net +453 bytes Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ab0405ad22 |
test(#4514): migrate git-adjacent workflow checks to named timeout constants
Batch 3 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/ci-rebase-check.test.cjs, tests/gsd-validate-commit-crash-policy.test.cjs, tests/pr-branch-planning-filter.test.cjs, tests/reapply-verify-hunks.test.cjs, tests/ship-notes-wedged-pr.test.cjs, tests/slug-derivation-drift-guard.test.cjs, and tests/worktree-safety.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 7 files from the rule's allowlist. Mints a new shared class norm, QUICK_SPAWN_TIMEOUT_MS (10000ms), in tests/helpers/timeouts.cjs: 5 sites across 4 of this batch's files had independently arrived at the same value for the same shape (a cheap, trivial subprocess/hook invocation with no real git/network/fan-out work). No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
fd37b6a171 | Merge pull request #4560 from open-gsd/test/4513-batch2-git-plumbing | ||
|
|
137f3115a1 |
fix(#4492): index the suffix window instead of pattern-matching it (#4539)
* fix(#4492): index the suffix window instead of pattern-matching it `MSG_SUFFIX="${CMD#*"$MSG_MATCH"}"` is quadratic in the -m message. bash tries every prefix length and compares the whole matched literal at each, and MSG_MATCH is BASH_REMATCH[0] — the entire `-m "..."` — so the cost grows with the thing being scanned. Measured on the real hook: 10.0s at 64KB, 22.0s at 96KB, 30.2s at 112KB, 40.1s at 128KB. `bash -x` with an EPOCHREALTIME PS4 attributes 10.116s of a 10.2s run to that one expansion, which computes an empty string. Conforming and non-conforming cost the same, so this is the path every commit takes, and Claude Code blocks on PreToolUse hooks. MSG_PREFIX on the line above has already located the match, so the suffix is arithmetic rather than a search. Same first-occurrence assumption both expansions always made — MSG_MATCH is a literal substring of CMD by construction. Equivalence checked across 480 comparisons on bash 3.2.57 and 5.3.15 under C, UTF-8 and SJIS locales, including multibyte text, repeated matches, metacharacters and invalid bytes. Three regression rows, all deliberately on the RESOLVE=1 path so they pin the suffix scan alone and do not depend on the separate #4429 SIGPIPE fix: non-conforming and conforming 112KB heredocs, plus a suffix-window row whose padding sits before the heredoc opener's newline so the COMMAND is large while the message stays small. Red against the true base — all three killed at the 10s bound with the head -1 sites still present — and green with only this change. Fixture sizes stay under Linux MAX_ARG_STRLEN (131072 on a 4KB-page kernel). Above it execve fails, the classifier cannot launch and the hook fails open, so a larger fixture measures the argument limit rather than the suffix scan; an earlier 131225-byte draft passed on base AND head for exactly that reason. Every row asserts empty stderr, which is what separates "validated" from "failed open". The bound is enforced by killing the process GROUP, not the direct child: the hook spawns a node classifier that inherits stdout, so killing only bash can leave the pipe open and `close` never arrives. `local/no-elapsed-assertion` forbids asserting on elapsed time, and `{ timeout }` is inert on a synchronous body, so the rows are async and the kill is the signal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte * chore(#4492): add changeset Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
42c02a00c0 |
enhance(#3418): report real codebase drift instead of the whole repository (#4124)
* fix(#3418): write the codebase-drift baseline from code instead of agent prose writeMappedCommit shipped correct and callerless, so no full map-codebase run ever wrote last_mapped_commit. The gate then read null and diffed HEAD against the empty tree, reporting every tracked file as newly added on every run. Adds the stamp-codebase-map leaf verb and calls it from the map-codebase workflow and the execute-phase auto-remap path, replacing the prose instruction that asked the mapper agent to stamp its own output. An agent that concludes its work is already done skips a prose step silently, which is the failure the stamp exists to detect. The gate now reports an absent or unresolvable baseline as skipped, with reason no-mapped-commit or unresolvable-mapped-commit, rather than as whole-repo drift. Files under .planning/ are excluded from the diff so the map's own commit does not read as seven new directories on the next run. Emitted-Drift-Ack-Growth: map-codebase.md — adds the stamp_codebase_map step and its rationale, new workflow content this change requires * test(#3418): cover the stamp writer and the absent-baseline gate * docs(#3418): document how the drift baseline is written and skipped * docs(#3418): note that a manual stamp reflows the map's whitespace writeMappedCommit writes through platformWriteSync, which normalizes markdown whitespace on .md targets. Run in its workflow position the stamp lands on documents the mapper just wrote, so the normalization is folded into the same commit, but a hand-run stamp over an already-committed map reflows that map as a side effect. Reported on the issue thread. * chore(#3418): add changeset fragment Typed Changed to match the enhancement route the linked issue's label sets. The docs-required lint is satisfied by the ARCHITECTURE.md update already in this branch. * fix(#3418): anchor the planning-artifact filter to the repo root git diff --name-status prints repo-root-relative paths whatever the cwd, so computing the exclusion prefix against cwd yielded ".planning/" while git printed "sub/.planning/" and the filter silently matched nothing from a subdirectory. * fix(#3418): derive the planning prefix from git, not from path arithmetic Anchoring the exclusion prefix with path.relative() against `rev-parse --show-toplevel` broke on Windows, where os.tmpdir() hands back the 8.3 short form and git resolves the long one, so relative() produced a "../.." chain that matched nothing. `rev-parse --show-prefix` gives the cwd's root-relative prefix from the same producer as the diff paths, so the two sides cannot disagree. * fix(#3418): take the planning lock around the codebase-map stamp Stamping seven documents is seven frontmatter read-modify-writes, and two stampers can run at once: the full map-codebase run and the execute-phase auto-remap. Wrap the write loop in withPlanningLock, the same lock the other .planning/ writers take, so a concurrent pair cannot lose an update. Also corrects the path-arithmetic comment, which read as if the Windows short-path hazard applied to the .planning half of the prefix. It applies to the rejected --show-toplevel alternative; both sides of the surviving relative() call are the same cwd string. * fix(#3418): read HEAD and the map file list under the planning lock The stamp resolved HEAD and listed the present codebase-map documents before it acquired the planning lock, so a stamper that then waited on the lock could write its now-stale sha over a newer one, or recreate a document deleted while it waited as a frontmatter-only stub. Both reads now happen inside the lock, matching the read-and-write-in-one-lock pattern config.cts and phase.cts already use. An empty --files value is refused as well instead of silently widening the stamp to all seven documents. * fix(#3418): narrow the map stamp to the documents an update run refreshed An "Update - only update specific documents" run reached the new stamp step with no --files narrowing, so the six documents the user did not select were stamped at HEAD and read as freshly mapped. The selection now threads through to --files, the same way the auto-remap path already does. A bare --files (an unquoted empty shell variable drops the token) parsed to null, indistinguishable from an absent flag, so it skipped the empty-filter refusal and stamped all seven. Presence is now read off argv. * fix(#3418): require the drift baseline to resolve to a commit, not any object `git cat-file -t` exits 0 for a tree or blob sha and for a ref name, and `git diff <tree> HEAD` is valid, so an exit-code-only probe accepted a baseline that is not a commit and reported the resulting diff as real drift. Check the reported type instead of the exit code alone, which routes every non-commit stamp to the same `unresolvable-mapped-commit` skip. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3ad75a6d59 |
enhance(#4285): resolve context-monitor fire-points from .planning/config.json (#4366)
* enhance(#4285): resolve context-monitor fire-points from .planning/config.json
The monitor's WARNING (35%) and CRITICAL (25%) fire-points were module
constants, so the only way to tune them was editing gsd-context-monitor.js —
a file in the MANAGED hooks registry, whose body the next install re-stages,
silently discarding the edit. The alternative was turning the safety net off.
Both are now readable from the config block the hook already opens:
hooks.context_warning_threshold and hooks.context_critical_threshold. Absent
keys resolve to today's 35/25, so every existing project is byte-identical.
Resolution is total and never throws — this hook must not block the tool call
it rides in on. A value is usable only if Number.isFinite (type-strict, so the
string "30" and true are rejected) and inside the 0-100 domain of the
remaining_percentage it is compared against; anything else falls back to the
default. The PAIR falls back together: critical >= warning has no coherent
reading, and honouring one side silently picks which of the operator's two
numbers to discard. That also covers a single override contradicting the other
key's default.
config-set validates the domain per key so accept and honour agree, but
deliberately does not enforce the pair — it writes one key per call, so a
two-step retune is transiently inconsistent on disk and refusing it there
would block a legitimate configuration.
Registration follows the statusline.show_git precedent: schema manifest plus
src/config.cts validation, not config-defaults.manifest.json and not
buildNewProjectConfig — emitting 35/25 into every new project would pin the
defaults at creation time for a setting nobody has tuned.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2
* enhance(#4285): address Codex review — per-key fallback docs, discriminating tests
Codex full-PR review (gpt-6-astra, read-only) returned five findings. Each was
verified against source before acting; all five are real.
1. docs/CONFIGURATION.md described the wrong fallback. An out-of-domain value
falls back PER KEY; both defaults apply only when the RESOLVED pair violates
critical < warning. warning 150 with critical 30 resolves to 35/30, not
35/25 — at remaining 28 that difference changes the severity emitted. The
table now states the two rules in the order they compose, and
docs/context-monitor.md gains the same worked example.
2. The inconsistent-pair test could not prove the CRITICAL side reverts: its
pair was 20/25, and 25 is already the default, so an implementation that
reset only `warning` passed it. A 45/50 pair — both halves away from their
defaults — now pins each side with its own reading, and an equal 45/45 pair
pins that the rule is strict (`<`, not `<=`).
3. The rejection table's rows could not tell rejection from acceptance: an
accepted -5 pairs with the default critical 25, trips the pair check, and
produces the same silence. Two rows now separate those: a below-domain
critical must escalate remaining 20 to CRITICAL (proving -5 was rejected,
not honoured), and an unusable critical beside a usable warning 45 must
still fire WARNING at remaining 40 (proving per-key fallback rather than
reset-both). The over-claiming comments are narrowed to what each row
actually shows.
4. Scope, reproduced rather than assumed: config-set writes through
planningDir(), so under GSD_WORKSTREAM it lands in
.planning/workstreams/<name>/config.json while this hook reads only
<cwd>/.planning/config.json. That is the pre-existing root-only scope
hooks.context_warnings has always had, but this PR advertises the setter
route, so both docs now say the keys are root-project settings.
5. Four other English docs still stated 35/25 as fixed: the REQ-CTX-02/03
requirements fragment, ARCHITECTURE.md's hook table and threshold table,
and INVENTORY.md's hook row. All now name them as defaults and point at the
config keys; docs/FEATURES.md is regenerated from its fragment via
scripts/gen-features.cjs --write, not hand-edited.
Four new mutations, each reverted after: resetting only the warning half on an
inconsistent pair (1 red), resetting both on any unusable key (1), dropping the
>= 0 bound (1), and accepting critical == warning (1). perf-317 is 116/0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2
* enhance(#4285): tighten claims after Codex round 2 — scoped paths, one more discriminator
Confirmation round found no runtime defect and confirmed the five round-1 fixes
landed. Four precision items, all real, all fixed here.
1. The scoped-write note named the wrong path for GSD_PROJECT. planningDir()
composes three distinct shapes, confirmed by running config-set under each:
.planning/<project>/config.json, .planning/workstreams/<ws>/config.json, and
.planning/<project>/workstreams/<ws>/config.json. docs/context-monitor.md
now tabulates all four cases instead of collapsing them into one.
2. The 45/50 silence row asserted empty stdout without pinning the exit code.
runMonitorRaw turns a spawn failure, a non-zero exit or a timeout into empty
stdout as well, so the row could have passed on a dead child. It asserts
exitCode === 0 first now, like the equal-pair row already did.
3. The sibling row's message claimed it proved critical fell back to 25. It
does not: coercing '30' to 30 yields WARNING at remaining 40 too, so the row
pins the WARNING side surviving and nothing more. Message narrowed, and a
new row reads the same config at remaining 28, where the two candidate
resolutions diverge — rejected gives (45, 25) and WARNING, coerced gives
(45, 30) and CRITICAL. Mutation-verified: swapping Number.isFinite for the
coercing global reds it.
4. "Accept and honour must agree" was too absolute in the src/config.cts and
tests/config.test.cjs comments. The agreement holds on the DOMAIN and per
key: an accepted value can still lose to the hook's pair check at read time,
and a scoped write never reaches the hook at all. Likewise a two-step retune
only CAN be transiently inconsistent — 35/25 to 20/10 is valid throughout if
critical moves first — so the docs now say what a setter-side pair check
would actually cost: rejecting that intermediate write and forcing an order.
The same over-absolute phrasing is in b7d179c89's message, which is left as
written rather than rewriting history; this commit and the PR body carry the
precise claim.
perf-317 117/0, config 192/0, config-field-docs 47/0, features-index-gate 84/0,
lint:ci clean cold.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2
* chore(#4285): add changeset
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DsUAawHKUy9pCpnye1Jzd2
* enhance(#4285): address review — planning-config rows, resolveThresholds properties
Two Minor findings from the maintainer review, no behaviour change.
Minor 1: gsd-core/references/planning-config.md's "Hook Fields" table gains
rows for hooks.context_warning_threshold and hooks.context_critical_threshold,
in that table's 5-column form, carrying the same per-key-fallback,
pair-reversion and root-config-scope claims docs/CONFIGURATION.md already
makes. hooks.workflow_guard's absence from that table is pre-existing and
out of scope here.
Minor 2: resolveThresholds() gets fast-check property coverage, which ADR 456
requires of a threshold/limit contract. Reaching it needed a require-time
seam: the resolver was previously observable only by spawning the hook, and a
subprocess per case cannot drive 200 runs — the same conclusion CONTEXT-INDEX
records for the ROADMAP Requirements parser. The stdin adapter therefore moves
into main() behind `require.main === module`, mirroring
gsd-cursor-subagent-start.js and gsd-statusline.js, and module.exports exposes
the resolver plus both default constants so a test asserts fallback against
the source of truth rather than a second copy of 35/25. Spawned behaviour is
unchanged: the 10s stdin timeout still arms per invocation (stdinTimeout is
now a module-scope let assigned in main(), still cleared by the end handler),
and the try/catch crash(ON_CRASH) path is untouched.
Seven properties: totality, ordering, exactness, togetherness, non-vacuity,
per-key fallback, non-object argument. Exactness is stated PER KEY — a mixed
result (one key honoured, one fallen back) is legal and is the documented
contract; the property falsified a per-pair phrasing of it in 4 runs.
Verified: cold lint:ci 0; perf-317 file 125/0; seven mutations killed and
restored, one of which (upper bound widened to 120) is invisible to the 17
hand-written cases and caught only by a property.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte
* enhance(#4285): close the Codex-found gap in the property coverage
Codex whole-PR review of round 3 returned no Blocker and no Major. Two items,
both in the tests added this round, both verified against source before acting.
Minor — the per-key fallback property was asymmetric: it required a usable
warning to survive an unusable critical, but never the reverse. A resolver
that reverted BOTH keys the moment warning was unusable passed all seven
properties. Reproduced exactly: that mutant answers 35/25 for
{warning: 150, critical: 30} where the resolver answers 35/30, and the file
stayed green at 125/0. The mirrored property closes it — with the mutant
re-applied it is now the single failing row, and it is the only row that
fails, so it is load-bearing rather than incidental.
Nit — the ordering property's comment credited it with catching a
half-honoured pair, which it does not: 45/50 "repaired" by resetting only
critical yields 45/25, perfectly ordered. That case belongs to togetherness.
The same comment claimed the behavioural rows sample an inconsistent pair at
exactly one point; stale — they cover 20/25, 45/50 and the 45/45 equality
boundary. Both claims corrected in place.
Verified: cold lint:ci 0; perf-317 file 126/0; the mutant above killed by the
new property alone and the hook restored byte-identical afterwards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UZw5UhR474YLyE4knjHrte
* enhance(#4285): name the installed-monitor prerequisite; close the negative-critical gap
Second Codex whole-PR pass, run because the base moved: the author's three
"Update branch" merges pulled ~26 upstream commits in, so the previously
reviewed diff sat on a base that no longer exists. No Blocker, no Major, two
Minor — both verified against source before acting.
Minor 1, and only reachable because of what the merge brought in: #2586
(
|
||
|
|
6c5e11049b |
fix(#4383): require phase before planned-phase writes (#4534)
* fix(#4383): require phase before planned-phase writes * chore: add changeset for #4534 --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c9d3e66631 |
fix(#4341): reference-count the git-config sandbox the three guard suites share (#4389)
* fix(#4341): reference-count the git-config sandbox the three guard suites share node:test evaluates every describe body during collection, before any test runs, so the three isolateGlobalGitConfig() calls happened back to back and each captured the PREVIOUS call's temp path as its "original": A captured undefined, B captured A's path, D captured B's. Suite A's after() fires first and restored undefined — deleting GIT_CONFIG_GLOBAL outright — so suites B and D ran against the developer's real ~/.gitconfig for the rest of the file. A also cleanup()'d a directory B still pointed at. One sandbox now, with the true original captured once and released when the last holder lets go; each returned restorer is idempotent, so an extra call cannot release someone else's hold. The three call sites are unchanged. Reproduced with a global core.hooksPath (via a fixture HOME carrying a .gitconfig, so the developer's real one is never touched): next: ℹ pass 291 ℹ fail 7 — "core.hooksPath is set to ...; a hook written to .../pre-commit would never run" branch: ℹ pass 296 ℹ fail 3 The 3 that remain are two suites (pr-subrepo, #3776 query commit --files) that never called isolateGlobalGitConfig at all — the same class, a different gap, and outside this issue's scope. Noted on the PR. D0 is the regression guard and is deterministic on every lane: suite A's after() runs before this suite's tests, so on the old helper GIT_CONFIG_GLOBAL is already gone by then regardless of what the host's git config contains — which is what makes it fail on CI, where the core.hooksPath that exposed the defect is absent. Verified: full-file run reds on the old helper, greens on the new one. * test(#4341): clean filtered gitconfig sandboxes on exit * test(#4341): retain exit cleanup until release succeeds |
||
|
|
342bd8ca6b |
fix: backfill changeset fragment PR number for #4560
Forgot the pr:0 -> real-number backfill step from CONTRIBUTING.md's documented changeset workflow after opening PR #4560, which broke changeset-lint and docs-lint (fail_invalid_fragment / fail_malformed_fragment). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bba20b51fd |
fix: pin transitive hono dependency to >=4.13.5 (moderate advisory)
A moderate-severity advisory chain (GHSA-gqvv-2mrq-wpjv, GHSA-g6gw-c38x-mqfc, GHSA-crvj-82cr-hjcx) was newly published against hono <4.13.5, pulled in transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk. Discovered blocking tests/npm-integrity-gate.test.cjs while verifying #4513 (unrelated to that batch's diff); fixed inline per this repo's no-defer policy. Adds "hono": ">=4.13.5" to package.json's existing overrides block (same pattern already used for qs, body-parser, @hono/node-server). npm audit --omit=dev now reports 0 vulnerabilities. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b33df03726 |
enhance(#4089): add minimum-solution reasoning check (#4118)
* enhance(planning): add minimum-solution reasoning check * chore: add changeset for planning guidance * chore: bind changeset to PR 4118 * docs: document planning sufficiency check * docs: distinguish planning sufficiency guidance --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
95e6a58fd4 |
test(#4513): migrate git plumbing batch to named timeout constants
Batch 2 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/git-base-branch.test.cjs, tests/commit-files-pathspec.test.cjs, and tests/git-fixture.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 3 files from the rule's allowlist. This batch introduces a violation shape not seen in Batch 1: several sites are pinned-value test assertions verifying the EXACT timeout production code hardcodes (not bounds on this suite's own subprocess calls). Named as four separate constants even where values coincide, so the tests keep catching independent production drift instead of silently tolerating it. No src/bin file touched, no numeric value changed anywhere -- verified site-by-site by two independent isolated review passes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ff5c5782ee |
test(#4512): migrate process-seam batch to named timeout constants
Batch 1 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/process-seam.test.cjs, tests/helpers-process-isolation.test.cjs, tests/run-with-timeout.test.cjs, and tests/helpers.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 4 files from the rule's allowlist. Adds SEAM_DEFAULT_TIMEOUT_MS to tests/helpers/timeouts.cjs (shared across 2 batch files, mirroring process-seam.cjs's own un-exported default). File-local constants elsewhere for values not shared across files or not a bench-derived class norm. No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c9016b2e35 |
docs(#4554): add changeset fragment
changeset-lint flagged the PR as touching user-facing shipped content (execute-plan.md, a workflow instruction file) with no fragment. Type Fixed per CONTRIBUTING.md — Fixed/Security changesets are exempt from the docs/-touch requirement Added/Changed/Deprecated/Removed carry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8bcf633e21 |
fix(#4554): trim execute-plan.md 21 bytes under its DEFAULT size-tier cap
gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes, tests/workflow-size-budget.test.cjs, ADR-1610) at next@a27cb6b2fa — introduced by #4540's call-site wiring for the summary.md/ user-setup.md .compact.md variants, which nobody caught crossing this exact margin before merge. This trips next's own Tests run on every shard/OS combination, which in turn blocks the repo's Base branch health PR gate (#4422/#4428) for every open and future PR regardless of that PR's own diff. Two meaning-preserving trims in the <success_criteria> block: a repeated parenthetical ("— unless parallel mode (orchestrator handles)", appearing twice) replaced with a "— same exception" back-reference on its second occurrence, and one redundant qualifier ("prominently") dropped — its behavioral content (surface the USER-SETUP.md warning at the TOP of output) is already fully specified earlier in the same file. 40,981 -> 40,940 bytes, 20 bytes of headroom under the cap. No procedural content lost. Fixes #4554. Emitted-Drift-Ack-Hash: gsd-core/workflows/execute-plan.md — deliberate content trim to clear the DEFAULT size-tier cap (#4554); not a regeneration artifact, a hand-authored byte reduction. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a27cb6b2fa |
enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b (gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4 (gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user: two independent, complete files per covered path (canonical + .compact.md sibling), with the gate picking which one gets Read at the call site. This is a different shape from Phase 5's spine+detail partition, and is safe here specifically because these files are already reached only by a runtime Read — a missed Read already means zero overlay content today, with or without workflow.compact_content, so selecting between two independently-complete files introduces no new failure mode (documented in gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section). Disposition, after inspecting every candidate rather than trusting a byte-size threshold (same rigor Phase 5 applied to review.md): - Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing reference doc emitted verbatim, not orchestrator instruction). The other 9 size-threshold candidates are dominated by fail-closed guards, exact CLI invocations, or output-format contracts (AskUserQuestion blocks) — recorded not-worth-compacting, same reasoning as Phase 5's review.md. - Stream 4: a ground-truth reachability audit replaced the initial size-only candidate list. Two files (summary.md, user-setup.md) got compact variants; a third (spec.md) was drafted, then dropped after discovering its only two call sites are eager @-includes, not a runtime Read — stream-1 material hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md itself has 3 eager call sites and only 1 genuine runtime-Read call site (execute-plan.md); only that one was wired, so the compact variant's savings apply to the sequential single-plan execution path only. - Discovered while auditing reachability: 12 gsd-core/templates/** files with zero references anywhere in workflow/agent/command prose, compiled source, or tests — dead scaffolding predating this phase. Deleted in this same PR per this repo's no-defer policy, after re-verifying against a computed path.join(...) pattern (not just a plain-string search) that nearly caused two genuinely load-bearing templates (user-profile.md, dev-preferences.md) to be misclassified as dead. New checker (tests/helpers/compact-content-variant.cjs): registration, reachability, protected-content-preserved, size-smaller — replacing Phase 3/5's disjointness/completeness checks, which assume a partition rather than two deliberately-overlapping documents. The reachability check's own "unprefixed match" guard had a real bug (rejected the repo's own `~/.claude/gsd-core/...` convention), caught by running it against the already-wired help/modes/full.compact.md pair rather than only synthetic fixtures — fixed to anchor on the nearest `gsd-core` path segment instead. Template consumer parity (tests/compact-content-template-variant-parity.test.cjs): proves each compact variant's `## File Template` fenced block — the actual output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed against — is byte-identical to the canonical file, then runs the one real deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent, backing `gsd-tools uat classify-coverage`) against content built from that shared contract. Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs) rather than extending the existing spine/detail one — different data shape, and the existing script's own contract deliberately isolates it from a test-only helper's shape changing. Emitted-drift acknowledgement: not needed. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before reaching the ack-lookup branch (same reasoning Phase 5 verified for its own diff). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * enhance(#4406): address code-review findings on the variant-swap gate - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content entries described only the spine+detail mechanism (Phase 5) and were missing this phase's variant-swap mechanism and its benchmark:compact-content-variants script entirely — required since this PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets" rule). Both now describe both mechanisms and which call sites are wired. - Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved: a canonical file with zero <!-- gsd:protected --> blocks must be a no-op, not a violation — the only branch of that function the existing fixtures didn't exercise. - Collapsed findCompactFiles/findMarkdownFiles in tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix helper — the two were identical recursive walks differing only in the extension predicate (minor Duplicated-Code finding). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification gsd-test caught this, not static analysis: 10 real failures in tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs, and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install path, which does fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md')) after copying gsd-core/templates/** into the target project, then merges it into both .github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the repo-root bin/install.js — a separately maintained installer bundle outside the src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around that read degrades to a silent skip rather than a crash when the template is missing, which is why this surfaced only once the real E2E install test ran, not from any static check. Re-verified the remaining 11 deleted filenames against bin/install.js specifically (plain substring and quoted-filename search) before trusting that list — all 11 have zero hits there, confirmed dead by the same standard this one file failed. Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md) and the phase design doc from 12 to 11 deleted files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule * docs(#4406): backfill changeset PR numbers Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion CI's own full-test matrix (not gsd-test's matrix, which does not run this check) caught 4 more false-positive dead-template classifications via tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs — a literal, word-boundary basename check across .github/workflows/, gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every file a PR deletes. It has no semantic awareness, so a deleted template's basename colliding with something else entirely still fires: - claude-md.md: gsd-core/templates/README.md had a stale table row claiming /gsd-profile reads this template to generate CLAUDE.md. Verified false (no code reads it anywhere, same search that already covered bin/install.js) — fixed the row to *(inline)*, matching every other command-generated artifact in that table. File stays deleted. - codebase/testing.md: collided with docs/guides/testing.md, an illustrative example row in docs-update.md's sample output table (an unrelated real generated-docs path). Swapped the example topic to "contributing" — the row is illustrative, any topic works. File stays deleted. - codebase/architecture.md, codebase/stack.md: collided with docs/reference/ planning-artifacts.md's directory listing of a user's own generated .planning/codebase/architecture.md and stack.md output — the same semantic mismatch already investigated and dismissed as unrelated earlier in this phase's audit, now caught by a gate instead of judgment. That listing repeats across 5 locale copies of the doc. - continue-here.md: collided with the real .continue-here.md pause-work artifact, referenced across 15+ locale and workflow files. For the last two, the lint's own error message offers "restore the file or update every consumer in the same commit." Rewording 15+ files across languages I cannot verify translation quality for, to shave 2 already-tiny templates that were merely presumed dead, is disproportionate to this PR's actual scope — restored codebase/architecture.md, codebase/stack.md, and continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates section accordingly. Final confirmed-dead set: claude-md.md, codebase/concerns.md, codebase/conventions.md, codebase/integrations.md, codebase/structure.md, codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next node scripts/lint-removed-but-needed.cjs now passes clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes * fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the user asked to be actually fixed, not just re-run past: PR #4497 (landed 2026-09-07, one day before this PR's CI run) isolated tests/codex-config.test.cjs into its own dedicated chunk because its measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe to share a chunk with any other file. That isolation was necessary but not sufficient — even alone, with zero companion-file contention, the file's real Windows execution time sits right at the 600s per-chunk ceiling. Two independent CI runs on two unrelated PRs (this one and #4154) were both killed within ~1.4s of the identical 600000ms mark — not random contention, a deterministic near-miss the isolation fix couldn't address because it never reduced the file's own cost, only removed the risk of a companion file's cost stacking on top of it (which the PR #4497 comment explicitly anticipated: "if a future profiling pass genuinely speeds up codex-config.test.cjs itself, this isolation can be revisited"). The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79 describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760, #3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more), several of which are explicitly documented as "folded" in from separate files that were never actually split back out ("Verified non-duplicate against both the pre-existing target and the other three folded sources"). Split into 4 files by top-level AST statement boundaries (never a naive column-0 regex — an early attempt at that overcounted 79 apparent "describe(" matches when only 21 are genuinely top-level; the rest are nested inside a handful of large folded-in blocks, which a regex can't tell apart from real top-level statements). Verified lossless twice: the split script asserts byte-for-byte reconstruction of every source character, and independently, total test()/describe() call counts match exactly between the original file and the sum across all 4 new files (433/79 both sides). Each new file carries the complete original shared header (imports/helpers) for safety; per-file unused-import warnings from that duplication are resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard form for an intentionally-unused destructured binding — never a bare `{ _foo }`, which would destructure a different, nonexistent property). No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its pinned test in tests/run-tests-harness.test.cjs: the file that keeps the original name (tests/codex-config.test.cjs) is now only ~28% of the original's size and safely isolated in its own chunk as before; the other three new files re-enter normal weight-balanced packing, none individually close to disproportionate. Confirmed no other file hardcodes the hardcoded filename anywhere that would silently stop these tests from running (the CI test-selection scripts determine scope algorithmically, not by literal filename). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
66dbb104a0 |
fix(#4459): anchor update_codebase_map's diff base on the phase directory (#4549)
* fix(#4459): anchor update_codebase_map's diff base on the phase directory execute-plan.md's update_codebase_map step derived its diff base from: git log --oneline --grep="feat({phase}-{plan}):" ... --reverse | head -1 A phase number is unique within a MILESTONE, not a repository (#3995). `--reverse | head -1` deliberately selects the OLDEST matching commit subject, so on a milestone that reuses a phase number, the diff base lands in the PREVIOUS milestone's same-numbered phase -- silently widening the file list that then drives which .planning/codebase/*.md files get amended, with no warning and nothing downstream that would notice. This is the same defect class already fixed at two other sites in this repo (code-review.md, structural-pre-pass.md) via a phase-DIRECTORY anchor instead of a commit-subject grep: PHASE_START = the first commit that ADDED anything under the phase directory, diffing from its parent (or the commit itself on a root commit). Mirrored that exact pattern here rather than inventing a new one. Added tests/execute-plan-update-codebase-map-diff-base.test.cjs: static regression guards (old grep gone, new #3995-shaped anchor present) plus a real-execution test reproducing the issue's own scenario -- two milestones reusing a phase number with a real constructed git fixture, extracting and running the step's actual bash fence, asserting the resulting diff is scoped to the current milestone's files only. Emitted-Drift-Ack-Growth: execute-plan.md — #4459 replaces the unbounded commit-subject grep with the phase-directory anchor already used by code-review.md/structural-pre-pass.md, net +687 bytes Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4459): cite this issue in the new file's allow-test-rule marker gsd-test caught tests/lint-allow-test-rule-refs.test.cjs failing: the new test file's `// allow-test-rule: source-text-is-the-product` comment (copied from the two sibling precedent files) was missing the required issue-ref suffix -- ADR-456 requires a NEW exemption to cite an issue via `#NNN` on the same comment line. Added `(see #4459)`. Verified via `node scripts/lint-allow-test-rule-refs.cjs` directly (clean) before re-running the full suite. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4459): backfill changeset PR number pr: 0 -> pr: 4549 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
511c900052 |
fix(#4458): reuse detectSubRepos for new-project.md's sub-repo detection (#4548)
* fix(#4458): reuse detectSubRepos for new-project.md's sub-repo detection new-project.md's Step 5.1 (Sub-Repo Detection) ran its own bash predicate: find . -maxdepth 1 -type d -not -name ".*" -not -name "node_modules" \ -exec test -d "{}/.git" \; -print `test -d` requires .git to be a DIRECTORY. A linked git worktree's .git is a FILE (a `gitdir: <path>` pointer), so this predicate silently excluded valid linked-worktree children while still finding ordinary clones. src/core-utils.cts's detectSubRepos(cwd) already handles this correctly (fs.existsSync, type-agnostic) but had zero callers anywhere in the codebase -- orphaned logic the workflow never actually used, despite duplicating a narrower version of the same check inline. Wired detectSubRepos into cmdInitNewProject's JSON output as a new sub_repos_detected field (matching the file's existing pattern of similar directory-scan-derived fields like has_existing_code/ is_brownfield/has_codebase_map) and replaced the workflow's raw find fence with a gsd_run query init.new-project call reading that field -- removing the duplicate, narrower detection logic entirely rather than patching it in place, per the issue's own "reuse a central policy" framing. Added the missing .git-as-FILE test case to the existing tests/core-utils.test.cjs detectSubRepos coverage (proving the helper was already correct -- the defect was entirely in the unwired workflow predicate) plus CLI-level end-to-end coverage in tests/init-manager.test.cjs using a REAL `git worktree add` fixture, matching the issue's own reproduction steps, alongside an ordinary child-clone case and a non-repository-directory negative case. Refreshed tests/fixtures/compact-content-benchmark-baseline.json (gsd-test caught the drift from new-project.md's byte-count change; the benchmark script itself always exits 0 -- report, not gate -- but the wrapper test enforces the committed baseline stays in sync). Emitted-Drift-Ack-Growth: new-project.md — #4458 replaces the raw find predicate in Step 5.1 with a gsd_run query call reading the new sub_repos_detected field, net +140 bytes Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4458): backfill changeset PR number pr: 0 -> pr: 4548 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4458): bound the new git worktree add spawn with a named timeout CI's lint-tests caught two ESLint findings my local gsd-test run couldn't see (gsd-test's matrix doesn't run npm run lint:ci -- same gap already observed on #4456's PR): - local/no-unbounded-spawn: the new execFileSync('git', ['worktree', 'add', ...]) call had no timeout, an indefinite-hang risk. - local/no-adhoc-timeout-literal: my first fix (a bare `timeout: 15_000` literal) was itself flagged -- two independent hardcoded copies of a guessed timeout can silently drift or collide (this repo hit exactly that on 2026-09-06, PR #4428). Fixed by importing GIT_FIXTURE_TIMEOUT_MS from tests/helpers/timeouts.cjs -- `git worktree add` checks out files into a new working tree, the same "construction" weight class as init/config/add/commit that constant already covers, not plain plumbing (GIT_TIMEOUT_MS's class). Verified via `npm run lint` directly (clean) before re-running gsd-test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4458): reduce redundant git subprocess overhead in new tests CI's full-test Windows shard 2/3 failed: chunk 1/9 (306 files) exceeded its internal 600s budget and was force-killed, with an unrelated file (codex-config.test.cjs) in flight at the moment of the kill -- meaning the chunk's AGGREGATE runtime, not any single hang, blew the budget. This PR's own three new tests each independently called createTempGitProject() (git init + a commit), and one of them also runs git worktree add -- real subprocess spawns, each Defender-scanned on Windows CI (tests/helpers/timeouts.cjs's own documented rationale for why Windows spawn classes get generous budgets). That's a genuine, quantifiable overhead addition to the exact chunk that timed out, not something to wave off as unrelated flake without checking. Two of the three tests never actually needed a real git repo -- detectSubRepos only inspects a CHILD directory's own .git, never the root's git state, and the existing SUBCOMMANDS loop earlier in this same file already proves `init new-project` succeeds against a plain, non-git createTempProject() fixture. Switched those two tests to the lighter fixture, leaving only the one test that genuinely needs a real repo (git worktree add requires one) on createTempGitProject(). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
00b7e622e1 |
test(#4457): cover init.new-milestone's project_exists/project_path under an active workstream (#4547)
#4457 reports init.progress and init.new-milestone reporting project_exists:false (and a workstream-scoped project_path) for a root-shared PROJECT.md after a normal migration. Direct read of current src/init.cts shows this is already fixed on next by PR #4543 (#4455's self-discovered follow-up): withProjectRoot, buildInitCompletenessFields, cmdInitProgress's project_exists, and cmdInitNewMilestone's project_exists/project_path all already resolve PROJECT.md via the root-aware planningDir(cwd, null). PR #4543 added parametrized regression coverage for ingest-docs/resume/progress/ new-project (plus a dedicated milestone-op test), but its own comment explicitly deferred new-milestone's coverage to "#4456's own new-milestone.md workstream-forwarding work" (PR #4543). PR #4545 (#4456) never added it -- its test additions only exercised the *workflow's* --ws argv forwarding through a stubbed gsd_run, never cmdInitNewMilestone's real output. That gap is exactly what #4457's own acceptance ask requests ("Add a migration-to-init regression test with a root PROJECT.md and no workstream-local copy"). Folds 'new-milestone' into the existing SUBCOMMANDS loop in tests/init-manager.test.cjs (it has no manager-style readiness precondition, and exposes both project_exists and project_path, so it fits the loop's existing assertion shape without a dedicated test). No production code change. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
147c89a9b8 |
fix(#4456): forward --ws to every downstream new-milestone.md call (#4545)
* fix(#4456): forward --ws to every downstream new-milestone.md call new-milestone.md's Step 1 parses --ws <name> into GSD_WS, but each workflow step's bash fence is a separate shell invocation — GSD_WS set in Step 1 never survived to Steps 5, 6, or 7. Four call sites never forwarded it: init.new-milestone (both calls), state.milestone-switch, and both phases.clear branches. Under GSD_WORKSTREAM env or a stored session pointer differing from the explicitly requested --ws, every downstream operation silently operated on the wrong workstream (or root) instead of the one the caller asked for. Confirmed --ws is a universally-parsed CLI flag (gsd-core/bin/gsd-tools.cjs: 4867, resolveActiveWorkstream) — stripped from argv and written into process.env.GSD_WORKSTREAM for the rest of that process, so appending it to ANY gsd_run query call works uniformly. Fixed by persisting GSD_WS to .planning/.gsd-ws-arg right after Step 1 parses it (mirroring the established .gsd-outgoing-milestone round-trip idiom this same file already uses for the identical cross-fence problem), reading it back in each later step, and appending it unquoted (matching the ${GSD_WS} splicing convention documented in workstream-flag.md). Cleaned up after its last use in Step 7. Bundled, in-scope fixes found while implementing the above (per this repo's no-defer policy): - Step 6's phase-archive `git add .planning/milestones/ .planning/phases/` hardcoded literal ROOT paths — both directories are workstream-scoped (matching cmdMilestoneComplete's established #1911 precedent), so under a workstream this staged nothing real. Added phases_dir/archive_dir fields to cmdInitNewMilestone and resolved through them instead. - Step 6's milestone-start commit hardcoded .planning/STATE.md — also workstream-scoped, so it would commit the wrong (or a stale) file under a workstream. Resolved through init.new-milestone's existing state_path field instead; PROJECT.md correctly stays a literal-shaped-but-resolved root path (shared, per the #4455 follow-up already merged). - cmdInitNewMilestone's config_path field: config.json is ALSO a shared file (marked `# Shared` in workstream-flag.md's directory diagram, same as PROJECT.md) but was resolved via the workstream-aware planningDir — fixed alongside cmdInitNewProject's identical instance of the same bug (found via grep, matching the precedent from the #4455 follow-up of fixing every occurrence of an identically-evidenced bug uniformly). Verified: direct CLI invocation confirms phases_dir/archive_dir/state_path resolve into the workstream while project_path/config_path stay root under GSD_WORKSTREAM=alpha. Manual bash-fence execution of every modified fence (Step 1 parse+persist, Step 5 forwarding, both phases.clear branches, the git add fence, Step 7's forward+cleanup, the commit fence) confirms correct behavior in both flat and --ws modes, including two flags composing together (--archive-version + --ws; --reset-phase-numbers + --ws). Rewrote the pre-existing "step 6: commit stages PROJECT.md" test, which asserted the literal (buggy) --files string verbatim — it now asserts the resolved paths via a JSON-returning stub, and gained isolated per-test tmp dirs (the prior version ran with no explicit cwd, at real risk of writing a stray .gsd-ws-arg into this repo's own .planning/ once Step 1's fence started performing a real write). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4456): Steps 9/10 also commit workstream-scoped files via literal root paths A fresh isolated code-review pass on the first version of this fix found the identical bug in two more places, missed in the initial sweep: - Step 9's requirements commit (`gsd_run query commit ... --files .planning/REQUIREMENTS.md`) and Step 10's roadmap commit (`--files .planning/ROADMAP.md .planning/STATE.md .planning/REQUIREMENTS.md`) both hardcoded literal ROOT paths for files that are workstream-scoped. - Worse: `.planning/.gsd-ws-arg` was being deleted at the end of Step 7, but Steps 9 and 10 run AFTER Step 7 and still needed to re-read it — the round-trip mechanism this fix builds was already gone before its two remaining consumers ran. Fixed by moving the `.gsd-ws-arg` cleanup to Step 10 (its true last consumer, after the roadmap commit) and adding the same fetch-then-_gsd_field-extract pattern already used in Step 6 to Steps 9 and 10, resolving `requirements_path`/`roadmap_path`/`state_path` through `init.new-milestone $GSD_WS_ARG` instead of literal paths. Also fixed (MEDIUM, same review pass): Step 1's `.gsd-ws-arg` write had no `2>/dev/null || true`, unlike every other round-trip write in this same file — brought into line with the established idiom. Verified: reproduced the pre-fix bug directly (Step 9/10 fences echoing the literal root paths regardless of --ws), confirmed both fences now resolve the workstream-scoped paths correctly, and confirmed the round-trip file survives Step 7 and is only removed after Step 10. Updated the Step 7 test that previously asserted premature cleanup (inverted to assert the file survives); added new coverage for Steps 9 and 10 in both flat and --ws modes. Two remaining LOW/pre-existing findings from the same review pass, deliberately left as-is: `phase_archive_path` (src/init.cts, untouched by this diff) resolves via the same root-only `getLatestCompletedMilestone` this fix's earlier commit already declined to touch, for the same genuine-product-intent-ambiguity reason (workstream-scoped vs project-pooled "latest completed milestone" is not resolvable from the code alone). `.planning/research/` staying root-scoped in the #222 self-heal prose is consistent with the existing (unchanged) `research_dir` field, not a new inconsistency. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4456): revert wrong config_path change; fix isolated-cwd test env gsd-test caught two real regressions from this fix's earlier commits: 1. config_path is NOT shared like PROJECT.md. The prior commit's grep-and-replace ("fix six more functions with the identical bug") also touched cmdInitExecutePhase's config_path (a fourth call site beyond the two I'd manually checked) — but tests/init.test.cjs's pre-existing, ADR-0006-governed "init handlers honor GSD_WORKSTREAM" coverage explicitly asserts config_path IS workstream-scoped for execute-phase/plan-phase/phase-op/milestone-op. workstream-flag.md's "# Shared" marking for config.json is stale (the same class of staleness already found for milestones/ during the #4455 follow-up); ADR-0006 plus its real, passing tests is the authoritative source. Reverted config_path to the plain workstream-aware planningDir(cwd) in all four functions it was wrongly changed in. 2. Isolating cwd to a tmpDir (needed once Step 1's fence started performing a real .gsd-ws-arg write) broke the runtime-launcher preamble's own gsd-tools.cjs discovery — no git repo at an isolated tmpDir, no global gsd_run on the CI bench's PATH. Fixed by passing RUNTIME_DIR explicitly in every isolated-cwd test's env, matching the preamble's own documented override precedence. Verified: direct CLI invocation confirms execute-phase's config_path is workstream-scoped again under GSD_WORKSTREAM=wsx; the RUNTIME_DIR fix confirmed against a stripped PATH (no global gsd_run), matching the bench condition that surfaced the original failure. Emitted-Drift-Ack-Growth: new-milestone.md — #4456 forwards --ws to every downstream gsd_run call across 7 fences (Steps 1/5/6x3/7/9/10), adding a persisted round-trip file plus resolved-path fetches that replace several hardcoded literal paths Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4456): backfill changeset PR number and correct final scope pr: 0 -> pr: 4545, and removed the changeset's claim that config.json is a shared file -- that was the change this same PR later reverted after gsd-test caught it contradicting ADR-0006's established, workstream-scoped config_path contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4456): baseline the 10 new SC2086 findings from --ws forwarding The lint-tests CI job failed with a hard exit 1. Diagnosis (not assumed): the log's two `fatal: ambiguous argument 'origin/next...HEAD'` git errors (lines 244/248) are a red herring — both belong to lint-removed-but-needed.cjs, which prints its own "could not resolve origin/next, skipping" message and exits gracefully, exactly like the already-handled two-dot-form error from lint-fix-has-regression-tests earlier in the same log. Neither contributes to the actual failure. The real cause is lint-workflow-shellcheck: this fix's new fences append $GSD_WS_ARG unquoted to gsd_run calls (deliberately, so it splits into 0 or 2 argv tokens — the same idiom gsd-core/workflows/verify-work.md already uses for ${GSD_WS} and already has baselined). ShellCheck correctly flags each as SC2086, and lint-workflow-shellcheck.cjs's baseline is a deliberate ratchet (#4109) requiring new findings to be explicitly accepted, not auto-passed. new-milestone.md previously had zero baselined SC2086 findings, so all 10 new (correct, intentional) occurrences were reported as new and failed the gate. Added 10 {file, code, message} entries to scripts/lint-workflow-shellcheck-baseline.json for gsd-core/workflows/new-milestone.md's SC2086 findings, matching the established, already-accepted precedent for the identical pattern in verify-work.md. No source or workflow file changed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0ebc3cf279 |
fix(#4455): PROJECT.md is shared across workstreams, not workstream-scoped (#4543)
* fix: PROJECT.md is shared across workstreams, not workstream-scoped (#4455 follow-up)
Self-discovered regression, found while diagnosing #4456: #4455 (PR #4542,
merged as
|
||
|
|
c6df4e1e46 |
fix(#4455): autonomous.md and complete-milestone.md resolve STATE/ROADMAP/MILESTONES/PROJECT/REQUIREMENTS through the workstream-scoped init fields (#4542)
* fix(#4455): thread workstream-scoped paths through autonomous and complete-milestone workflows autonomous.md and complete-milestone.md read/wrote hardcoded literal `.planning/STATE.md` / `.planning/ROADMAP.md` / `.planning/milestones/...` paths in their shell fences, bypassing workstream scoping entirely. With GSD_WORKSTREAM=alpha set, planningDir(cwd) correctly resolves into workstreams/alpha/, but a literal `cat .planning/STATE.md` still read the ROOT file (or silently returned empty if root state was absent) -- reproduced deterministically in the issue's own repro. Root cause: each workflow step's bash fence is a separate shell invocation, and cmdInitManager/cmdInitCompleteMilestone's JSON payloads never carried resolved state_path/roadmap_path/archive_dir fields for the workflows to extract -- unlike cmdInitPlanPhase, which already does this correctly and is the pattern this fix mirrors. - src/init.cts: cmdInitManager and cmdInitCompleteMilestone now emit state_path/roadmap_path (workstream-scoped via planningDir(cwd), existence-checked, toPosixPath'd, null when absent -- identical to cmdInitPlanPhase's existing contract) and archive_dir (the milestone archive directory, composed the same way milestone.cts's already-correct archive helper does per #1911). - autonomous.md: discover_phases and iterate now extract state_path via the already-fetched INIT_MANAGER payload instead of hardcoding `.planning/STATE.md`; iterate's second, previously-separate hardcoded read is folded into the same fence (no double-fetch); lifecycle step 5b checks the resolved archive_dir instead of a hardcoded milestones path. - complete-milestone.md's reorganize_roadmap_and_delete_originals step (which previously called no init command at all) now fetches init.complete-milestone and uses the resolved roadmap_path/state_path/ archive_dir for the backlog read, the write-guard sentinel's armed content, the Write-tool target for the reorganized ROADMAP.md (the sentinel fence now echoes the resolved path so the executing agent can see it), and the safety-commit --files list. `.planning/MILESTONES.md` and `.planning/PROJECT.md` stay literal root paths -- documented shared files, per the issue's explicit "not a blanket replacement" scope. Regression tests extract and execute the real bash fences (with a stubbed gsd_run) rather than string-matching the markdown, covering flat mode (unaffected), an active workstream (the issue's own repro shape, now correctly resolving), the no-double-fetch requirement, and a dedicated guard locking MILESTONES.md/PROJECT.md as shared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): add changeset for workstream-scoped autonomous/complete-milestone fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): close write-guard gap on workstream-scoped curated paths Isolated security review of the #4455 fix (workstream-scoped STATE/ ROADMAP/milestone-archive path resolution in autonomous.md and complete-milestone.md) flagged that hooks/gsd-write-guard.js's CURATED_PATTERNS only matched root-level .planning/ paths, never .planning/[<project>/]workstreams/<ws>/... — meaning the catastrophic- shrink guard silently never engaged for a workstream-scoped write. This is directly relevant here: the #4455 change makes a workstream- scoped ROADMAP.md Write reachable via complete-milestone.md's own explicit sentinel-hatch instructions, which assume guard protection that did not actually exist for that path shape. Extended CURATED_PATTERNS with the three workstream-scoped equivalents; consumeSentinelFor's own path-derivation logic needed no change since it derives from the actual write target. Verified empirically (a 293->16 line workstream ROADMAP.md shrink now correctly returns exit 2 / decision:"block") and with 5 new regression tests. Also addressed a code-review nit on the core #4455 fix: cmdInitCompleteMilestone called planningDir(cwd) three separate times instead of caching it once. Accepted as-is (not fixed): complete-milestone.md's reorganize_roadmap_and_delete_originals step re-fetches `gsd_run query init.complete-milestone` three times across its fences rather than merging the first two (no state-changing Write between them, unlike autonomous.md's iterate step which does merge). This is an efficiency nit, not a correctness bug — merging risks disrupting the step's prose flow and its existing binding test for a non-functional gain. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): add changeset for the write-guard workstream-scope fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): fix gsd-test-surfaced regressions from workstream-path fix Running gsd-test against the full #4455 diff (including the write-guard security fix and the cmdInitCompleteMilestone caching nit) surfaced four real, non-flaky failures, all direct consequences of editing gsd-core/workflows/autonomous.md and complete-milestone.md: 1. tests/autonomous-converge.test.cjs pinned the OLD hardcoded `STATE_CONTENT=$(cat .planning/STATE.md ...)` read in both discover_phases and iterate. That is exactly the literal-path behavior #4455 fixes, so the test needed updating to assert the new init.manager-resolved `STATE_PATH` read instead (with an explicit doesNotMatch guard against regressing to the old literal). 2. tests/workstream-scoped-paths.test.cjs's own "no-double-fetch" test counted gsd_run invocations via a shell variable incremented inside the stub function — but `INIT_MANAGER=$(gsd_run ...)` runs gsd_run inside the command-substitution SUBSHELL, so that increment never survives back to the parent shell and the counter always read 0. Switched to a file-based call log (one byte appended per call), which survives the subshell boundary. 3. tests/compact-content-partition-guard.test.cjs's disjointness check flagged the reorganize_roadmap_and_delete_originals step's new `INIT_CM=$(gsd_run query init.complete-milestone)` fetch (added 3x, per the accepted-as-is disposition in the prior commit) as byte-identical to a pre-existing, unrelated fetch already present in complete-milestone/detail/elaboration.md's handle_branches section (§2). Same idiom, same conventional variable name, coincidentally colliding across the spine/detail split boundary. Renamed the new step's local variable to INIT_REORG — a distinct, purpose-specific name is arguably better practice anyway for two logically unrelated fetches, and it removes the literal collision honestly rather than restructuring the split. 4. tests/benchmark-compact-content.test.cjs reported real byte-count drift in the committed baseline (autonomous.md and complete-milestone.md both grew from the #4455 content). Refreshed via `node scripts/benchmark-compact-content.cjs --write`. Verified: node scripts/benchmark-compact-content.cjs --check now reports the baseline up to date; a standalone invocation of checkDisjointness() against the real repo state now reports zero violations across all 6 registered splits; manual bash-fence execution of both the autonomous.md iterate fence (call count = 1) and the complete-milestone.md backlog fence (with INIT_REORG) confirms correct behavior. Emitted-Drift-Ack-Growth: autonomous.md — #4455 workstream-scoped STATE.md path resolution replaces hardcoded literal reads Emitted-Drift-Ack-Growth: complete-milestone.md — #4455 workstream-scoped STATE/ROADMAP/archive path resolution replaces hardcoded literal reads Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): MILESTONES.md/PROJECT.md/REQUIREMENTS.md are workstream-scoped too, and so is project-only mode Fresh isolated code-review and security-review passes against the full diff (run after the previous gsd-test-surfaced fixups landed) each found one real, confirmed defect: Code review: the safety-commit `--files` list and the REQUIREMENTS.md `git rm` step both hardcoded `.planning/MILESTONES.md`, `.planning/PROJECT.md`, and `.planning/REQUIREMENTS.md` as literal root paths — but src/milestone.cts's cmdMilestoneComplete writes MILESTONES.md via `planningPaths(cwd).planning` (the workstream base) and PROJECT.md/REQUIREMENTS.md resolve the same way through `planningPaths().project`/`.requirements` (src/planning-workspace.cts). Only `todos` is the documented root-scoped exception (#4256); an earlier version of this fix wrongly generalized that exception to MILESTONES.md/PROJECT.md too, and the now-corrected test previously enshrined that wrong behavior as intended. Under an active workstream, the safety commit would have silently missed the actual files `milestone complete` just wrote, and the git-rm step would have targeted the wrong (root) REQUIREMENTS.md entirely. Fixed by exposing `milestones_path`/`project_path`/`requirements_path` from init.complete-milestone (src/init.cts) and resolving all three through them, the same pattern already used for state_path/roadmap_path/ archive_dir. The four remaining literal MILESTONES.md/PROJECT.md mentions elsewhere in complete-milestone.md (lines ~12-13, ~441, ~607, ~662) are display-only prose in status/summary message templates, not actual file operations — left as-is; they are a cosmetic path-display inaccuracy under an active workstream, not a data-integrity bug like the two fixed here. Security review: confirmed the write-guard fix from the prior commit is correct and complete for workstream scoping, and independently surfaced the same project-only gap the code-review pass above also caught structurally: `CURATED_PATTERNS` had no pattern for `.planning/<project>/...` (GSD_PROJECT set, GSD_WORKSTREAM unset) — planningDir(cwd) supports that shape independently of workstream nesting, so it is reachable, not hypothetical. Fixed by adding three more patterns, verified empirically (a project-scoped 292->16 line ROADMAP.md shrink now correctly returns exit 2 / decision:"block") and with 6 new regression tests. Verified: manual bash-fence execution of the corrected commit-files and requirements-rm fences (both flat mode and GSD_WORKSTREAM=alpha) resolves to the right paths in both cases; a standalone invocation of checkDisjointness() against the real repo state still reports zero violations; the benchmark baseline was refreshed again for the further size change (already covered by the existing Emitted-Drift-Ack-Growth trailer on complete-milestone.md two commits back — that trailer is read over the whole merge-base..HEAD range, not per-commit, so it still applies here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): backfill changeset PR numbers and correct final scope pr: 0 -> pr: 4542 for both fragments, and updated both bodies to reflect the final fix scope (MILESTONES/PROJECT/REQUIREMENTS are workstream-scoped too, not shared-root exceptions; the write-guard fix also covers project-only scoping, not just workstream nesting). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): lifecycle-5b archive-path assertions use the fence's own separator, not path.join PR CI's windows-latest shard 3/3 failed: "expected ls to find the root archive file, got: ...\milestones-root/v1.0-ROADMAP.md". The autonomous.md lifecycle step 5b fence composes the checked path with a literal bash `/` (`"${ARCHIVE_DIR}/v${milestone_version}-ROADMAP.md"`), which on Windows yields a MIXED-separator path — Windows backslashes from archiveDir plus one trailing `/`. My test's assertion used path.join(archiveDir, 'v1.0-ROADMAP.md') instead, which on a Windows Node process produces an all-backslash path that never matches the fence's mixed-separator output. Both assertions in that describe block now mirror the fence's own literal `/` concatenation (`${archiveDir}/v1.0-ROADMAP.md`) instead of path.join — matching the style the other two describe blocks in this same file (safety-commit --files list) already used correctly for the identical archive-dir pattern, so this brings the one outlier into line rather than introducing a new idiom. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): write-guard sentinel comparison now realpath-resolves the token, not just the target PR CI's macos-latest full-test shard 2/3 failed a #4455 test: "the sentinel hatch ... unblocks a workstream ROADMAP.md write" got status 2 (still blocked) instead of 0. Root cause, unrelated to the Windows fix in the previous commit: hooks/gsd-write-guard.js's main flow realpath-resolves the Write TARGET before the curated-pattern match (round 9 Minor 1's symlink-before-match fix, `filePath = fs.realpathSync(filePath)`), but consumeSentinelFor resolved the sentinel TOKEN's absolute path via plain path.resolve() with no realpath step. On macOS, os.tmpdir() resolves through a /var -> /private/var symlink, so a test's cwd (lexically under /var/folders/...) and its realpath'd target (/private/var/folders/...) diverge — an armed, correct sentinel then never matches the realpath'd target string, and the guard stays incorrectly blocked. This is not macOS-specific in principle: ANY cwd sitting under a symlink (a symlinked project checkout, a symlinked worktree) hits the same asymmetry — gsd-test's Linux bench runs never caught it because /tmp there is not a symlink. Fixed by applying the same fs.realpathSync (with the same keep-lexical-on-failure fallback the caller already uses) to the token's resolved path before comparing. The named file is already known to exist at this point (the caller only reaches consumeSentinelFor after successfully reading the target), so realpath is expected to succeed in the legitimate case; a garbage/mismatched token still fails safe (verified — falls back to the lexical path, still mismatches, stays blocked). Verified: reproduced the exact bug locally (macOS) via os.tmpdir() before the fix, confirmed it resolves after; the negative case (sentinel armed for a DIFFERENT file) still correctly blocks; the pre-existing relative-token sentinel tests (predating #4455) still pass; a garbage/non-existent token still fails safe. Added a deterministic, cross-platform regression test using an explicit symlink (skipped on Windows, matching the existing round-9 symlink test's own skip condition) so this class of bug is caught by gsd-test's Linux bench too, not only by a real macOS CI run. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b8c70a33c0 |
fix(#4454): surface skipped explicit --files paths instead of silent success (#4538)
* fix(#4454): surface skipped explicit --files paths instead of silent success cmdCommit's staging loop deliberately skips a --files path that no longer exists on disk when the caller passed --files explicitly (#2014 guards against staging an unwanted deletion for a temporarily-absent file), but recorded nothing about the skip. The final result reported unqualified committed: true, so a caller had no way to distinguish "everything in --files landed" from "some paths were silently excluded because they didn't exist" -- a real file deletion meant to be committed (e.g. phase.complete removing .planning/milestone.lock) would silently never land, with git status still showing it unstaged after a "successful" commit. Tracks skipped paths in a skippedFiles array and surfaces them as skipped_files (matching this result family's existing snake_case precedent, timed_out) in the success result AND the nothing_to_commit result (reachable when every named path was missing), included only when non-empty so the common case's payload shape is unchanged. Also extended to the SECOND nothing_to_commit result (reached when git itself reports "nothing to commit" after the nothingToCommit guard was false -- the residual partial-skip + partialCommitRefused window the surrounding comments already document) for the same consistency. The #2014 staging/deletion guard itself is untouched -- purely a visibility fix, exactly as the issue requested. Regression tests cover: existing-plus-missing file (the issue's own repro shape, with the missing path genuinely tracked-then-deleted so the #2014 assertion is meaningful, not vacuous against an never-tracked path); only-existing files (no skipped_files key at all); only-missing files (nothing_to_commit with skipped_files); default mode unaffected; and multiple missing files reported in order. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: recover ack trailers buried by squash-merge concatenation Discovered while verifying an unrelated PR (#4454): tests/emitted-attribution.test.cjs failed citing pr-branch.md's growth as unacknowledged, even though PR #4537 (#4447) had genuinely acked it via an Emitted-Drift-Ack-Growth commit trailer. Root cause: git's %(trailers:...) placeholder (used by readAckTrailers) only recognizes a trailer block that is the TRUE TERMINAL block of a commit message. GitHub's squash-merge commit body concatenates every constituent commit's subject+body in order, then unconditionally appends its own `---------` separator and Co-authored-by trailers. Confirmed directly against a real squash commit: %(trailers) returns ONLY GitHub's own two Co-authored-by lines -- not even the LAST original commit's own trailer survives, since GitHub's appended suffix breaks the backward scan before it ever reaches past `---------` to any original commit's content, last or not. An ack trailer added on any non-final commit of a PR (the common case -- more commits routinely land after an ack, e.g. a lint fix or a changeset) is therefore silently invisible forever to any future comparison against a base predating that squash. Fix: readAckTrailers now runs a second pass (extractSquashBuriedTrailers) alongside the existing whole-message read. For each commit in range, if its raw body contains GitHub's squash-suffix signal, the pre-suffix text (everything before the LAST such signal -- an earlier bullet's own body may legitimately contain a markdown horizontal rule that looks the same, so anchoring on the first occurrence would truncate too early and miss a later bullet's real ack) is split on squash-bullet (`* <subject>`) boundaries, and each chunk is independently trailer-parsed via `git interpret-trailers --parse` -- the same underlying algorithm as %(trailers:...), but runnable against arbitrary text rather than only a real commit object. This finds a trailer buried in ANY bullet, not just the last one. Scoped tightly to avoid reintroducing the false-positive class %(trailers:...) was originally chosen to prevent (a mid-body MENTION of trailer syntax must stay inert): the sub-chunk pass activates only on commits matching the squash-suffix signal, so an ordinary commit whose body happens to contain markdown bullets is completely unaffected, and each chunk still goes through git's own strict per-chunk terminal-block detection. Verified end-to-end against a real squash commit (recovers the buried trailer) and four adversarial fixtures now pinned as regression tests: ordinary bullet prose with no squash suffix (stays empty); a squash-shaped commit where one bullet's body merely mentions trailer syntax mid-paragraph (stays inert, the "row 32" false-positive class, now verified at per-chunk granularity); and an earlier bullet's own markdown horizontal rule not truncating the scan before a later bullet's real ack (the last-match-not-first-match case an isolated review pass caught during this same fix). This overrides one-concern-per-PR per CLAUDE.md's Defects & Warnings policy -- a genuine defect discovered mid-work is fixed inline, not deferred to a separate issue (spawn_task for this was correctly blocked by the no-defer guard). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4454): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
78013b3b74 |
fix(#4447): classify mixed structural+transient planning commits as a 5th arm (#4537)
* fix(#4447): classify mixed structural+transient planning commits as a 5th arm pr-branch.md's analyze_commits step computed only NON_PLANNING and STRUCTURAL per commit -- never a total planning-file count -- so its four classification arms assumed every planning-only commit was either wholly structural or wholly non-structural. A commit touching both a structural .planning/ path and a transient/other one matched no arm, and the ambiguous prose let an LLM executing the workflow silently drop it, breaking STATE.md's per-commit revision chain in default mode. Adds an explicit PLANNING_COUNT variable and rewrites the four arms into five, each with an exact computable condition. The new "mixed planning commit" arm (structural + transient/other, no code) gets the same treatment mixed code+planning commits already get: INCLUDE, relying on create_pr_branch's existing universal per-commit filter to strip the transient/other paths -- no new filtering logic needed. tests/helpers/pr-branch-filter.cjs's classifyCommit already returned 'include' for this shape (no upper bound on its structural check); the defect was entirely in the workflow's own prose spec, which is what an executing agent actually reads. New tests pin both: classifyCommit's already-correct behavior (tests 49-50), and a failing-first assertion that analyze_commits computes an explicit planning-total signal (test 51, fails against the pre-fix text). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4447): address code-review findings on the mixed-planning arm - Correct the mixed-planning arm's prose: create_pr_branch's universal filter only strips the TRANSIENT_DIRS subset, not the "other" bucket (config.json, intel/, etc.) -- that subset is preserved, not filtered, same as default mode already does for it on any commit. - Fix the "Mixed planning commits" display line to use the same mode-conditional bracket form as "Structural planning commits" -- it was hardcoding "included" even though the arm is EXCLUDE in strict mode, which would have misled a strict-mode user. - Tighten test 51's regex from unanchored /PLANNING_COUNT=/ to /^PLANNING_COUNT=\$\(/m so it requires the real shell-assignment shape, not just the substring appearing anywhere in prose. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: pr-branch.md — growth is this PR's own #4447 fix (5th classification arm with explicit computable conditions), not incidental drift * fix(#4447): bound test 51's regex quantifier (local/no-unbounded-quantifier) lint:ci flagged the unbounded [\s\S]*? over readFileSync content as a catastrophic-backtracking risk (CWE-1333 class). Bounded to {0,20000}, comfortably larger than the analyze_commits step's actual size. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4447): add changeset for the pr-branch mixed-planning classification fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: eliminate SIGPIPE race in gsd-validate-commit.sh subject/config extraction Discovered while verifying an unrelated PR (#4447): tests/hooks-opt-in.test.cjs's "a git-GENERATED subject is never measured against the supplied message (round 7)" test intermittently got r.status===141 instead of the expected 2 for the --fixup=HEAD case, on a run where the identical code had passed cleanly moments earlier -- confirming a timing race, not a deterministic bug in the test's own assertions. Root cause: gsd-validate-commit.sh runs under `set -euo pipefail` and extracted the commit subject via `SUBJECT=$(echo "$MSG" | head -1)` (two call sites) and the opt-in ENABLED flag via `$(printf '%s\n' "$CONFIG_OUT" | head -1)`. `head -1` closes its read end as soon as it has one line; a real commit message or multi-command-type CONFIG_OUT is multi-line, so the writer can receive SIGPIPE (exit 128+13=141) if its write lands after that close. Under pipefail this is NOT suppressed -- it aborts the whole hook instead of the intended exit-2 rejection. Fix: replace both patterns with pure bash parameter expansion (`${VAR%%$'\n'*}`) -- zero subprocesses, zero pipe/race surface, and behaviorally identical to `head -1` for single-line, multi-line, and trailing-newline input (verified directly). The third similar pipe (`tail -n +2` feeding a `while read` loop that drains to EOF) is a different, race-free shape and was left alone. Regression test is a static, by-construction assertion (per this repo's policy against forcing scheduling races to reproduce deterministically): the vulnerable pipe patterns must be absent from the shipped script, and the parameter-expansion forms must be present. This overrides one-concern-per-PR per CLAUDE.md's Defects & Warnings policy -- a genuine defect discovered mid-work is fixed inline, not deferred to a separate issue. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: add changeset for the gsd-validate-commit.sh SIGPIPE race fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: harden SIGPIPE-race regression test against reformatted reintroduction Code review found the test's original exact-string regexes would miss a cosmetically-reworded reintroduction of the same dangerous head-1 pipe (extra whitespace, an appended 2>/dev/null). Broadened to content-tolerant but still $(...)-wrapped regexes (bounded quantifiers per local/no-unbounded-quantifier) -- verified against both the current file (no false positive, including the fix's own explanatory comments that quote the bare unwrapped pattern in prose) and a synthetic reformatted reintroduction (correctly caught). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4447): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e03921c7d8 | enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) | ||
|
|
18c899def5 |
enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
733bec3ad1 |
enhance(#4261): report size-cap headroom on every run, with a reserved margin (#4418)
* enhance(#4261): report size-cap headroom on every run, with a reserved margin The tier caps are red lines and none of them moves here. What was missing is everything below the red line: a passing run said nothing, so a contributor at 99.6% of a cap and one at 60% got identical feedback, and the density that produces merge-time collisions was invisible to the people creating it. Two levels, matching the shape execute-phase.md already carries by hand (a hard ceiling plus a lower margin "so minor future edits don't re-trip the gate") and which was until now the only capped file with one: 1. a headroom census printed every run, green included, sorted least-headroom-first, and appended to the GitHub job summary 2. a 95% reserved margin that names the files inside it and REPORTS rather than fails The margin deliberately does not fail. A cap breach is a red line; a file at 96% is not broken, it is a file whose next contributor should extract before adding. Failing there would create a second red line and force exactly the +N bumps the policy forbids. Neither level asserts a count, so this adds no snapshot to regenerate — the per-file size baseline was deleted by #2724 for conflicting on 7 of 7 PRs that touched it. Also deletes rather than refreshes the per-tier high-water comments in both guard files. They were measured once and then diverged from the tree: the LARGE line still claimed "gsd-executor 42,342 -> ~6.8 KB headroom" while the real high-water sat at 99.6% of that cap, so the comment documenting the margin was itself why nobody noticed the margin was gone. As measured on next by the new census, the pressure has grown since the issue was filed: gsd-plan-checker.md has 9 bytes of headroom, gsd-verifier.md 21, and plan-phase.md 14. * chore(#4261): add changeset for the size-cap headroom census * test(#4261): exercise reserved-margin boundaries --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
95529ef153 |
fix(#4291): isolate ship:pre and empty-points e2e CLI tests from real HOME (#4321)
* fix(#4291): isolate ship:pre and empty-points e2e CLI tests from real HOME tests/loop-hooks-ship-pre-e2e.test.cjs's runTools() and tests/loop-hooks-empty-points-e2e.test.cjs's spawnGsd() spawned gsd-tools without sandboxing HOME, so a capability genuinely installed on the host (e.g. beads, markdown-linting) leaked into the resolved registry and inflated their exact-count/empty-array assertions. Mirrors the installSpawnEnv() fix already applied to loop-hooks-verify-post-e2e.test.cjs in #4204/#4293. * fix(#4291): credit GSD_HOME blanking in spawnGsd docblock installSpawnEnv() blanks GSD_HOME as well as sandboxing HOME, and capability-loader's overlayRoots checks GSD_HOME before falling back to os.homedir() — sandboxing HOME alone would leave that env var reaching the real machine. The sibling loop-hooks-ship-pre-e2e.test.cjs docblock already states both halves; this one credited only HOME. Per trek-e's review on open-gsd/gsd-core#4321 (Nit, isolated adversarial pass). --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
423f38e655 |
fix(#4444): honor config-set --dry-run instead of silently ignoring it (#4504)
* test(#4444): failing-first regression coverage for config-set --dry-run config-set --dry-run is currently parsed nowhere -- routeConfigSet (gsd-core/bin/gsd-tools.cjs) never checks args for it, and cmdConfigSet has no dry-run parameter, so the flag is silently swallowed and the command always writes for real. Reproduces the issue's own repro (sequential --dry-run calls where the second's previousValue proves the first persisted), plus coverage for validation-still-runs, secret-masking, and the sibling unset (config-set <key> null) branch, which has the identical defect. This commit adds the regression coverage only; the fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4444): honor config-set --dry-run instead of silently ignoring it routeConfigSet (gsd-core/bin/gsd-tools.cjs) never read args for --dry-run, and cmdConfigSet had no dry-run parameter at all -- so the flag was silently accepted (as any unrecognized trailing argument is) and the command always wrote for real. A second "dry run" then showed previousValue reflecting the first one, proving it had persisted. Threads a dryRun option through cmdConfigSet, gating BOTH mutating branches: the null/unset path (unsetConfigValue) and the real-set path (setConfigValue) -- the unset branch had the identical defect, undiscovered until auditing every mutation site while designing this fix. Each gains a previewConfigValue/previewUnsetConfigValue counterpart that reuses the real function's exact traversal/creation logic (_setNestedValue/_unsetNestedValue) on a throwaway in-memory config copy that is never written -- so the preview can never diverge from what the real write would compute. All validation (unknown key, enum/number/boolean checks, secret masking) runs identically whether or not --dry-run is passed; only the final write is skipped, replaced with a `{ dry_run: true, would_update / would_unset: true, ... }` preview payload matching the precedent established by `milestone complete --dry-run` (#2118) and `todo complete --dry-run` (#4096/#4325). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(#4444): extract loadConfigJson to stop a 5th copy-paste of the same load/parse block Code review flagged that setConfigValue, unsetConfigValue, setConfigValues, and the two new preview functions each repeated the identical "load .planning/config.json, JSON.parse, catch -> CONFIG_PARSE_FAILED" block -- exactly CLAUDE.md's own "Generative Fix Divergence" known-defect pattern. Extracted a single loadConfigJson(cwd) helper; behavior is unchanged (verified: build, tsc, and the dry-run/real-write smoke test all pass byte-identical to before). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4444): changeset for the config-set --dry-run fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4444): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom install-minimal-hooks.test.cjs (weight=24.45, the heaviest file in the suite) sits alone in its own chunk yet still occasionally brushed the 600000ms per-chunk ceiling on Windows -- observed on PR #4504's first CI run for this change (passed clean on rerun, consistent with the "legitimately too slow for the budget" cause the chunk-timeout diagnostic already names, not a leaked handle). Raised RUN_TESTS_CHUNK_TIMEOUT_MS's default from 600000ms to 800000ms: ~33% more margin, still comfortably below the 900000ms regen:derived fixture timeout that fragment-single-edit-propagation.install.test.cjs deliberately keeps ABOVE the chunk ceiling, and far under the 45-minute job cap -- Windows shards currently finish in ~19-20 minutes total, so there is ample headroom. Updated every dependent mirror/assertion in lockstep (tests/helpers/emitted-runtime.cjs's duplicated CHUNK_TIMEOUT_CEILING_MS constant, its lock test in tests/emitted-attribution.test.cjs, the Windows-skip prose in fragment-single-edit-propagation.install.test.cjs, and docs/TESTING-SUITES.md's reference table) so nothing describes a stale value. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Revert "fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom" This reverts commit 394aadaf6f5af6fd700bf0f444c9fbd686285a4f. * test(#4444): consolidate redundant installer spawns in install-minimal-hooks.test.cjs This file's real, unrelated pre-existing cost (dated 2026-09-06, PR #4428) is what tipped a Windows CI shard over the per-chunk timeout backstop on PR #4504 (issue #4444's own diff never touches this file or the installer). Rather than raise the timeout, cut the file's actual spawn count: several describe blocks independently re-installed the IDENTICAL runtime/scope/flag configuration just to assert different things about the same install output. Merged each such group onto a single shared install, with every original assertion preserved: - --help x3 -> x1 - the three per-runtime/scope --minimal E2E loops (global, local, and on-disk-matches-manifest) merged into one loop over SKILL_RUNTIMES x [global, local]: 44 spawns -> 22 - the --minimal manifest-mode/backcompat triple-install -> one shared, memoized install via sharedMinimalManifestInstall() - .sh hooks existence checks (5 tests) -> 1, executable-bit check (its own Windows-conditional skip) left separate - Codex #4087 hook-helper tests (3) -> 1 - Windsurf #4087 hook-helper tests (2) -> 1 - pi shared-hooks-bundle tests (3 per scope) -> 1 per scope Net: ~65 real installer spawns in this file down to ~29, no assertion dropped or weakened. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93e141a006 |
enhance(#4139): Phase 4 — measure the window instead of asserting it (#4502)
* enhance(#4404): add offline token benchmark for compact-content splits ADR-4139 Decision 2 requires the finite-attention justification for workflow.compact_content to be measured, not asserted. `npm run benchmark:compact-content` computes, per registered spine/detail split discovered under gsd-core/workflows/, the token count with the split active (spine alone) vs inactive (spine + all detail parts read back in), using gpt-tokenizer (pinned exact devDependency — Anthropic publishes no tokenizer for Claude 3+, so every output surface labels this a PROXY-TOKENIZER comparison: the on/off delta is exact under one tokenizer applied identically to both sides, the absolute counts are not Claude's real ones). Reporting-only by design and verified so: --check diffs the live recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json) and prints drift, but never exits non-zero for a drifted or missing baseline — the only thing allowed to fail this script is a genuine I/O error reading a source .md file it's measuring. Not wired into lint:ci or pretest. Discovery is deliberately reimplemented rather than importing tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a scripts/ reporting tool from depending on a test-only module. tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to prove the benchmark makes no network call, monkeypatching http/https/ net/dns/fetch to throw rather than relying on sandboxing. docs/CONFIGURATION.md documents the new benchmark against the workflow.compact_content key to satisfy this repo's docs-required gate for an Added-type changeset. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4404): address orthogonal review findings on the token benchmark Standards axis found two hard violations against documented rules: - CLAUDE.md's Generative Fix Divergence rule requires a parity assertion for shared discovery logic maintained in two places. Added a test comparing benchmark-compact-content.cjs's own discoverRegisteredSplits against tests/helpers/compact-content-split.cjs's version on the real repo tree, so the two can never silently drift apart. - The changeset body closed its bold span with a period and continued as a second sentence, instead of the canonical `**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents. Spec axis found the "network disabled + identical output across two runs" Done-when criterion was verified as two separate properties (determinism tested without network denial, offline survival tested as a single run) rather than as one combined property. Added a test that runs the benchmark twice under the deny-network preload and asserts byte-identical stdout. Security axis found tests/fixtures/deny-network.cjs didn't patch dns.promises (a separate binding from the callback dns API), tls.connect, or http2.connect — inert today since nothing in the benchmark calls them, but a silent gap in what the preload's own header claims to guarantee. Patched all three. CLAUDE.md's Property-Based Testing rule also requires a fast-check test for budget-limit arithmetic; added one for computeAggregate's off/on summation (true sum over N splits, never NaN/Infinity, never exceeds 100% when off >= on for every split). Standards axis's remaining two findings (a Data Clumps observation on the {offTokens, onTokens, reductionPct} triple, and mild duplication in formatDriftReport's three line-formatters) are left as judgement calls: introducing a named type for a 3-field local tuple, or a formatter abstraction for three short lines, would be exactly the premature abstraction CLAUDE.md's engineering guidance warns against for a script this size. All changes verified directly (parity logic, the fast-check property, and the three newly-denied network surfaces actually throwing under the preload) via node -e before committing; full npm run lint:ci passes with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4404): backfill changeset pr number to 4502 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8dcdcb253e |
fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema (#4501)
* test(#4443): failing-first regression coverage for hooks.commit_types config key isValidConfigKey('hooks.commit_types') currently returns false and config-set hooks.commit_types rejects with "Unknown config key", because the key was never added to config-schema.manifest.json's validKeys when it shipped (#3811/#4340, 1.13.0) despite being documented (docs/COMMANDS.md) and consumed by hooks/gsd-validate-commit.sh. This commit adds the regression coverage only; the manifest fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): avoid false-positive docs-guard registration trip The assert message for the new hooks.commit_types test mentioned "docs/COMMANDS.md" literally, which happened to land between two unrelated pre-existing backticks and tripped lint-docs-guard-registration.cjs's template-literal co-occurrence detector (a known, documented false-positive shape for that lint). Rephrased to drop the literal docs/ path from the message; the test's intent (documenting why the key must be valid) is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema config-schema.manifest.json's validKeys never got hooks.commit_types added when the feature shipped (#3811/#4340, 1.13.0) despite it being documented (docs/COMMANDS.md) and consumed by hooks/gsd-validate-commit.sh -- so config-set hooks.commit_types rejected with "Unknown config key", and the only way to configure a documented feature was hand-editing .planning/config.json. While auditing every hooks.* key actually read by shipped code against validKeys (CLAUDE.md's no-deferrals rule: a defect found anywhere in the tree while working an issue is fixed in the current change, not filed separately), hooks.community -- gsd-validate-commit.sh's own opt-in gate -- turned out to have the exact same gap. Both are added here; an audit of every hooks.* read site confirmed these are the only two missing entries. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): e2e coverage for hooks.community + changeset Closes the coverage-rigor gap the Standards review flagged: hooks.community had only a unit-level isValidConfigKey assertion, not the same real config-set CLI round-trip hooks.commit_types already got. Also adds the changeset fragment the same review flagged as a missing hard requirement. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): use PROBE_TIMEOUT_MS instead of a bare 15000 literal local/no-adhoc-timeout-literal (lint:ci) correctly flagged both new spawnSync calls' bare timeout: 15000 -- this call class (a short CLI probe against a temp fixture) is exactly what tests/helpers/timeouts.cjs's PROBE_TIMEOUT_MS documents. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4443): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a0f8f956c4 |
enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior" section lists four checks; the ADR's Decision 5 and its own phase table ("partition rules + the five checks") list five — the same four plus "boundary moves are declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR implements all five. docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule itself, the protected-content list and <!-- gsd:protected --> sentinel syntax (relocated unchanged from gsd-core/references/compact-content-protected-content.md, now deleted — it was never referenced by any runtime workflow Read, only by the predecessor test as documentation, so nothing at runtime regresses, and removing it from gsd-core/references/ also drops it from all 19 installed-project shipped-content trees for a file nothing ever read), and the five checks explained for a human reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped content". tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no registry file, a pair is registered by existing on disk), line normalization (carries forward Phase 2's bare-label-line isTrivial fix and the canonical gsd_run-launcher-preamble exclusion), sentinel extraction, and a Boundary-Move-Declared commit-trailer reader that is a direct structural port of tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader (ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range, same dedupe/conflict rules. tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3 (registration + size cap) run unconditionally against every registered split. Checks 1 (completeness, fires once per split on the PR that introduces a new detail/ path), 4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5 (boundary moves declared) are PR-diff-scoped against the resolved base ref and skip cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a deliberate asymmetry from check 5's trailer reader, which must throw rather than silently pass when ITS range is uncomputable, since that function is answering "did this PR declare its moves" rather than "is there even a diff to look at". Each of the five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair, built against synthetic temp files or real throwaway git repos, per this repo's rule that a guard nobody has seen go red is not yet a guard. Building the real fixtures caught and fixed one real bug before it shipped: check 4's line-presence test was using the trivial-line-filtered normalizer, so a byte-identical spine falsely reported its own protected code-fence line as "deleted" — fixed with a non-filtering membership check. Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md is a fourth workflow sub-file kind that already existed on disk and was already known to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same pattern, same PR, rather than deferred. Also, mechanically required by the new fourth sub-file kind: - scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS alongside modes/steps/templates — a detail/<part>.md inherits its parent's response_language coverage through the same per-file proof, not a parallel one. - tests/workflow-size-budget.test.cjs: explicit regression test locking that detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers — true by construction (measureWorkflows/listWorkflowStems don't recurse), made explicit per the issue's own Done-when item rather than left true-by-omission. - scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates` NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked); docs/INVENTORY.md's table updated to four kinds. Verified: `npm run lint:ci` clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): review findings + a real gsd-test failure in the new guard Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a separate security review ran against the prior commit. Fixed everything each surfaced: - Security (Low, path-traversal existence oracle): checkRegistration's dangling-reference check extracted detail-path-shaped substrings from spine PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot and probed fs.existsSync with no containment check — a spine file containing `../../../etc/detail/passwd.md`-shaped text could make the guard test file existence outside the repo. Added a path.relative-based containment check before the fs.existsSync call; anything that resolves outside repoRoot is now reported as a dangling reference directly, never probed on disk. - Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case CLAUDE.md's TEST RULES require (limit-1/limit/limit+1). - Standards (Property-Based Testing): extractProtectedBlocks (a sentinel parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict parser, bijective-shaped) had no fast-check property test. Added three: a render/parse bijectivity property for the trailer parser (mirroring the exact ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for the same parser, and a well-formed-sentinel-round-trips property for extractProtectedBlocks. Then dispatched gsd-test on the resulting commit. It found a real bug the reviews couldn't have caught (none of them can run inside gsd-test's sandbox): checks 4/5's real-repo assertion failed against plan-phase's own split, reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" — content Phase 2 (#4402) legitimately moved into detail/elaboration.md months before this PR's Boundary-Move-Declared mechanism existed to require a trailer for it. Root cause: `resolveBase()`'s own doc comment already documents that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox container, and its fallback candidate (a bare `next` branch) can resolve to a point in history that predates an already-merged, already-reviewed split — making that split look "newly introduced" from the sandbox's vantage point. Check 1 (completeness) already scopes itself correctly to only genuinely-new detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping, so a stale base made them re-litigate a settled split retroactively. Fixed by having checks 4/5 skip any split name check 1 already counted as newly-split — their own premise ("did an EXISTING split shed/undeclare something") does not apply to a split that is, from the resolved base's vantage point, brand new; that is check 1's domain alone. Verified locally (25/25 tests pass via a direct `node -e` require, since `node --test` is blocked in this repo) and via re-reasoning through the exact real-repo scenario the gsd-test failure showed. Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the prior commit's deletion of gsd-core/references/compact-content-protected-content.md was never reflected there, which is what golden-install-tree.test.cjs's other 19 failures in the same gsd-test run were. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4403): backfill changeset pr number to 4497 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs (weight 17.87, by far the chunk's dominant cost) packed alongside 39 other files. Traced, not assumed: - scripts/run-tests.cjs's own timeout-headroom comment for the OUTER per-shard timeout documents that "adding one test file reshuffled 115 of 268 unit files between shards" — shard/chunk composition is architecturally known to be unstable to single-file additions, which is exactly what this PR's own new tests/compact-content-partition-guard.test.cjs is. - A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own CI), already documents the SAME chunk hitting the SAME 600s backstop with the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of 17.87 — not a stale-table miss") dominating it — the fix then was cutting the Windows per-chunk budget from 60 to 40. That cut clearly was not enough: two documented incidents in two days, at two different budget settings, both centered on one file that alone consumes ~45% of even the reduced Windows budget. - tests/test-timings.json's own header confirms its source data (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the recorded figure" — the packer's weight-balancing is working off data that is both stale (table last regenerated 2026-08-07) and known to underestimate the platform where the failure occurs. Given codex-config.test.cjs is disproportionately heavy AND every companion sharing its chunk is decided by a packing algorithm already documented as reshuffling unpredictably on any new file, tuning the shared budget a third time only moves the marginal line to wherever the next new file happens to land — it does not remove the gamble. Isolating codex-config.test.cjs into its own dedicated single-file chunk, unconditionally and on every platform, removes it at the source: the file never enters the pool packChunks balances, so no other file's packing changes, and no future single-file addition (mine or anyone else's) can silently reintroduce this exact failure by landing in its chunk. Extracted as a small pure function, partitionIsolatedFiles (mirroring this file's existing pattern of pulling packing/analysis logic out of main() for in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents), with 6 new tests in tests/run-tests-harness.test.cjs covering basename matching across path separators, near-miss non-matches, the empty-list case, and the isolated-set contents. Root cause is now closed rather than papered over with a retry: this failure is a property of one specific heavy file's chunk placement, not something that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped up, this isolation can be revisited — this is a packing-side mitigation for a known file's cost, not a claim the cost is irreducible. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ac6ed6201d |
fix(#4257): harvest only prose phase references; W002 names its workstream scope (#4486)
* test(#4257): W002 harvest precision + workstream-scoped warning regression rows Tests-only RED commit: A-rows pin the command-mention/code-span harvest precision on statePhaseTokens, B-rows drive W002 under root and workstream scope (scope clause asserted, root grammar byte-identical), C-rows pin the additive workstream snapshot field. All fail on next; fix follows. * fix(#4257): harvest only prose phase references; W002 names its workstream scope Sub-defect (a): the statePhaseTokens harvest was the verbatim #3309 relocation of verify.cts's unanchored, markdown-blind scan ([Pp]hase\s+(TOKEN) over the raw file), so GSD's own command names (/gsd-execute-phase 5, bare or quoted) and any token inside a code span/fenced block were harvested as phase references and fired W002 on ledger rows. Now strips fenced blocks then inline spans via the canonical markdown-sectionizer seam (#2365 composition order) and matches with a left word boundary (?<![-\w]) so hyphen- or word-suffixed carriers are mentions, not references. Pinned tradeoff: a genuine reference written in backticks stops counting (a quoted literal is not a reference). Sub-defect (b): the valid set is workstream-scoped by construction (planningPaths under GSD_WORKSTREAM; per-workstream numbering is deliberate), but the message claimed 'only phases 1, 2 are declared' unqualified. New additive PlanningSnapshot.workstream field, sourced from planning-workspace's new resolveEnvWorkstream() — the ONE env discriminator planningDir itself applies — so the clause cannot disagree with the base the reads used. Root scope keeps the byte-identical message; the checker's scope is unchanged. * test(#4257): close the B2 quoted-literal code span (fixture typo) The B2 fixture wrote a single opening backtick — an unterminated span is literal text per CommonMark, so its content is prose and W002 correctly fired on it. The test's name, the A3 snapshot-level twin, and the B2 matrix row all intend a closed span; pre-fix this was indistinguishable because the unanchored harvest fired either way. * chore(#4257): changeset fragment (pr number to backfill) * chore(#4257): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
c3a18b5ba0 |
docs(#4440): stop telling agents to grep .env files the secret guard denies (#4500)
* docs(#4440): stop telling agents to grep .env files the secret guard denies verification-patterns.md's <environment_config> and user-setup.md's three per-service Verification examples documented reading .env/.env.local directly via grep. Every covered runtime's secret-read guard denies that (Claude Code deny-rules since #768/v1.4.0; the always-on gsd-secret-read-guard hook since #4236/#4221 in 1.13.0) -- verified by piping each documented command through the shipped hook. verification-patterns.md now checks the environment (printenv) instead of the file, with a case statement replacing a broken grep -v alternation (grep's BRE `|` is literal, so the old placeholder filter matched nothing -- PLACEHOLDER/TODO_fill values passed the "substantive" check as real). Verified under sh (dash) against real/placeholder/empty/ unset values. Existence check ([ -f ".env" ] || [ -f ".env.local" ]) is untouched -- it was never denied. user-setup.md's three grep <SERVICE> .env.local lines are removed outright rather than swapped for printenv: those examples describe a Next.js shape where the framework loads .env.local at runtime without exporting it to the shell, so a printenv substitute would wrongly report "not set" on a correctly configured project. Each block's existing service-level check (build/webhook/connection/email test) already verifies the setup. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4440): changeset for the secret-guard verification-examples fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4440): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e38d246015 |
fix(#4439): render pending-todo 'created' as date-only, matching [date] (#4494)
* test(#4439): failing-first regression coverage for pending-todo date rendering renderPendingTodoBullet currently echoes the todo's `created` frontmatter verbatim into the rendered bullet's [date] bracket, so a full ISO-8601 timestamp (the actual stored shape, per add-todo.md's create_file step) leaks through instead of the date-only format documented in docs/reference/state-md.md and docs/COMMANDS.md. This commit adds the regression coverage only; the renderer fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4439): render pending-todo 'created' as date-only, matching [date] renderPendingTodoBullet echoed the todo's created frontmatter verbatim into the rendered STATE.md bullet. That value is always a full ISO-8601 timestamp by design (add-todo.md's create_file step writes init.todos' timestamp field), but docs/reference/state-md.md and docs/COMMANDS.md document the bullet as `- [date] ...` — a short calendar date. Added pendingTodoDateOnly(), a pure display-only formatter: a value starting with a well-formed YYYY-MM-DD is shortened to just that; anything else ('unknown', a malformed string) passes through unchanged. The stored frontmatter and the JSON todos[].created field are untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4439): cover a non-4-digit-year near-miss for the date-only formatter Standards review flagged a gap: the malformed-date coverage exercised a non-padded month/day but not a non-4-digit year, the other way the input can look almost-but-not-quite like YYYY-MM-DD. Same pass-through code path as the existing malformed-date test; no behavior change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4439): changeset for the pending-todo date-only bullet fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4439): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8b7a0b696b |
enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot split for plan-phase.md, the largest of the 58 eagerly-@-included workflow files (98,290 bytes): the spine keeps every happy-path step, every protected-content block (planner/checker prompt templates, quality gates, the failing-direction few-shot example, the two ScheduleWakeup guardrail paragraphs — each marked with a <!-- gsd:protected --> sentinel), and condensed one-paragraph summaries of five rare/opt-in fallback paths (planner and checker filesystem-hang recovery, phase-split recommendation, source-audit gaps, the thinking-partner conditional, and plan bounce). The full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md (9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only when workflow.compact_content is false (the default) — the exact same resolution rule now stated once in the new shared gsd-core/references/compact-content-gate.md, which every future split references instead of restating. Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to this one split — Phase 3/#4403 owns the generalized guard): the union of spine + detail contains every non-trivial line the parent commit carried (0 missing), no non-trivial line is duplicated between them (0 duplicated), and every declared protected block is well-formed and non-empty. The spine shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a project that has NOT opted in. Verified live, end to end, twice, against this actual repo (not a synthetic fixture) — real gsd-planner and gsd-plan-checker subagent spawns, real PLAN.md output: - workflow.compact_content=false: planned a real disposable phase (a docs/how-to page for enabling the key itself); planner returned PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks against real repo state confirmed. - workflow.compact_content=true (detail.md never read): planned a second real disposable phase; planner returned PLANNING COMPLETE with frontmatter.validate and verify.plan-structure both clean, again fully grounded against real repo state. The five condensed fallback sections were independently re-read spine-only and confirmed sufficient to act on correctly without detail.md's elaboration. Also drafts gsd-core/references/compact-content-protected-content.md — the protected-content category list and <!-- gsd:protected --> sentinel syntax ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged once it lands there. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory Two independent review sub-agents (Standards and Spec axes of /code-review) caught the same structural defect: ADR-4139 Decision 6 mandates gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually skippable"), and this PR had shipped a flat plan-phase/detail.md instead, copying issue #4402's own (inconsistent) restatement rather than the locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and updating every cross-reference (the spine's step 0.5 gate pointer, the shared compact-content-gate.md's own resolution-rule wording, and the completeness test's path constants). Also, from the same review pass: - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content rows said "nothing branches on it yet" — no longer true now that plan-phase.md's spine does. Updated both to name plan-phase as the pilot and note the rest of the corpus is still pending. - Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures (npm run gen:install-tree) — the three new shipped files were missing from the installer emitted-tree goldens. - Found via a cache-busted `eslint . --max-warnings 0` (this repo's eslint --cache has produced false-greens before): the split test's `git show` call had a bare `timeout: 10000` literal, tripping local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS constant from tests/helpers/timeouts.cjs instead of a second guessed copy of the same class of timeout. Verified NOT needed, by tracing the actual mechanism rather than asserting (tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes every gsd-core/{workflows,references}/** path to itself as an identity source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before ever reaching the ack-lookup branch — there is no unattributed delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so the growth ratchet has nothing to ack either. Re-verified after these changes: the completeness/disjointness self-check (0 missing, 0 duplicated) still holds against the relocated detail file, and a full `npm run lint:ci` passes clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore literal content the pre-existing drift guards pin on The first gsd-test run against this split (19 failures) surfaced real regressions: several pre-existing structural guards pin the EXACT text of the sections this split condensed, and paraphrasing broke them. - tests/plan-phase-drift-guard.test.cjs expects the literal `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside plan-phase.md itself, not a prose description of the same check. Restored the exact line into both §9a and §11a's spine summaries. - tests/thinking-partner.test.cjs expects plan-phase.md to literally offer "No, I'll decide" as the skip option. Restored that exact phrase into the condensed thinking-partner paragraph. - Both restores would have duplicated the same text into plan-phase/detail/elaboration.md (which still carries the full elaboration). Removed the now-redundant restatements from the detail file instead of leaving them duplicated — the spine already computes DISK_PLANS before the detail elaboration is ever read, so the detail file references it rather than recomputing it. - Re-running scripts/sync-runtime-launcher.cjs after that edit found the canonical gsd_run preamble had also become an unintentional spine/detail duplicate (both files call gsd_run and each is required, by runtime-launcher-parity's own contract, to carry its own copy). That's sanctioned duplication under a DIFFERENT contract, not lost/copy-pasted content, so tests/plan-phase-compact-split.test.cjs now excludes it from the disjointness check the same way it already excludes trivial fences/headings. - Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's own isTrivial(): a blanket `line.length <= 15` cutoff silently swallowed real content (e.g. the 14-char `<quality_gate>` sentinel). Replaced it with a specific bare-label-line pattern (`Options:`, `Display banner:` etc.) — verified 0 missing / 0 duplicated against the actual split, an improvement over both the original cutoff and a naive full removal (which produces false-positive "duplicates" on generic recurring labels). - gsd-core/references/planning-config.md's own workflow.compact_content row used `/gsd-plan-phase` (hyphen). That file is Claude-facing source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs requires in colon form; docs/CONFIGURATION.md's use of the hyphen form is correct as-is since docs/ is human-facing and outside that test's scanned directories. Fixed to `/gsd:plan-phase`. - tests/plan-phase-compact-split.test.cjs's own `git show` of the parent commit failed inside the gsd-test sandbox ("detected dubious ownership") because the checkout is mounted under a UID the invoking user doesn't own. Scoped `-c safe.directory=<repo-root>` to that one git invocation rather than touching global git config. - docs/INVENTORY.md still had one outstanding "detail.md part" wording fix from the earlier adversarial-review pass, staged now. Re-verified locally against the exact assertions in all four affected test files (all pass) before dispatching a fresh gsd-test run — no change here should have broken any of the other 18 gates; `npm run lint` is clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore the full marker enumeration to §9a's spine trigger line The isolated Spec-axis review flagged that §9a's "Triggered when" line was condensed to "Agent() returns but the return contains no recognized marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` / `## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine- parsed structural headings" category compact-content-protected-content.md lists as protected. The load-bearing use of that same list (the gsd_stall_watch call and the Handle Planner Return bullets a few lines above) was never touched — only this one descriptive restatement was genericized — but leaving any instance of a protected category unsentineled is the silent erosion ADR-4139 Decision 4(c) warns sufficiency isn't machine-checkable enough to catch on its own. Restored the full enumeration into the spine. That reintroduced an exact duplicate into plan-phase/detail/elaboration.md, which still stated the same trigger sentence verbatim. Reworded the detail file's version to reference the spine's trigger condition instead of restating it, since the spine is now the single place that sentence lives in full — mirroring the DISK_PLANS/"already computed above" pattern from the previous commit. Re-verified locally: completeness/disjointness (0 missing, 0 duplicated) and all previously-fixed literal-content assertions still hold. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4402): backfill changeset pr number to 4471 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93abf3a7b6 |
fix(#4276): read the quoting, not just the digits, before eating an IO number (#4420)
POSIX recognizes an IO number only when the digit run is unquoted, but tokenize()'s redirection branch tested `cur` — the token's characters — and never `curMask`, their quoting. `"2"` and `2` were indistinguishable to it, so a correct `gsd_run query commit ... --files "2">out` had its value consumed as an IO number and scored as an unscoped invocation: the #2269 guard reddening on documentation that was right. The mask was already maintained in the same loop and thrown away on this branch. One conjunct reads it: '0' marks a bare character, so /^0+$/ asks precisely whether every character of the digit run was unquoted. The regression arms come in both directions. The quoted rows (glued, detached, single-quoted, and the partially-quoted 2"3" that a some-character-unquoted test would get wrong) must be scoped; the existing unquoted `--files 2>&1` row is the negative control and must stay unscoped, so an implementation that simply stopped consuming IO numbers altogether fails instead of passing. No live instance exists in any of the six scan roots, so this is latent rather than urgent — a false positive, never a silent miss. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
476394689a |
fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix The new suite executes the shipped supplied-root-pin guard against real git fixtures (drifted primary-checkout cwd halts before the write and the FATAL names both roots; matching cwd permits it; unexpanded/empty pins halt; normalization forms; submodule and sibling boundaries; metacharacter quoting; drive-letter form gate) and locks the dispatch contract across execute-phase.md, its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772 per-plan serialization assertion retargets to the fragment that now carries those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring. * fix(#4254): pin sequential executor to the orchestrator's validated root Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its own cwd; every existing guard is worktree-mode-only or self-referential, so an executor spawned with a drifted cwd committed onto the wrong checkout silently. - worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard, composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT (git-vs-git comparison on both sides — representation-safe on Windows, the #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule allowance, warn-and-proceed only when the dispatch carries no pin block. - execute-phase.md sequential branch: build-time embed of the bound <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the wave serialization rules move with the fragment, verbatim in substance) plus the per-write/commit pin instruction in <sequential_execution>. Worktree-mode dispatch untouched (its self-derived toplevel IS correct there). - INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens regenerated for the new fragment; changeset added. * chore(#4254): backfill changeset PR number * fix(#4254): accept backslash-separated Windows drive pins CI on windows-latest showed every permit-path test failing with "Actual root: <none>": pins composed from Node's path.join arrive in the backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form gate rejected before the cwd-side root was ever computed — a legitimate matching pin could never pass. The gate now accepts either separator ([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names. * fix(#4254): portable drive-form gate for MSYS bash The bracket class [\\/] that accepted backslash drive pins parses inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins — every permit-path test red with "Actual root: <none>"). Replace it with standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* — the escape form is version- and build-portable. Verified across all forms: both drive spellings accepted; bare "C:", relative, empty, and unexpanded rejected. * fix(#4254): runtime-generated backslash comparator + self-describing FATAL The Windows CI legs failed every #4254 permit-path row with 'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*). Stage misattribution: <none> appears whenever the FATAL fires BEFORE the cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired. Mechanism: the test harness spawns bash -c <script> through the Windows command-line boundary; that round-trip applies one extra shell-quoting pass with double-quote semantics — a backslash written twice in the script text arrives halved, while a lone backslash survives (the pin displays intact; row 9's pure-bash gate independently showed the halved pattern rejecting C:\ while C:/ still passed its surviving arm). On windows-latest every pin carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...), so the gate ate every pin before the actual root was ever computed. Fix, robust by construction: - the drive-form gate generates its backslash comparator at RUNTIME (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now contains no doubled backslash anywhere, enforced by a regression assertion on the extracted guard text; - the FATAL self-describes: Guard stage (pin-unbound / form-gate / actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line carrying git's own stderr for capture failures and both compared values for mismatches — future platform failures name their stage in the log; - row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296 Minor 1 duplication smell) is replaced by driving the SHIPPED guard and asserting the stage; rows 2/4 pin the new stage machinery. Validated on darwin across drift/match/relative/unbound/empty/bare-drive/ forward-and-backslash drive forms, each also re-run under a simulated Windows transit (every doubled backslash halved) with identical outcomes. * fix(#4254): close the empty-comparator fail-open seam in the drive-form gate Self-review of the runtime-generated backslash comparator: if printf's octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical fail-open path. Fail closed with a self-describing diagnostic instead of trusting the shell's printf. --------- Co-authored-by: sim <sim@local> |