c6df4e1e463c3e4d0844ef7d4624c45477bd7a13
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c6df4e1e46 |
fix(#4455): autonomous.md and complete-milestone.md resolve STATE/ROADMAP/MILESTONES/PROJECT/REQUIREMENTS through the workstream-scoped init fields (#4542)
* fix(#4455): thread workstream-scoped paths through autonomous and complete-milestone workflows autonomous.md and complete-milestone.md read/wrote hardcoded literal `.planning/STATE.md` / `.planning/ROADMAP.md` / `.planning/milestones/...` paths in their shell fences, bypassing workstream scoping entirely. With GSD_WORKSTREAM=alpha set, planningDir(cwd) correctly resolves into workstreams/alpha/, but a literal `cat .planning/STATE.md` still read the ROOT file (or silently returned empty if root state was absent) -- reproduced deterministically in the issue's own repro. Root cause: each workflow step's bash fence is a separate shell invocation, and cmdInitManager/cmdInitCompleteMilestone's JSON payloads never carried resolved state_path/roadmap_path/archive_dir fields for the workflows to extract -- unlike cmdInitPlanPhase, which already does this correctly and is the pattern this fix mirrors. - src/init.cts: cmdInitManager and cmdInitCompleteMilestone now emit state_path/roadmap_path (workstream-scoped via planningDir(cwd), existence-checked, toPosixPath'd, null when absent -- identical to cmdInitPlanPhase's existing contract) and archive_dir (the milestone archive directory, composed the same way milestone.cts's already-correct archive helper does per #1911). - autonomous.md: discover_phases and iterate now extract state_path via the already-fetched INIT_MANAGER payload instead of hardcoding `.planning/STATE.md`; iterate's second, previously-separate hardcoded read is folded into the same fence (no double-fetch); lifecycle step 5b checks the resolved archive_dir instead of a hardcoded milestones path. - complete-milestone.md's reorganize_roadmap_and_delete_originals step (which previously called no init command at all) now fetches init.complete-milestone and uses the resolved roadmap_path/state_path/ archive_dir for the backlog read, the write-guard sentinel's armed content, the Write-tool target for the reorganized ROADMAP.md (the sentinel fence now echoes the resolved path so the executing agent can see it), and the safety-commit --files list. `.planning/MILESTONES.md` and `.planning/PROJECT.md` stay literal root paths -- documented shared files, per the issue's explicit "not a blanket replacement" scope. Regression tests extract and execute the real bash fences (with a stubbed gsd_run) rather than string-matching the markdown, covering flat mode (unaffected), an active workstream (the issue's own repro shape, now correctly resolving), the no-double-fetch requirement, and a dedicated guard locking MILESTONES.md/PROJECT.md as shared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): add changeset for workstream-scoped autonomous/complete-milestone fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): close write-guard gap on workstream-scoped curated paths Isolated security review of the #4455 fix (workstream-scoped STATE/ ROADMAP/milestone-archive path resolution in autonomous.md and complete-milestone.md) flagged that hooks/gsd-write-guard.js's CURATED_PATTERNS only matched root-level .planning/ paths, never .planning/[<project>/]workstreams/<ws>/... — meaning the catastrophic- shrink guard silently never engaged for a workstream-scoped write. This is directly relevant here: the #4455 change makes a workstream- scoped ROADMAP.md Write reachable via complete-milestone.md's own explicit sentinel-hatch instructions, which assume guard protection that did not actually exist for that path shape. Extended CURATED_PATTERNS with the three workstream-scoped equivalents; consumeSentinelFor's own path-derivation logic needed no change since it derives from the actual write target. Verified empirically (a 293->16 line workstream ROADMAP.md shrink now correctly returns exit 2 / decision:"block") and with 5 new regression tests. Also addressed a code-review nit on the core #4455 fix: cmdInitCompleteMilestone called planningDir(cwd) three separate times instead of caching it once. Accepted as-is (not fixed): complete-milestone.md's reorganize_roadmap_and_delete_originals step re-fetches `gsd_run query init.complete-milestone` three times across its fences rather than merging the first two (no state-changing Write between them, unlike autonomous.md's iterate step which does merge). This is an efficiency nit, not a correctness bug — merging risks disrupting the step's prose flow and its existing binding test for a non-functional gain. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): add changeset for the write-guard workstream-scope fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): fix gsd-test-surfaced regressions from workstream-path fix Running gsd-test against the full #4455 diff (including the write-guard security fix and the cmdInitCompleteMilestone caching nit) surfaced four real, non-flaky failures, all direct consequences of editing gsd-core/workflows/autonomous.md and complete-milestone.md: 1. tests/autonomous-converge.test.cjs pinned the OLD hardcoded `STATE_CONTENT=$(cat .planning/STATE.md ...)` read in both discover_phases and iterate. That is exactly the literal-path behavior #4455 fixes, so the test needed updating to assert the new init.manager-resolved `STATE_PATH` read instead (with an explicit doesNotMatch guard against regressing to the old literal). 2. tests/workstream-scoped-paths.test.cjs's own "no-double-fetch" test counted gsd_run invocations via a shell variable incremented inside the stub function — but `INIT_MANAGER=$(gsd_run ...)` runs gsd_run inside the command-substitution SUBSHELL, so that increment never survives back to the parent shell and the counter always read 0. Switched to a file-based call log (one byte appended per call), which survives the subshell boundary. 3. tests/compact-content-partition-guard.test.cjs's disjointness check flagged the reorganize_roadmap_and_delete_originals step's new `INIT_CM=$(gsd_run query init.complete-milestone)` fetch (added 3x, per the accepted-as-is disposition in the prior commit) as byte-identical to a pre-existing, unrelated fetch already present in complete-milestone/detail/elaboration.md's handle_branches section (§2). Same idiom, same conventional variable name, coincidentally colliding across the spine/detail split boundary. Renamed the new step's local variable to INIT_REORG — a distinct, purpose-specific name is arguably better practice anyway for two logically unrelated fetches, and it removes the literal collision honestly rather than restructuring the split. 4. tests/benchmark-compact-content.test.cjs reported real byte-count drift in the committed baseline (autonomous.md and complete-milestone.md both grew from the #4455 content). Refreshed via `node scripts/benchmark-compact-content.cjs --write`. Verified: node scripts/benchmark-compact-content.cjs --check now reports the baseline up to date; a standalone invocation of checkDisjointness() against the real repo state now reports zero violations across all 6 registered splits; manual bash-fence execution of both the autonomous.md iterate fence (call count = 1) and the complete-milestone.md backlog fence (with INIT_REORG) confirms correct behavior. Emitted-Drift-Ack-Growth: autonomous.md — #4455 workstream-scoped STATE.md path resolution replaces hardcoded literal reads Emitted-Drift-Ack-Growth: complete-milestone.md — #4455 workstream-scoped STATE/ROADMAP/archive path resolution replaces hardcoded literal reads Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): MILESTONES.md/PROJECT.md/REQUIREMENTS.md are workstream-scoped too, and so is project-only mode Fresh isolated code-review and security-review passes against the full diff (run after the previous gsd-test-surfaced fixups landed) each found one real, confirmed defect: Code review: the safety-commit `--files` list and the REQUIREMENTS.md `git rm` step both hardcoded `.planning/MILESTONES.md`, `.planning/PROJECT.md`, and `.planning/REQUIREMENTS.md` as literal root paths — but src/milestone.cts's cmdMilestoneComplete writes MILESTONES.md via `planningPaths(cwd).planning` (the workstream base) and PROJECT.md/REQUIREMENTS.md resolve the same way through `planningPaths().project`/`.requirements` (src/planning-workspace.cts). Only `todos` is the documented root-scoped exception (#4256); an earlier version of this fix wrongly generalized that exception to MILESTONES.md/PROJECT.md too, and the now-corrected test previously enshrined that wrong behavior as intended. Under an active workstream, the safety commit would have silently missed the actual files `milestone complete` just wrote, and the git-rm step would have targeted the wrong (root) REQUIREMENTS.md entirely. Fixed by exposing `milestones_path`/`project_path`/`requirements_path` from init.complete-milestone (src/init.cts) and resolving all three through them, the same pattern already used for state_path/roadmap_path/ archive_dir. The four remaining literal MILESTONES.md/PROJECT.md mentions elsewhere in complete-milestone.md (lines ~12-13, ~441, ~607, ~662) are display-only prose in status/summary message templates, not actual file operations — left as-is; they are a cosmetic path-display inaccuracy under an active workstream, not a data-integrity bug like the two fixed here. Security review: confirmed the write-guard fix from the prior commit is correct and complete for workstream scoping, and independently surfaced the same project-only gap the code-review pass above also caught structurally: `CURATED_PATTERNS` had no pattern for `.planning/<project>/...` (GSD_PROJECT set, GSD_WORKSTREAM unset) — planningDir(cwd) supports that shape independently of workstream nesting, so it is reachable, not hypothetical. Fixed by adding three more patterns, verified empirically (a project-scoped 292->16 line ROADMAP.md shrink now correctly returns exit 2 / decision:"block") and with 6 new regression tests. Verified: manual bash-fence execution of the corrected commit-files and requirements-rm fences (both flat mode and GSD_WORKSTREAM=alpha) resolves to the right paths in both cases; a standalone invocation of checkDisjointness() against the real repo state still reports zero violations; the benchmark baseline was refreshed again for the further size change (already covered by the existing Emitted-Drift-Ack-Growth trailer on complete-milestone.md two commits back — that trailer is read over the whole merge-base..HEAD range, not per-commit, so it still applies here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4455): backfill changeset PR numbers and correct final scope pr: 0 -> pr: 4542 for both fragments, and updated both bodies to reflect the final fix scope (MILESTONES/PROJECT/REQUIREMENTS are workstream-scoped too, not shared-root exceptions; the write-guard fix also covers project-only scoping, not just workstream nesting). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): lifecycle-5b archive-path assertions use the fence's own separator, not path.join PR CI's windows-latest shard 3/3 failed: "expected ls to find the root archive file, got: ...\milestones-root/v1.0-ROADMAP.md". The autonomous.md lifecycle step 5b fence composes the checked path with a literal bash `/` (`"${ARCHIVE_DIR}/v${milestone_version}-ROADMAP.md"`), which on Windows yields a MIXED-separator path — Windows backslashes from archiveDir plus one trailing `/`. My test's assertion used path.join(archiveDir, 'v1.0-ROADMAP.md') instead, which on a Windows Node process produces an all-backslash path that never matches the fence's mixed-separator output. Both assertions in that describe block now mirror the fence's own literal `/` concatenation (`${archiveDir}/v1.0-ROADMAP.md`) instead of path.join — matching the style the other two describe blocks in this same file (safety-commit --files list) already used correctly for the identical archive-dir pattern, so this brings the one outlier into line rather than introducing a new idiom. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4455): write-guard sentinel comparison now realpath-resolves the token, not just the target PR CI's macos-latest full-test shard 2/3 failed a #4455 test: "the sentinel hatch ... unblocks a workstream ROADMAP.md write" got status 2 (still blocked) instead of 0. Root cause, unrelated to the Windows fix in the previous commit: hooks/gsd-write-guard.js's main flow realpath-resolves the Write TARGET before the curated-pattern match (round 9 Minor 1's symlink-before-match fix, `filePath = fs.realpathSync(filePath)`), but consumeSentinelFor resolved the sentinel TOKEN's absolute path via plain path.resolve() with no realpath step. On macOS, os.tmpdir() resolves through a /var -> /private/var symlink, so a test's cwd (lexically under /var/folders/...) and its realpath'd target (/private/var/folders/...) diverge — an armed, correct sentinel then never matches the realpath'd target string, and the guard stays incorrectly blocked. This is not macOS-specific in principle: ANY cwd sitting under a symlink (a symlinked project checkout, a symlinked worktree) hits the same asymmetry — gsd-test's Linux bench runs never caught it because /tmp there is not a symlink. Fixed by applying the same fs.realpathSync (with the same keep-lexical-on-failure fallback the caller already uses) to the token's resolved path before comparing. The named file is already known to exist at this point (the caller only reaches consumeSentinelFor after successfully reading the target), so realpath is expected to succeed in the legitimate case; a garbage/mismatched token still fails safe (verified — falls back to the lexical path, still mismatches, stays blocked). Verified: reproduced the exact bug locally (macOS) via os.tmpdir() before the fix, confirmed it resolves after; the negative case (sentinel armed for a DIFFERENT file) still correctly blocks; the pre-existing relative-token sentinel tests (predating #4455) still pass; a garbage/non-existent token still fails safe. Added a deterministic, cross-platform regression test using an explicit symlink (skipped on Windows, matching the existing round-9 symlink test's own skip condition) so this class of bug is caught by gsd-test's Linux bench too, not only by a real macOS CI run. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
69e7afd0c7 |
chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags an unbounded */+/{n,} quantifier over a broad character class ([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the exact #2128-fixed shape) applied to a regex whose match target is data-flow-traced to readFileSync content. eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer shared with no-crlf-fragile-split (Phase 2) rather than a second copy — no-crlf-fragile-split refactored onto it with zero behavior change, parity-tested. Real triage, not 798 mechanical edits: the ADR's census (2026-08-08) screened every unbounded quantifier in the tree unscoped. Correctly scoped to readFileSync-derived content (matching Phase 2's own G2/G3 scoping), the rule found 162 real hits across two detection waves — the second wave (93) surfaced only after a genuine off-by-one bug in this rule's own first draft was caught while writing its RuleTester tests and fixed (the bug silently missed every directly-quantified [\s\S]* with no gap before the quantifier — exactly the class this rule exists to catch). 3 hits landed in production src/ (commands.cts, milestone.cts, roadmap.cts) and were each empirically timed against adversarial input (matching #2128's own measured-not-assumed precedent) — all confirmed linear-time/benign, left unbounded with a measured-evidence comment rather than mechanically bounded. The remaining 159 are test-file fixture parsing (test-author-controlled, fixed-size content, not adversarial input) — each suppressed with a specific, non-generic reason. Zero functional behavior changed anywhere in this diff. tests/no-pending-3212-markers.test.cjs locks the epic's own closing invariant (ADR §7: "assert zero pending #3212 markers remain") — ground truth confirmed trivially true today (no phase left any such marker behind), now regression-locked going forward. Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): correct rule category mislabel, add CI test-scope entry An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs mistakenly carried meta.docs.category: 'Portability', copied from a sibling rule without realizing what that implied: docs/contributing/cross-platform- portability-rules.md governs an ADR-1703 rule family under a hard "zero escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic — and its eslint-disable-next-line suppressions (159 of them, added earlier this same phase after empirical benign-verification) are an intentional, correct design, not a bypass. Corrected to category: 'Best Practices', matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the same epic, which is also correctly outside PROTECTED_RULES), and the rule's own docstring now states this explicitly so a future reader doesn't have to re-derive it. Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their own test suites under targeted CI selection — was previously unregistered and invisible to that fast-path (this PR's own gsd-test checkpoint runs the full suite regardless, so this only affects future narrowly-scoped PRs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic) Security review found the rule meant to catch algorithmic-complexity bugs had one of its own: hasUnboundedBroadQuantifier's negated-class inner scan walked from each `[^` occurrence to the next `]` (or EOF) with no bound, while the outer loop only ever advanced by one character — O(n²) total work on a pattern with many unclosed `[^` runs. Runs unconditionally inside checkPattern on any `new RegExp('literal string')` argument in any linted file, before the (cheap) readFileSync data-flow gate — so a single crafted string literal, no valid regex syntax required, could make `npm run lint` / CI hang. Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/ 16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n, quadratic); extrapolated, the 300000-char repro from the finding would run ~165s. Post-fix (bail the inner scan once units exceeds the rule's own 1-2-unit scope, rather than continuing to hunt for a closing `]`), the same 300000-char input runs in 8.7ms via the real rule module, independently reconfirmed at 18ms via a fresh Linter.verify() call. New regression row in tests/no-unbounded-quantifier.rule.test.cjs asserts the RuleTester run on a 50000-char adversarial pattern completes and returns a defined result — no wall-clock assertion (CLAUDE.md Clock Seams / local/no-elapsed-assertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch next merged 12 more PRs during this PR's review. Two consequences: - tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own workflow .md content — the same Class A pattern as the ~159 sites already triaged elsewhere in this PR. Suppressed with the same established reason. - lint-allow-test-rule-refs' ratchet ceiling needed re-raising again (301 -> 303) for the same reason as the two prior bumps: organic growth from unrelated, already-reviewed PRs landing concurrently, not a defect in this branch's own diff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5fd5c81042 |
test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync, each returning a typed discriminated union { outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }. Every call is timeout-bounded; there is no unbounded path. runGsdTools becomes an adapter over the seam. Its legacy { success, output, error, exitCode } shape and retry-once-on-kill behaviour are preserved byte-identically, so none of its 136 caller files change. Outcome discrimination was corrected against probed runtime behaviour rather than assumption: a timeout and a maxBuffer overflow are identical on both status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs ENOBUFS). Overflow is therefore classified before timeout. This fixes a live defect — the previous isKilled() treated an overflow as a kill, retried it for a second full 60s run, and then reported "host OOM or scheduler contention" for a child that had merely printed too much. Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs, which brought 31 previously unlinted shared helpers under the same rules their sibling test files already obey, and fixes the 5 violations that surfaced — including a bare npm invocation without shell:true in tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now routed through the existing portable runNpm helper. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): migrate every local spawn wrapper onto the process seam Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name, parameter list, return shape and post-processing (JSON parse, ANSI strip, env sanitising, field extraction) — only the spawn mechanism changes, so no test assertion moves. The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather than a fourth primitive; it is explicit rather than inferred from the file extension, because guessing an interpreter from a path fails silently when a script's name does not match its shebang. Seven wrappers were previously unbounded and now carry an explicit timeout sized to what each actually runs, not the seam default. Two of those seven (gsd-write-guard, lint-docs-command-form) were absent from the issue's inventory entirely and were found by scanning after the migration. Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md reference section covering the three primitives, the discriminated union, and the two rules the seam enforces. Scope disclosure recorded in the phase design notes: the issue scoped three identifier names. A scan for local helpers that spawn AND return the spawn result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded direct git call sites. This change bounds 25 of those. The remaining surface is the same defect class and is NOT closed by this PR. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify an externally-killed child as KILLED, not EXITED Blocker found in this branch's own diff, independently confirmed by an isolated reviewer. A child killed by an external signal — a genuine bench OOM kill — makes spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field. The seam's "no error implies EXITED" rule therefore classified it as a clean exit, and runGsdTools returned { success: false, exitCode: 1 } without retrying. That silently defeated the #969 kill-discrimination for precisely the case it was built for: the old isKilled() fired on `signal != null`, retried once, then threw a labelled resource-starvation error. A real OOM would have been reported as an ordinary assertion failure. Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes the adapter retry on TIMED_OUT or KILLED — reproducing the old `killed || signal != null || code === 'ETIMEDOUT'` condition exactly. SPAWN_FAILED still does not retry (matching the old behaviour, where signal was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate divergence: the old code retried it because signal was SIGTERM, burning a second 60s run on a child that had merely printed too much. All five outcomes verified against the live runtime rather than assumed: SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT), >1MB stdout -> buffer_overflow (ENOBUFS). Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): address standards-review findings on this branch Three findings from the standards axis of the review, all in this branch's own diff. The CONTEXT.md glossary entry this branch introduced was already stale on the branch's own last commit: it enumerated a 4-member OUTCOME while the code had 5, because the KILLED fix did not update it. That is precisely the drift the "module changes update Domain-terms" gate exists to catch, so the entry now lists all five and explains KILLED. api-coverage-gate-e2e compared an outcome against the raw string 'exited' rather than OUTCOME.EXITED, the only such outlier; the enum is now imported and used. A sweep for the other four outcome literals found no further comparison sites. Three call sites hand the literal bash flag '-c' to the seam's first parameter, which the JSDoc described as an absolute script path. Rather than add a fourth primitive, the contract is corrected to match reality: the parameter is renamed `target` and documented as the first argv element handed to the interpreter — normally a script path, but for an interpreter invoked with an inline program it may be that interpreter's own flag. No behaviour change. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): assert the cross-platform timeout contract, not the macOS one The remote runner failed on both Linux lanes (node 22 and node 24, identical) while the same tests passed locally on macOS. Two assertions encoded a platform-specific behaviour as a cross-platform guarantee. When spawnSync times out, macOS preserves the child's partial stdout/stderr; Linux discards it and returns empty strings. Verified on node v26.5.1 both ways. The seam passes through whatever spawnSync hands it and cannot manufacture output that was discarded, so the production code was correct — the tests were wrong. Both tests now assert the guarantee the seam actually makes on every platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being strings rather than undefined or a Buffer. The partial-content assertions are retained behind an explicit process.platform === 'darwin' guard so the macOS coverage is not lost, and the first test is renamed to say what it now guarantees. This is the failure mode the remote matrix exists to catch: local macOS verification would have shipped it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout Windows CI caught two defects the Linux matrix could not. tests/context-predicates-query.test.cjs passes a 32K-char argv value. On Windows that exceeds the argv limit and spawnSync fails with code ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise, status === null implies TIMED_OUT" — swallowed it, so the adapter retried a spawn that can never succeed and then threw the resource-starvation error. The old isKilled() returned false for that shape and returned an ordinary failure result. TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set. Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG, E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform whose timeout errno differs classified correctly, so the greedy catch-all is no longer needed. The second defect is a contract regression I introduced and had claimed otherwise. That same test asserts `typeof r.exitCode === 'number'`, and toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the assertion failed on type. The old code returned `err.status ?? 1` on every non-retried failure path. The adapter now returns 1 again for both, and the comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is retracted: the seam keeps the richer truth (exitCode null plus a distinct outcome), the legacy adapter keeps the old numeric contract its callers actually depend on. Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT -> SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL -> KILLED; clean exit -> EXITED. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c61dd49d95 |
enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks (decision: 'block', exit 2) a whole-file Write collapsing a curated .planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md, STATE.md) below 40% of its on-disk line count. Files under 40 lines are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message) bypasses for legitimate milestone resets. Fix 3 of #973 — the only defense independent of per-agent tool config. Registered on the Claude plugin surface (hooks.json), settings-json runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and INVENTORY regenerated; regression tests negative-controlled (16/16 RED with the hook absent, 16/16 GREEN with it present). * chore(#2255): backfill changeset pr number to 2301 * enhance(#2255): address review — fail-closed reads, typed block output, registration, property test Review fixes for trek-e's CHANGES_REQUESTED on PR #2301: - Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES (no-shipping-drift test). - Blocker 3: update the always-on hook enumerations in ADR-766 and CONTEXT.md from six to seven. - Major 4: fail CLOSED on non-ENOENT read errors — only a missing file (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a typed readError field and the override still honored. Tested, with a negative control against the pre-fix hook. - Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the floor; sub-floor always exempt), boundary examples pinned. - Major 6: block output now carries typed oldLines/newLines/ overrideEnvVar fields; tests assert on those instead of regexing the free-form reason string. - Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS bypass on macOS/Windows); limit+1 boundary tests added for both the floor and the ratio. * enhance(#2255): engage the write guard on Kimi's native payload shape The guard shipped with Claude-vocabulary checks (tool_name 'Write', tool_input.file_path), which #2304 showed leaves a guard dormant on Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli forwards its native payload verbatim — tool_name 'WriteFile' (bare or module-qualified) and tool_input.path per its tool schemas (src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown name, and exited 0. Apply the same per-guard normalization PR #2326 gives the three sibling guards (name + field mapping, inlined — hook scripts stage as standalone files), and write the block reason to stderr as well as stdout JSON: Kimi feeds stderr, not stdout, back to the model on exit 2, so a stdout-only reason blocks without telling the model why or naming the documented override. Regression tests pipe Kimi-shaped payloads (engage, qualified-name, stderr-reason) plus exemption pins (StrReplaceFile stays out of scope by design; non-curated paths pass) — verified red against the pre-fix guard, green after. * enhance(#2255): rebase onto next; regenerate golden-parity fixtures * enhance(#2255): wire the escape hatch into complete-milestone's reorganize step Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP reorganize — the tree's only legitimate milestone reset and the exact caller GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the rewrite through a shell write with the hatch set on the command (a hook inherits the runtime env, so a bare Write cannot carry a per-step override), and a binding test derives the env var name from the guard's typed output and asserts (a) the workflow step sets it and (b) the guard passes the identical catastrophic payload under it — so the next complete-milestone.md edit cannot silently re-break the wiring. * enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string reconstruction were unreachable-by-effect — the guard exits 0 for any tool_name !== 'Write', so nothing ever read the fields they set, leaving guaranteed-surviving mutants against the Stryker bar. The map now carries only WriteFile -> 'Write'; the StrReplaceFile exemption test message states the fall-through it actually exercises. * enhance(#2255): review minors — American spellings; writeSync before exit(2) Minor 1: normalised/normalise -> American house style. Minor 2: the two block paths wrote stdout+stderr via async pipe writes then exit(2) — async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block payload durable. * enhance(#2255): assert stderr equals the typed reason, not raw prose Minor 3: the last raw-text match in the suite pinned override-name prose on stderr. The contract is "stderr carries the reason Kimi feeds back" — now asserted as stderr non-empty and byte-equal to the parsed stdout.reason. * enhance(#2255): bind the write-guard's Kimi normalization into the parity test Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with nothing binding it. This extends PR #2326's kimi-guard-normalization-parity test (same path and helpers, authored as a superset so either merge order resolves cleanly): sibling byte-parity is existence-gated zero-or-all — trivially green until #2326 lands, full-strength after — and the write-guard copy is bound semantically (map is the value-inverse of convertKimiToolName; the Kimi name for Write must map, or the guard is dormant on Kimi; the path -> file_path half must be present). Byte-parity is deliberately not asserted for this copy: it legitimately omits the Edit-class mapping (Major 1 — dead code in a Write-only guard). * enhance(#2255): refresh golden-parity fixtures for revised guard + workflow * chore(#2255): regenerate golden fixtures after rebase onto next The committed fixture hashes were generated against a tree predating next's latest 11 commits, which independently modified the same install-parity surface. Rebased onto next and regenerated with `npm run gen:golden`. Verified: against upstream/next the regenerated fixtures differ by exactly this PR's own entries -- hooks/gsd-write-guard.js (new), hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and gsd-core/workflows/complete-milestone.md. No unrelated drift. * fix(#2255): regenerate workflow size baseline for complete-milestone `complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2 review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step, but tests/workflow-size-baseline.json was never regenerated. The per-file workflow baseline test (issue #1074) failed on ubuntu-latest/22 and both macOS shard 1/3 jobs. The growth is justified: it is the escape-hatch binding requested in review round 2 (the guard must not hard-block the tree's only legitimate milestone reset), not incidental bloat. Regenerated via `npm run size:baseline`; the diff is exactly the one entry. * chore(#2255): regenerate golden fixtures and size baseline after rebase onto next * enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert (no PreToolUse hook exists on Bash in this family; the write succeeded by dodging the guard, not by the override firing) and the protection was prose. The hatch is now a transport code consults: complete-milestone's reorganize step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the Write tool as the sanctioned path, and the guard — at the block point only — verifies the sentinel is fresh (15 min) and names the pending target, then CONSUMES it and allows that one write. Path-bound + single-use + freshness keep it from becoming a standing unlock. The env var remains as the interactive transport, where it can actually reach the hook. Regression tests written first (negative control: 3 failed pre-fix): the armed-sentinel Write passes and consumes; stale does not exempt; a token for a different file neither exempts nor is consumed; the binding test now takes the sentinel name from the guard's typed output (overrideSentinel), asserts the step arms it, and asserts the step no longer routes the rewrite around Write via a shell pipe. Also in this commit, same file: - m2: block emission is exception-safe — emitBlock() wraps both writeSync sites in their own try/catch that still exits 2, so an EPIPE can no longer convert fail-closed into the outer catch's fail-open. - Header discloses the two reviewed design limits (cumulative sequential shrink; lexical match vs symlinked paths) per round-5 scoping. * docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4) USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the runtime-hooks-surface registration comment, and the changeset now describe both hatches — the single-use sentinel for workflow steps and the env var for interactive use — instead of implying a per-step env can reach a hook. The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)` parenthetical the repo's fragments use (round-5 m4). * chore(#2255): regenerate derived families on the rebased tree (full sweep) Full generator sweep after rebasing onto next @ the body-parser-patched lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every regen delta verified to be either a PR-owned entry (gsd-write-guard.js, complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's committed value for entries our arbitrary-side conflict resolution had left stale (all 18 runtime fixtures checked mechanically). * test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY retry budget; the sentinel disarm now rides it like every other teardown. * chore(#2255): regenerate derived families after rebase onto next Full sweep on the rebased tree (build -> gen-inventory-manifest -> gen:golden -> size:baseline). Every delta is either a PR-owned entry (hooks/gsd-write-guard.js, its registration surfaces hooks/managed-hooks-registry.cjs and the two plugin buses, gsd-core/workflows/complete-milestone.md) or exact convergence to next's committed value across all 18 runtime fixtures. * chore(#2255): regenerate derived families after rebase onto next @ |