5d804dd2877a506735e4959341e0efbf33658809
1082 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5d804dd287 |
fix(#3709): clear the context-monitor warn sentinel on PreCompact (#3808)
* fix(#3709): clear the context-monitor warn sentinel on PreCompact The monitor's per-session warn sentinel survived a compaction, so once the first CRITICAL of a session had fired, `lastLevel` stayed pinned at 'critical' for the rest of the run. The hook was already wired to PreCompact (#772), but read the event only at the very END, and solely to pick an output envelope. Two documented behaviours died as a result: - "First warning always fires immediately" — the first warning of the post-compaction cycle was debounced instead. - "Severity escalation (WARNING -> CRITICAL) bypasses debounce" — computed as `lastLevel === 'warning'`, which can never be true again, so every later CRITICAL waited out the full five-tool-use debounce, exactly when an immediate warning matters most. `criticalRecorded` was equally sticky: a session that compacted and later truly ran out kept a /gsd:resume-work breadcrumb (#1974) describing the earlier near-miss rather than the exhaustion that ended the run. Reproduced first, with the issue's own literal repro, including the detail that the compaction consumed a debounce slot (callsSinceWarn 0 -> 1). The reset runs BEFORE the metrics read, deliberately: a post-compaction reading is healthy again, so the ENOENT / stale / above-threshold branches would all exit first and never reach it. Returning early also stops the compaction from eating a slot of the cycle it was meant to restart. The event name is now read once through a shared `readEventName()` helper, so this reset and the #2289 output allowlist cannot drift on what counts as "no event name". Seven rows against a real sequence (the defect is state carried ACROSS calls, so they need their own driver — the existing helpers delete the sentinel after each invocation). Reverting the reset turns SIX of them red; the seventh is the non-vacuity row asserting a NON-compaction event must not clear the sentinel, which correctly passes either way. AC4 initially passed with and without the fix — asserting `criticalRecorded === true` is vacuous when the seeded stale sentinel already carries it. It now seeds a `staleProbe` marker that can only survive if the sentinel survives, so its absence is what proves the state was rebuilt. hooks/dist/ is gitignored and regenerated by build:hooks, so no committed dist copy needs syncing. Verified: `npm run lint:ci` exit 0; acceptance criteria 1-6 driven end-to-end against the real hook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): reset ahead of the config gate, and pin the placement itself Codex review of the #3709 fix, before opening the PR. Three findings, all in this change's own new code. 1. `context_warnings: false` prevented the reset. The config early-exit sits ABOVE where the reset was placed, so a session that disabled warnings, compacted, then re-enabled them mid-session resurrected the stale sentinel and the original bug with it. Config is re-read per invocation, so that sequence is supported rather than hypothetical. The reset now runs ahead of the config gate: clearing the sentinel is CLEANUP, not a warning — state that must not outlive a compaction should not outlive it merely because warnings are switched off right now. It cannot emit anything from there, so the disabled contract is untouched. 2. Nothing pinned the "before the metrics read" placement. Every row wrote a fresh metrics file, so the reset could have been moved below the metrics read, the stale check, or the healthy-threshold exit with all seven rows still green — while a REAL PreCompact, which carries no fresh metrics and follows a recovery to healthy usage, silently kept its sentinel. Three rows now pin it: no metrics file at all, usage recovered to healthy, and warnings disabled. Each catches a distinct wrong placement — moving the reset below the config check reds the third; below the metrics read reds all three. 3. The absent-sentinel row proved nothing. `assert.doesNotThrow` was vacuous because the driver caught every child exit, so a hook that exited 1 on the ENOENT unlink would still have passed. The driver now returns the exit code and the row asserts it is 0. Also corrected the `readEventName` comment: it said the event is "read once", which is not literally true — there are two call sites. The point is one DEFINITION of what counts as an event name, so the reset and the #2289 allowlist cannot drift; the comment now says that. Verified: 60 rows in tests/perf-317-context-monitor-fs.test.cjs, 0 fail, with both placement mutations driven to red and reverted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#3709): backfill changeset pr number The fragment shipped with the documented `pr: 0` placeholder, which the changeset lint treats as always-silent, because the PR number does not exist until the PR is opened. Backfilled to 3808 now that it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): a compaction clears the stale reading too, not just the state Review round 1. Major 1 was right and it mattered: clearing only the sentinel traded a warning that never fires for one that fires when it must not. The statusline bridge still holds the PRE-compaction reading, and STALE_SECONDS is 60, so for up to a minute it still reads fresh and still says the context is exhausted. With the sentinel gone, firstWarn is true, so the next PostToolUse emitted a spurious CONTEXT CRITICAL immediately after the compaction that FREED the context — and flipped criticalRecorded, spawning a false context-exhaustion breadcrumb. That is the same breadcrumb inaccuracy #3709 exists to fix, re-entered from the other side. Reproduced before fixing, exactly as the review described. A compaction now invalidates the warning state AND the reading that produced it. Removing the bridge loses nothing: the statusline owns that file and rewrites it on every render, and its absence is already the "no reading yet" state a fresh session starts in, which exits silently. Two things my own verification caught while fixing it: - The first attempt did NOTHING. metricsPath was declared below the PreCompact block, so referencing it hit the temporal dead zone, threw, and the outer catch swallowed it into a silent exit 0. The probe still printed "silent", which looked like success but was the old debounce. metricsPath is now hoisted beside warnPath. - The new Major 1 row was VACUOUS. The driver's `metrics: false` DELETES the bridge, but the defect is a bridge that is still there and still reads fresh, so the row passed on the ENOENT early-exit rather than on the fix. Only the sentinel-only mutation exposed it. The driver grew a `metrics: 'keep'` mode that leaves the stale file in place; both Major 1 rows now red under that mutation. Also from the review: - Minor 1 — the compaction-abort path is now stated in the source rather than left silent, including why a conditional reset (SessionStart source "compact") is out of scope for this fix. - Minor 2 — docs/context-monitor.md completed: PreCompact wiring and the early return under How It Works, a table of all three things the reset clears, the breadcrumb guard, the warnings-disabled interaction, and the never-block property under Safety. - Minor 3 — changeset trimmed from ~1,400 chars of implementation narration to the user-visible change. - Nit 1 — a failed unlink (Windows EPERM/EBUSY) no longer leaves the bug silently intact: the file is neutralised in place instead, with a shape safe for each (an empty sentinel, a timestamp-0 bridge). - Nit 2 — reviewer-process narration removed from shipped test source. The remaining "Codex" mentions are pre-existing and name the RUNTIME. - Nit 3 — the debounce-slot row now asserts the observable consequence (the first post-compaction warning fires) rather than repeating AC1's assertion. - Nit 4 — the file docblock now lists the folded-in blocks and asks the next contributor to extend it. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): the unlink-failure fallback truncates to empty, matching deletion The fallback wrote well-formed neutral values, and neither was equivalent to the deletion it stood in for: '{}' parses, so firstWarn was false and the first post-compaction warning was debounced — AC2 undone on exactly the path the fallback exists for — and '{"timestamp":0}' was never stale (the guard is `metrics.timestamp && ...`), so the flow reached emit with remaining === undefined and injected a literal 'Usage at undefined%'. Truncating to '' makes JSON.parse throw on both reads: the sentinel read keeps firstWarn true, the bridge read falls to the outer catch and exits 0 silently (review of #3808, Blocker 1). The branch is now executed for real: an EPERM is injected into the child's fs.unlinkSync via --require preload — method monkeypatching, never chmod 0o000, which root bypasses under Docker/CI (Blocker 2). Both rows proved failing-first against the neutral-value fallback. The boundary trios at WARNING=35 / CRITICAL=25 are completed on the emit path with 34, 26, and 24 (Major 3); 36/35/25 were already pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): the truncation fallback refuses to follow a planted symlink The per-session files live in a shared sticky tmpdir, where an unlink failing EPERM is exactly what another user's planted file produces — and a planted SYMLINK would make the fallback's plain truncating write empty out its TARGET, weaponising the hook against any file its own user can write. Open with O_WRONLY|O_TRUNC|O_NOFOLLOW instead: a symlink fails ELOOP into the same give-up arm. On Windows the constant is absent and '|| 0' keeps the fallback alive there, where the held-handle case it exists for occurs and temp dirs are per-user. Found by Codex review; the new row proved failing-first against the writeFileSync fallback (victim file truncated to zero bytes). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): refuse non-regular files everywhere, not only where O_NOFOLLOW exists Codex round 2: '|| 0' removed the no-follow protection exactly where it cannot be expressed as an open flag — Windows, whose tmpdir is NOT guaranteed per-user (TEMP/TMP overrides, system-temp fallback). An lstat isFile() guard now rejects symlinks and every other non-regular shape on all platforms before the truncating open; O_NOFOLLOW stays, as the lstat->open substitution-race backstop where the platform has it. The symlink row additionally asserts the planted link SURVIVES the call, so a preload match that stops engaging can no longer pass the row vacuously off a successful unlink. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(#3709): tolerate the Windows give-up, still outlaw neutral values Both windows-latest CI lanes fail the two EPERM rows deterministically: the runners hold freshly written files with a share mode that allows DELETE (every real-unlink row passes) but refuses a truncating write-open, so the fallback's give-up arm engages — which is the fallback working as designed, not the defect the rows exist to catch. The rows are now platform-aware: POSIX still requires exact truncation and the behavioural follow-ons; Windows accepts truncated-or-untouched but still rejects the Blocker-1 regression class (a parseable neutral value is never legal anywhere), with the follow-ons gated on the truncation actually landing. Also corrects the hook comment: libuv defines O_NOFOLLOW as 0 on Windows — a no-op, not an absent constant. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): a compaction watermark closes the window bridge deletion only narrowed Round-3 Major 1: the statusline is an uncoordinated process that re-writes the bridge on every render, so a render landing between the PreCompact clear and the compaction's completion re-created the PRE-compaction reading under a CURRENT timestamp — past STALE_SECONDS, into a spurious post-compaction CRITICAL and a false exhaustion breadcrumb: the exact failure the deletion was added to prevent. PreCompact now also writes claude-ctx-<id>-compacted.json ({at}) and the metrics read drops any reading not STRICTLY newer than it — which also covers unstamped/zero timestamps once a compaction happened. Written unlink-then-O_EXCL so a planted file or symlink is never followed; failure degrades to the old narrowing. Docs and changeset now describe the watermark instead of overclaiming for the deletion. Round-3 Major 2: DEBOUNCE_CALLS and STALE_SECONDS get their trios — the gate increments BEFORE comparing, so seeds 3/4/5 pin 4-debounced, 5-emits, 6-emits; ages 59/60/61 pin the strict >. The child's clock is pinned via a --require preload (a wall-clock boundary row would flip on one second of startup delay). timestamp-0's falsy bypass is pinned directly as characterized behaviour. Mutation-proven: dropping O_NOFOLLOW, <= for <, and >= for > each red exactly one row. Minors: the symlink row's comment now names the lstat guard it actually pins, and a preload-blinded-lstat row drives the O_NOFOLLOW substitution -race backstop for real (3); absence assertions use warnRaw so a corrupt leftover cannot pass as deleted (4); the Windows give-up is an explicit t.skip, never a silent if (5); readEventName is total via String(), keeping #2289's side-effects-always-run contract for malformed event names, with a row (6); the PreCompact rationale lives once in docs/context-monitor.md with the code keeping only line-level constraints (9); the changeset is release-note-sized (10). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): the grace window covers the compaction's duration, not just its start Codex on the first watermark cut: the watermark stamps the compaction's START, so a statusline render one second later — still mid-compaction, still the old reading — passed 'strictly newer' and re-fired the false CRITICAL. Readings inside COMPACT_GRACE_SECONDS (60) past the watermark are now dropped: the window covers the compaction's own duration, a healthy reading dropped there behaves identically to an accepted one (it exits above-threshold anyway), and a genuine exhaustion warning is delayed at most one window after a compact. A watermark stamped ahead of the reader's clock is ignored — a clock step backwards or a stray file must degrade to plain staleness, never mute the monitor indefinitely. Both proven failing-first. readEventName is strict about TYPE, not coerced: String() rendered ['PreCompact'] as 'PreCompact' and would run the reset off a malformed payload. typeof: every non-string is 'no event' — silent, side effects intact — with rows for the number, hostile-object, and array-wrapped cases. The lstat-claim preload arm now writes an engagement marker the substitution-race row asserts on, so a match string that silently stops matching can no longer let the row pass off the real lstat guard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: retrigger CI — the previous wave was cancelled by an Actions outage, zero job failures Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): drive the compaction rows on the clock, not on a future stamp Round-4 review raised three majors, all in the test scaffolding around the fix rather than in the fix itself. Major 2 (taken first — it is the cheapest and it unblocks Minor 6): call() passed process.env to the child unmodified, so two rows depended on ambient GEMINI_API_KEY. The preserved Gemini fallback is `eventName === "" && !!process.env.GEMINI_API_KEY`, and readEventName returns "" for every malformed name, so with the key set the malformed-event row's `stdout === ''` assertion failed outright — reproduced by running it under GEMINI_API_KEY=x. call() now takes an explicit env, the way the sibling runMonitor helper in this file always has, and both rows pin the variable unset. (The array row survived an ambient key only because its reading was debounced — incidental, not independence, so it is pinned too.) Major 1: the AC2/AC3 rows drove the hook with a bridge stamped 62 seconds in the FUTURE — a shape hooks/gsd-statusline.js cannot produce, since it always stamps Math.floor(Date.now()/1000) on the same clock. They proved "the sentinel was cleared" while their assertion messages claimed the documented immediate-warning behaviour, which is gated behind the grace window and went unexercised. Both rows now run the real sequence on the clock-pinning preload this PR already added for the STALE trio: PreCompact at a fixed instant, then a normally-stamped render one second past the window. Verified non-vacuous — stubbing the sentinel unlink reds both. Major 3: COMPACT_GRACE_SECONDS, the one constant this PR introduces, was the only threshold without a limit-1/limit/limit+1 trio, in a PR that adds full trios for four pre-existing ones. The seeded offsets were +0, +1 and +61; the boundary itself (+60) and limit-1 (+59) were untested. Added, driven by advancing the reader's clock rather than post-dating the reading, so the reading is never ahead of the reader and only the grace gate can drop it. Verified against three mutations — `>` to `>=`, the constant to 59, and the constant to 61 — each of which reds exactly one row of the trio. No production code changed. Verified: 85/85 in this file, lint:ci exit 0, and the two Minor-6 rows now pass under GEMINI_API_KEY=x as well as unset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): harden the watermark read and pin the thresholds it introduces Codex review of the full PR found four majors. All reproduced here against the real hook before fixing. MAJOR — the watermark was write-hardened but read-untrusted. PreCompact already refuses to follow or overwrite a planted object (unlink-then-O_EXCL), but the read was a bare readFileSync, so anything the write side gave up on was followed by every later invocation. In a shared sticky os.tmpdir() that is a mute primitive — a planted recent watermark suppresses monitoring — and a symlink to a FIFO stalls a synchronous read. Measured against the pre-hardening file: a symlink to a planted watermark WAS honored and muted the monitor. The read now uses the same lstat + O_NOFOLLOW pair the sentinel path uses, plus a size bound; symlink, directory and oversized cases are all refused, with a plain-file control proving watermarks still work. MAJOR — the `now + 5` skew tolerance was an unnamed, untested threshold. It is now WATERMARK_SKEW_SECONDS with a +4/+5/+6 trio, verified against two mutations (`<=` to `<`, and the constant to 6), each of which reds one row. This is the same class as round 4's Major 3, one layer up. MAJOR — the malformed-event row shared one session across both subcases, so the hostile-object iteration's `assert.ok(s.warn())` passed off the sentinel the `42` iteration left behind. A regression throwing before the bookkeeping would have kept it green — vacuous for exactly the subcase it exists for. Fresh session per subcase, with an explicit no-sentinel precondition. MAJOR — the stale-reading row's non-vacuity is an artifact of call()'s future stamp: with a production stamp the watermark suppresses the same reading, so the row cannot isolate bridge deletion. The two guards genuinely overlap inside the window, so no end-to-end row can separate them; the comment now says so and points at the direct pin (s.metrics() === null) instead of claiming an isolation it does not have. Docs corrected where measurement contradicted them: the window NARROWS the race rather than covering the compaction's duration, and the delay is not bounded by the window alone — first recovery is watermark+61s with no skew but watermark+66s at the accepted +5s skew. Aborted compactions are muted the same way. The truncation fallback is documented as best-effort, which is what the code and the Windows rows already do. Verified: 89/89 in this file, lint:ci exit 0, symlink/directory/oversize all refused where the pre-hardening file honored them, both new trios mutation-checked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): move the PR's two new exits onto the declared-policy vocabulary #3911 / ADR-3889 migrated this hook off raw process.exit() while this PR was in review, replacing every exit with hooks/lib/hook-exit.js's allow(), which forces each call site to name its crash policy. The PreCompact reset and the watermark gate are added by THIS PR, so they did not exist to be migrated and came through the merge as the only two raw exits left in the file — caught by the new local/require-registered-exit rule. Both are ALLOW: a compaction is never blocked by this hook, which is the policy the rest of the file declares. Caught only in CI, not locally: `npm run lint` runs eslint with --cache, and the cached entry for this file predated the new rule, so a warm local cache reported clean. Re-verified with the cache cleared. allow() terminates rather than throwing, which matters for the watermark call site because it sits inside a try/catch — a throwing helper would unwind into that catch and silently drop the grace-window mute. Verified behaviourally, not by reading: the grace trio, the skew trio and the non-regular-file rows all still pass. Verified: lint:ci exit 0 with a cold eslint cache, full suite exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): harden the routine sentinel writes and the read beside them Round 7 ruled that the three routine debounce-accounting writes to the warn sentinel must match the three writes this PR already hardened: leaving the fourth unhardened beside them is the asymmetry that invites the defect back. They now go through one writeSentinel() helper using the compaction watermark's own unlink-then-O_EXCL shape, rather than a second policy — the unlink removes any existing object, and O_EXCL then refuses to create through one, so a write can only land on a fresh regular file this process made. The routine READ beside them was the last bare readFileSync on warnPath, and the same rationale applies to it verbatim; the watermark's read was hardened in round 4 for exactly this reason. Same lstat + O_NOFOLLOW + size bound. Its scope is stated in the test rather than overclaimed: lstat establishes that the sentinel is a plain regular file, not that it is trustworthy, so a cross-owner regular file at the predictable path is still read and is left as a disclosed pre-existing residual. Also fixes an accept-direction regression this PR introduced and six rounds of review missed. readEventName collapsed an ABSENT event name and a MALFORMED one onto the same '', and the preserved Gemini fallback keys off eventName === "", so with GEMINI_API_KEY set a malformed payload began emitting an AfterTool envelope. At the merge-base, data.hook_event_name.trim() threw on a truthy non-string after the side effects and nothing was ever emitted. Measured base-vs-head with a fresh sentinel per run: 42, ['PreCompact'] and {} all went silent -> EMITS, while an absent name and 'PostToolUse' were unchanged. readEventName now returns '' only for an absent name and null for a present-but-non-string one; both call sites compare for equality only, so every well-formed payload behaves identically. Five new rows, each proven fail-first with the mutations attributed separately: reverting the writes reds the write-through and non-regular rows, reverting the read reds the mute and non-regular rows, and reverting the absent/malformed split reds the Gemini row. The changeset's "behaves like a fresh session" is narrowed to name the 60-second suppression window and the best-effort reset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9 * test(#3709): pin both 4096-byte read bounds at their boundaries Round 8 asked for limit-1/limit/limit+1 coverage on the size bound the round-7 sentinel read-hardening introduced (gsd-context-monitor.js:335). The existing refusal row pads to 8192 -- a full 4096 bytes clear of the fence -- so `>` vs `>=`, or an off-by-one in the constant itself, was invisible to it. Covers the sibling bound too. The identical check guards the round-4 WATERMARK read at :278 and its refusal row pads to 8192 in exactly the same way; the review's own rationale (this file already holds WATERMARK_SKEW_SECONDS to a boundary trio, so an uncovered bound is the odd one out) applies to it unchanged. That half is a class sweep of a pre-existing bound and is test-only -- say the word and it comes out without touching the rest. Both trios assert on observable hook output rather than an internal error. Sentinel: an honored {callsSinceWarn:1,lastLevel:'warning'} keeps the debounce arm taken at remaining=30, so nothing is emitted, while a refused one falls back to first-warn defaults and emits. Watermark: honored mutes (stdout empty), refused leaves the warning. The 4097 row is the non-vacuity control for the two accept rows. Payloads are sized by measurement, with Buffer.byteLength asserted to equal the target, not by arithmetic on an assumed prefix width. Proven fail-first in both directions, with the hook restored after: `> 4096` -> `>= 4096` reds both trios (94/96); `> 4096` -> `> 4097` reds both trios (94/96); restored, 96/96. Under both mutations only the two new rows fail -- the pre-existing 8192-padded rows stay green, which is the review's fencepost claim demonstrated rather than assumed. * fix(#3709): correct the changeset's mute-window claim and a superseded comment Both from the pre-push Codex pass on the full PR. The changeset said readings are "suppressed for up to 60 seconds after a compaction starts". That is false at the accepted skew boundary, and this repo's own docs/context-monitor.md already carried the accurate figure: first recovery is watermark+61s with no skew and watermark+66s for a watermark at the +5s skew limit. Measured independently at +64 silent, +65 silent, +66 warning. The changeset now states the window plus the accepted skew, matching the doc rather than contradicting it. A comment in the malformed-event row still described readEventName as returning "" for every malformed name. Round 7 superseded that: a present-but-non-string name returns null and only an ABSENT one returns "", so a malformed payload can no longer reach the Gemini fallback at all. Marked as historical and corrected. The GEMINI_API_KEY pin stays -- the row is about readEventName's typing, not the fallback, and an ambient key would still change what it measures. Codex's three Major findings are not taken, on attribution rather than logic; the reasoning is in the PR reply. In short: the watermark does not exist at the merge-base at all (0 occurrences), so "base emits, HEAD mutes" compares a new feature against its absence rather than showing a regression; and the base sentinel read is a bare readFileSync, which blocks on a planted FIFO exactly as the hardened read would, so the TOCTOU stall is not introduced here. The underlying limits -- watermark provenance, and lstat->open races on a non-symlink substitution -- are real, pre-existing, and already offered to the maintainer as follow-ups. * fix(#3709): read both sentinels through one hardened helper; state the two limits precisely Round 9's Major, with a correction to its premise, and both Minors. The review names "watermark read/write helpers this PR adds" that a call site at :238-250 duplicates inline. There are no such helpers: this PR adds readEventName and writeSentinel, the latter a write-side primitive a read cannot call, and :238-248 is base code the diff never touched. What IS duplicated is the hardened READ. The watermark read (round 4) and the warnPath read (round 7) are the same ten lines twice -- lstat, isFile and a 4096-byte bound, O_RDONLY|O_NOFOLLOW, readSync, close -- differing only in the path variable and the error string, and that is two copies to keep in step by hand. Now one function, readSentinel(target), beside writeSentinel. Refusal throws; both callers already wrapped the read in a try/catch that degrades to "no file", so behaviour is unchanged by construction. Proven rather than assumed: with the helper replaced by a bare readFileSync in a complete scratch tree, exactly the five hardened-read rows in tests/perf-317-context-monitor-fs.test.cjs go red -- round 7's symlinked and non-regular sentinel and its size bound, round 4's non-regular watermark, round 8's watermark size bound -- so the helper carries both call sites' guarantees and the rows pin it. 96/96 with the helper in place. Minor, drop vs delay: the grace-window comment said "dropped" on one line and "delayed" three lines later, and docs/context-monitor.md said "delayed". A genuine exhaustion reading inside the window is skipped, not queued: its warning and its #1974 breadcrumb both fire on the next reading after the window, so both are delayed when a later reading comes and lost when none does -- a session ending inside the window records neither. Comment and docs now say exactly that, and that the loss is accepted over trusting a reading that may be the pre-compaction value under a fresh timestamp. Minor, ordering: the PreCompact unlink and the debounce writeSentinel(warnPath) are two writers with nothing serialising them; a debounce invocation that read pre-compaction state and lands its write after the unlink would resurrect the sentinel the reset removes. The hook relies on the host dispatching a session's hooks one at a time, which Claude Code does and the other runtimes are assumed to. Stated at the reset as an assumption, with the lock-file alternative named and not taken. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3709): write the compaction watermark through writeSentinel Review of #3808, round 10. The PreCompact watermark write was the block writeSentinel was lifted from in round 7, and it kept its own inline copy of unlink-then-O_EXCL a few lines below the helper. Round 9 flagged that write-side duplication; the round-9 reply misread it as the read side and unified only the reads. The write now calls the helper too, so the hook holds one copy of the hardened write, not two. Behaviour is unchanged: same unlink-then-O_EXCL sequence, same flags, same best-effort outer catch. The one difference is that writeSentinel closes the descriptor in a finally, where the inline copy leaked it if writeSync threw before closeSync. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7bQLXAKubb4EFLtCPiLiL * fix(#3709): read the statusline bridge through the same hardening as the sentinels Review of #3808, round 11. `metricsPath` is built one line from `warnPath` and `watermarkPath` — same tmpdir, same predictable `claude-ctx-{sessionId}` shape, same threat model this PR documents at length for its siblings — and it is the only one of the three read on EVERY invocation. It was also the only one still reached by a bare `readFileSync`, so the symlink-follow and the symlink-to-FIFO stall that rounds 4 and 7 closed on the other two stayed reachable here, on the file's highest-traffic path. It now goes through `readSentinel` like the rest. The 4096-byte bound is ample for it: the statusline writes four fixed fields (`gsd-statusline.js`), about 140 bytes with a UUID session id, so no legitimate bridge approaches it. A refusal lands in the same rethrow an unreadable or malformed bridge already did. The comment introducing `readSentinel` claimed the warn sentinel was "the one bare readFileSync". Read as scoped to `warnPath` that was true, but it reads as a claim about the file and it is not one — the bridge kept its own until this round. Corrected rather than left to mislead the next reader. Round 11 Minor: `readSentinel` discarded `fs.readSync`'s return value and assumed the buffer was full, so a file truncated between the `lstat` and the read left a zero-filled tail. It now refuses a short read. Stated plainly because it was measured: this guard has NO observable behavioural delta — deleting it leaves the new row green, because the NUL tail makes `JSON.parse` throw one line later and both paths degrade to "no sentinel". It is a consistency fix in a function whose purpose is refusing to trust what it read, and the test comment says exactly that rather than implying coverage it lacks. Five rows added: the bridge refusing a planted symlink (with an attacker-chosen reading that WOULD warn if followed, so silence is proof), a non-regular bridge, an oversized bridge, the shrink path end to end, and the direction that matters most — a healthy bridge still warns, so the hardening is not a mute. Proven by mutation: reverting the bridge to `readFileSync` reddens two rows. The shrink injection carries an engagement marker for the same reason the lstat-claim one does, learned the same way: the hook rewrites the sentinel later in the invocation, so a size check afterwards passes whether the truncation landed or not. An independent full-PR pass on this round added two more, both taken: `writeSentinel` discarded `fs.writeSync`'s return value, and a short write is permitted by the syscall — so a truncated sentinel could reach disk and every later read would reject it, silently losing the debounce accounting or the watermark this write exists to record. It now loops until the payload is written, as Node's own `writeFileSync` does, with an explicit no-progress guard. Pinned by a row that injects a one-byte first write; reverting the loop reddens it. The directory row's comment claimed it pinned the `lstat` isFile() check. It does not — measured: deleting that condition leaves the row green, because reading a directory fails on its own a line later. The comment now says the row pins the outcome, and names the symlink row as the one that pins isFile(). DISCLOSED, NOT FIXED HERE — a session id long enough to push the derived filenames past NAME_MAX. The bridge is `claude-ctx-{id}.json`; the sentinel and watermark add longer suffixes, so on a 255-byte limit the watermark stops fitting at a 230-character id and the sentinel at 233. Measured base-vs-HEAD at 233+: base is SILENT, HEAD emits the warning, because the bare `writeFileSync` base used threw ENAMETOOLONG out of the warning path while `writeSentinel` degrades best-effort and lets the warning through. That is an accept-direction delta and it is in the delivering direction — base swallowed a warning the user should have seen, which is this issue's own failure class. The underlying limit is a property of the per-session filename scheme, shared by two files that predate this PR, and bounding session ids belongs to whatever writes them (`gsd-statusline.js`), not to the sentinel logic. Happy to fold a length guard in here if you would rather have it in this PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5869febb16 |
enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change readVerificationStatus() now recomputes a deterministic sha256 fingerprint over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY, requirements, implementation files in the verified change set) and returns stale on any mismatch, fail-closed when a covered file is missing, unreadable, or escapes the project root. Legacy reports with no fingerprint metadata keep the prior SUMMARY-mtime staleness check unchanged. The verifier computes covered_digest via the new verification.fingerprint CLI command rather than by hand, since a digest is deterministic math, not an LLM-estimated value. * chore(#4155): backfill fork PR number in changeset * fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap * fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior Partial fingerprint metadata (one of covered_files/covered_digest present, the other missing or malformed) now fails closed to stale instead of silently downgrading to the legacy mtime-only check. computeCoveredDigest also canonicalizes with realpathSync before re-confining, so an in-root symlink whose target escapes the project root can no longer produce a matching digest. gsd-verifier.md restores the completeness requirement and checklist item trimmed by the earlier size-budget fix, within the LARGE tier byte cap. * chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap * fix(#4155): address gemini adversarial review findings computeCoveredDigest now threads the caller-supplied opts.fs seam through its confinement and read paths instead of always using raw node:fs — a caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP 2, #2790 follow-up) was silently bypassed for covered-input reads. The project-root anchor itself still canonicalizes through real fs (it is a trusted value the caller derived, not attacker-influenced covered-input data); only per-file candidate reads go through the injected seam. Covered-file paths are now canonicalized (./ prefixes, redundant slashes, internal .. segments) before becoming dedup/sort/hash keys or confinement subjects — closes both a spurious-stale false positive (two spellings of the same file hashing differently) and a confinement gap (an internal .. segment that doesn't start the string). gsd-verifier.md now states covered-file paths are project-root-relative, not phaseDir-relative, closing an ambiguity that would have made a real verifier agent's first fingerprint invocation fail closed. defaultFsImpl's methods now late-bind through fs.<method> rather than capturing function references at module load — the earlier direct-capture form was invisible to existing tests' t.mock.method(fs, 'statSync', ...) seams, a real regression caught by the full suite (not the reviewer). * fix(#4155): catch a plan/summary added to the phase dir after verification but never declared The content digest only recomputes hashes for paths the verifier actually declared in covered_files — it had no way to notice a plan or summary added to the phase directory after verification if that new file was never declared, silently regressing behind the legacy mtime check it replaces (which scans the live directory, not a declared list). findUncoveredCurrentArtifact re-scans the live phase directory for every current *-PLAN.md/*-SUMMARY.md and requires each to be represented in covered_files, closing that gap; a directory scan failure fails closed to stale rather than silently skipping the check. CONTEXT.md's Verification Module entry corrected to describe the fingerprint path's stricter fail-closed FS-error contract (routes to stale) instead of the module's original degrade-to-safe one (missing / not-stale), which only the legacy path still keeps. * refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/ sorted covered_files independently — one shared helper now backs both (gemini review's ponytail-lens finding). Adds one CLI-to-readVerificationStatus test against a genuine .planning/phases/NN-x/ project with an implementation file outside .planning/ entirely, closing the review finding that prior #4155 unit fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot falls back to phaseDir itself there) and never exercised real multi-level path resolution. * fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/ Two independent review rounds (opus critical-reviewer + opus ponytail + agy, run twice) found two instances of the same fail-open class: - computeCoveredDigest's per-file reads routed through the caller's injected fsImpl. planning-inspect.cts passes a `.planning/`-confined containment fs into readVerificationStatus's opts.fs, so any covered implementation file outside `.planning/` (mandatory per the issue) made the confinement wrapper throw, which was caught and turned into a stale digest -- reporting every fingerprinted phase permanently stale via `planning.inspect`, regardless of actual drift. Per-file reads now always use real node:fs, matching the pre-existing treatment of root canonicalization; the realRel-vs-realRoot check is the real confinement boundary for this data and needs no seam. - allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans reports readdir failures via a `scope` field, it never throws), so an unreadable nested plans/ dir was silently treated as "zero artifacts, all covered" instead of failing closed. Now branches on scope !== SCOPE.COMPLETE. Also, per ponytail's second-round findings: reverted an unwarranted FINGERPRINT_VERSION bump and digest length-prefix from the first fix (no v1 digest has ever existed -- the feature is unreleased -- and the prefix closed a collision that grants no capability beyond what a writer of covered_files already has more cheaply); removed a verifier-facing escape-hatch instruction whose own example was a case that should trigger staleness, not bypass it; corrected CONTEXT.md references to the renamed allCurrentArtifactsCovered and a stale "unconditional" rescan claim; simplified the isStale derivation, removed dead FsLike members, and tightened test coverage. Regression tests for both fail-open bugs are included and were each confirmed to fail against the pre-fix code before the fix landed. full test suite: 2558/2560 pass, 2 skipped, 0 fail * fix(#4155): trim gsd-verifier.md under the LARGE size cap Fork CI caught what my local runs missed: the superseded/nested-plans instruction added earlier pushed gsd-verifier.md to 49299 bytes, 147 over the LARGE tier's 49152-byte hard cap (tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's wording and dropped a redundant inline comment tag; no content lost. * chore(#4155): point changeset at the upstream PR number pr: 19 was the fork PR opened for internal review-lane CI; now that open-gsd/gsd-core#4290 exists, the changeset field must match it per CONTRIBUTING.md's release-notes convention. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5ad9a36f35 |
fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id was a hardcoded literal, so `--host` chose only the argv RENDERING while the LEVEL always came from the installed plan-checker's frontmatter — `low` under every shipped model profile. Every prompt-fed lane therefore ran at a fast structural verifier's effort, and because the rendered argument is a CLI config override it silently beat the effort the operator had configured for that CLI. At `low` a large source-grounded prompt makes a model end its turn with no final message, so the lane came back empty and its stub read as a crash. Effort is a property of the review, so the lane declares it. Two new fields on ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` — carried through each capability manifest and the generated registry, set on the three lanes with an argv effort channel and null on the other nine. A new pure `resolveLaneEffort()` resolves config key -> lane default -> nothing, where "nothing" emits no effort argument at all and the reviewer CLI's own configuration decides; `inherit` selects that path explicitly and an unrecognized level falls back to the lane default rather than being forwarded to a CLI that would reject it. The host's negotiated effortSurface still gates the rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving in-process also removes up to twelve subprocess spawns per review. The empty-output stub now names the effort the lane ran at and distinguishes a clean exit from a timeout kill, a non-zero exit, and a process that never ran — `status` is null for both a timeout and a signal, so those were indistinguishable before. The hint is hedged: a clean empty exit is most often a model stopping short, but it is also consistent with a CLI writing its output elsewhere. Also: the capability validator now knows both fields, rejects a malformed key or an out-of-vocabulary default, and rejects a default declared without a config key (a level the operator could never override). An existing end-to-end row in tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures the lane's own key and pins the decoupling in the same real spawn, with the agent execution tier set to a level that must not appear. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive. Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e8800287d5 |
enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target * fix(#4153): fail closed unresolved update target * test(#4153): require a concrete recovery installer * fix(#4153): use concrete unresolved recovery command * chore(#4153): bind changeset to fork PR * test(#4153): cover portable update diagnostics * fix(#4153): keep update diagnostics portable * fix(#4153): harden update version diagnostics * test(#4153): reject jq in update version checks * test(#4153): expose step-local parser gap * fix(#4153): keep JSON parsing step-local * docs(#4153): align update target guidance * test(#4153): expose workflow runtime fallback * test(#4153): expose resolver runtime fallback * fix(#4153): leave unknown workflow runtime empty * fix(#4153): stop inferring Claude for unknown targets * test(#4153): preserve Claude workflow targeting * test(#4153): preserve known runtime directory identity * fix(#4153): recognize Claude workflow paths * fix(#4153): reuse known runtime directory identities * chore(#4153): acknowledge emitted workflow growth The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes. Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation * test(#4153): expose missing Windsurf workflow contract * docs(#4153): document Windsurf update targets * chore(#4153): bind changeset to upstream PR * fix(#4153): gate unresolved-target exit before the VERSION-missing fallback The VERSION-missing bullet in get_installed_version sat before the UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version 0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy "proceed to install" without ever reaching the fail-closed exit this PR adds, reopening the ill-defined mutating path #4153 closes. Reorder so the unresolved-target gate runs first and scope the VERSION-missing bullet to require an already-resolved target. Also drop two vacuous mutationSpies entries: they checked '--sync'/ '--reapply' (commands/gsd/update.md content) against `step`, a slice of workflows/update.md — always -1 regardless of correctness. Those routes bypass get_installed_version entirely and are already covered by install.test.cjs, reapply-patches.test.cjs, and skill-frontmatter-contract.test.cjs. * chore(#4153): point changeset pr field at fork PR #10 for fork CI * test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS (src/update-context.cts) are two independently maintained copies of the same runtime->dir mapping with no parity check; add one so a future edit to either surface without the other fails loudly instead of silently drifting. Nit 2: call out in the changeset that a custom --config-dir matching no known runtime, marker file, or env var now resolves unresolved instead of silently defaulting to claude -- this narrowing is intentional, it's the fail-closed behavior #4153 asks for. * fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path agy (gemini-3.8-flash-high) adversarial review of the full PR: 1. check_latest_version's uc_field() copy-pasted get_installed_version's `${2:-$UC}` fallback, but every call site here passes $2 explicitly and $UC does not exist in this step's scope -- dead, misleading reference. Use $2 directly. 2. No unit test covered resolveUpdateContext's preferredConfigDir fast path returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS suffix, marker file, or env var (the exact fail-closed case #4153 adds). Added. A third finding (update.md:90 using /gsd:update vs docs using /gsd-update) was investigated and rejected: /gsd:update is the actual registered Claude Code command name (commands/gsd/update.md name: gsd:update) and is locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update is a separate, pre-existing, intentional prose convention used in audience-facing docs (README/INVENTORY/FEATURES). Not a defect. * chore(#4153): backfill changeset pr field to upstream PR #4237 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a262ad6b61 |
fix(#4148): dispatch wave-pre step hooks (#4185)
* fix(#4148): dispatch wave-pre step hooks External capabilities can render step hooks before a wave, but the execute workflow consumed only contributions and silently skipped every step. Reuse the shared dispatch contract before executor spawning and pin the capability-validator boundary with a red-first regression. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * test(#4148): pin wave-pre dispatch ordering * test(#4148): pin wave-pre dispatch contract * chore(#4148): bind upstream changeset PR * chore(#4148): restore fork changeset identity * fix(#4148): align wave-pre dispatch contract Mirror the sibling wave-post all-shapes clarification while pruning redundant prose so the rebased workflow remains below its frozen byte ceiling. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * chore(#4148): restore upstream changeset identity * fix(#4148): align wave-pre capability guidance * docs(#4148): identify wave-pre manifest input Name the third-party manifest trust origin at the wave-pre dispatch boundary so the reviewer-requested validation guidance matches wave-post. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * docs(#4148): preserve execute-phase byte budget Remove a redundant advisory label while retaining the non-blocking contract, keeping the reviewer-required trust-boundary wording at the enforced 93,400-byte ceiling. * fix(#4148): mark wave-pre manifest-input validation as security-relevant Reviewer nit on PR #4185: wave-pre's step-dispatch sentence had the (third-party manifest input) parenthetical but dropped the ⚠ marker that wave-post's parallel sentence (execute-phase.md:1044) carries, losing the visual flag that this validation is security-motivated. Trims the redundant "of one" from "not one shape of one" to reclaim the 4 bytes the marker adds — the ADR-857 byte-margin gate (tests/claude-orchestration.test.cjs) leaves zero slack at the 93,400-byte ceiling. * fix(#4148): trim wave-pre step-dispatch prose to clear ADR-857 byte ceiling Merging next's unrelated growth (#3990's TDD_APPLICABLE conditional) pushed execute-phase.md 116 bytes past the 93,400-byte ceiling, failing CI on all three platforms. The security-relevant ⚠ marker and ref.command validation call-out (added per prior reviewer nit) are preserved verbatim per the pinned regression test in capability-registry.test.cjs; only the non-pinned connective prose is trimmed. * fix(#4148): recalibrate execute-phase.md self-imposed margin, restore security marker next grew execute-phase.md by ~230 bytes across two unrelated merges during this fix (#3990's TDD_APPLICABLE conditional, then a further step-extraction commit), consuming this test's own self-imposed 93,400 safety buffer under ADR-857's actual, unmodified 93,600 ceiling (docs/adr/857-capability-system.md:22). The wave-pre step-dispatch sentence cannot shrink further without dropping one of the pinned substrings this same test file asserts on (kind=="step", loop-hook-dispatch, never blocks or redirects executor spawning, Validate `ref.command`). Raises the self-imposed margin to 93,550 (still 50 bytes under the real, untouched ADR ceiling) and restores the ⚠ marker the prior reviewer round required for the ref.command validation call-out, which byte pressure had dropped. * fix(#4148): restore full ref.command validation wording, drop self-imposed margin Adversarial review (agy/gemini-3.8-flash-high) flagged two issues in the prior CI-recovery commit: 1. Trimming "in-context before any shell use" from the step-dispatch warning weakened the inline operational instruction (the reader is told WHAT to validate but not the specific in-context-not-shell mechanism the referenced loop-hook-dispatch.md:45-51 threat model requires). Restored it - the merge with next since the last commit freed enough real margin (77 bytes under the untouched 93,600 ADR-857 ceiling) to afford it without any margin change. 2. The prior commit self-imposed margin bump (93400 to 93550) was, on reflection, the wrong lever: it is a number this PR invented, not an ADR value, and re-bumping it every time next grows execute-phase.md is a losing pattern (already needed twice in one session). Removed the redundant assertion; the same line existing bytes-under-93600 check against the real, frozen ADR-857 ceiling (docs/adr/857-capability-system.md:22) is the actual invariant and is untouched. workflow-size-budget.test.cjs tier hard cap (98304 bytes, extract-not-bump by design) remains the correct backstop for runaway growth. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
d29b50d696 |
fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md * fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding * fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch. * chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names * fix(#4051): review fixes — em-dash description style, split audit-fix route * chore(#4051): sync skill mirrors of execute-phase/phase descriptions * chore(#4051): add changeset (pr backfill to follow) * chore(#4051): backfill PR 4289 in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
b1b7cabfb5 |
docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry Adds one `type: "eos"` entry for a Qoder host integration and regenerates docs/registries/eos-registry.md. Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop). The integration depends on @opengsd/gsd-core, negotiates the ADR-1239 host-integration handshake, and projects GSD's agents, skills, and hook scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the China edition), merging GSD's lifecycle hooks into settings.json. Every axis is sourced from Qoder's own docs per the never-infer rule. `dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a frontmatter-declared, per-agent-definition property, and GSD's two isolation negotiation models both assume a per-dispatch injection point Qoder does not expose. Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in favor of the EoS path. Closes #4123 * docs(#4123): backfill changeset pr field |
||
|
|
75ee7b0214 |
enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate One query verb computes TDD-applicability for a plan (CLI flag, plan type: tdd frontmatter, a task's tdd="true" attribute, or the workflow.tdd_mode config default), mirroring phase.mvp-mode's precedence-cascade shape. Foundation for epic #4272 Phase 2, which wires both dispatch backends to consume it instead of restating the predicate independently. Also fixes workflow.tdd_mode, workflow.research, and workflow.nyquist_validation, which never reached cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone because loadConfig() never populates config.workflow — a dead accessor found while wiring this verb's own config read, fixed inline per the no-defer rule rather than left alongside it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4273): document phase.tdd-applicable's FEATURES.md entry Add a docs/features/ fragment for the new phase.tdd-applicable query verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left untouched: it documents /gsd-* slash commands only, and the sibling verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI reference entry anywhere in docs/ either -- only inline prose mentions in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md precedent to extend. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): stop whitelisting capability-owned config keys centrally workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are each already owned by their own first-party capability's federated config schema (the tdd/research/nyquist capabilities declare them under their own capability.json `config`), resolved via isCapabilityConfigKey. Adding them to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as the prior commit in this branch did (mirroring workflow.mvp_mode, which genuinely is central-only), declares the same key in two places at once. That collision breaks capability-loader.cts's loadRegistry composition: gsd-test caught this as 84-85 unrelated failures across capability-cli/capability-command-dispatch/capability-lifecycle test files, every one showing "unknown capability: <id>" for a freshly-installed third-party capability that should have resolved fine. Verified directly (not asserted): reverting only this file, keeping the config-loader.cts tdd_mode/research/nyquist_validation flattening and the init.cts call-site fixes from the prior commit, and re-running the exact capability install + capability set repro from tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the failure with the whitelist entries present and clears it without them. loadConfig() still surfaces all three flattened values correctly with no central whitelist entry (confirmed directly against the compiled module) — the whitelist additions were never required for the #4273 fix to work; they were an incorrect over-application of the mvp_mode precedent to keys that aren't central. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4273): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f4bf449296 |
fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat cmdAuditUat admits `human_needed` OR `gaps_found`, but parseVerificationItems had a body only for the first and returned an empty array for the second — standing on a comment deferring to `plan-phase --gaps`, a different command audit-uat never reaches. Since cmdAuditUat pushes a file into `results` only when `items.length > 0`, a `gaps_found` report did not under-report: it vanished, taking its phase's `by_phase` row with it, so a clean-looking total gave the reader no cue anything was skipped. Eligibility now has one owner (the caller) and parseVerificationItems reports what the file says. The closed-entry filter could not be built on extractFrontmatter: its array-item parser keeps only each `- ` entry's FIRST line and has no notion of nested key/value objects, so an entry's `status:`/ `resolution:` siblings never reach its output and a closed entry is indistinguishable from an open one downstream. Rather than grow a competing object-list parser — or change extractFrontmatter, whose blast radius is every frontmatter consumer in the repo — this reads the raw segment BEFORE the flattening, via the existing anchored sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps` machinery that already parses exactly this `- `-opened, indentation- continued shape. The human_needed path is byte-for-byte unchanged: same reader, same display names, same numbering, no resolved-entry filtering — pinned by a test and verified by identical CLI output on base and head. parseGapsItems keeps its narrower `status: resolved` rule so no *-UAT.md behaviour moves. Closes #3850 * chore(#3850): backfill changeset pr number for #3879 * fix(#3850): one parse per entry, one fence parser, one resolved-entry rule Adversarial review on #3879: B1, B2, M3, m5, m8 and n9. B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell 5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found` report vanished from the audit exactly as it did before this fix — this issue's own symptom, on a platform the repo already has a named defect class for. `extractFrontmatter`'s BOM+fence logic is now factored out as `frontmatterRegion` and shared. One fence parser, not two. B2 — the resolved-entry skip paired two DIFFERENT parsers by array index: `parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A block sequence written at its key's indent — ordinary, legal YAML — makes them disagree about entry count, and from the first disagreement every index names a different entry, so an OPEN entry inherits a CLOSED one's resolution and is silently dropped. That is the defect this PR exists to fix, reintroduced inside the fix. Display name and sibling fields now come from ONE parse of the raw slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as `parseYamlRegion` does, so the string is byte-identical to what `extractFrontmatter` produced. The flattened array remains the #2286 GATE, but is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment. M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both readers use, rather than two copies differing only in `result`. m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an acceptance criterion #3850 does not contain: the issue has no AC section, and its suggested fix (2) states the skip unconditionally, naming a file with 14 of 16 entries resolved. That file is `human_needed`, so the asymmetry left the reporter's own scenario over-reporting by 14. m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and says the column-0 boundary rule is now a cross-module contract. n9 — the vestigial bare block is gone and its body de-indented. Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF fixture (M4 — it survived by accident, now pinned) and the unified skip rule. Fail-first verified by running the new tests against the pre-fix build: the BOM, nested-sequence and unified-skip cases are red there. * fix(#3850): read the entries as objects, not as re-parsed display text Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1 (#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml: `parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry now flattens to `test: A, resolution: R` rather than to its first line. The original mechanism existed ONLY to work around that lossy first-line flattening — it sliced the raw frontmatter segment and re-parsed each entry by hand so a `resolution:` sibling was visible at all. With a real parser upstream that workaround is obsolete, so it is deleted rather than repaired: `sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the `splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are all gone. `frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` — the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence, same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping one step before the display flattening. `flattenObjectListItem` is exposed alongside it so a caller deriving a display name produces the byte-identical string `extractFrontmatter` would have. That collapses the review's blockers into properties of the parse rather than things this fix has to get right: - B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI. - B2 (index pairing) — there is no second reader. Display name and sibling fields come from one object. - M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers. - M4 (CRLF) — js-yaml's, not ours; verified through the CLI. Also confirmed on the rebased base, per review: #3850 still reproduces on `next` after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found` fixture), so this PR is still doing work #3707 did not do. Nothing was dropped as redundant. One behaviour note: `entryField` returns a present value verbatim and treats only whitespace-only as absent. Trimming would rewrite an author's `truth:` on its way to becoming the display name. * fix(#3850): keep every frontmatter list entry at its own row Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result to objects, and filtering COMPACTS: `parseHumanVerificationItems` then numbered the survivors by their position in the compacted array. On a list mixing object and non-object entries the non-object rows disappeared outright and the rest were renumbered — #3850's own vanishing-row defect, reached through entry SHAPE instead of file STATUS. Base never had it: it walked the display array, so every row surfaced at its own position. Renamed to `frontmatterListEntries` and it no longer filters (the name now matches what it returns). Deciding what a non-object entry MEANS is a caller's judgement; dropping it is nobody's. Both readers now walk the DISPLAY array — one element per row, the array #2286 already gates on — and consult the parsed array only for "does this entry carry a closure field?". `parsedEntriesFor` owns that pairing and checks the two lengths agree before trusting an index; all-null is the correct degradation, since over-reporting a closed row is recoverable and closing the wrong one is not. Names stay byte-identical to base for every entry shape, including a nested sequence (`[nested]`, not `["nested"]`). Same class closed in the gaps reader: a non-object `gaps:` entry surfaced nothing at all and now surfaces as `unknown`, which is this module's documented fail-safe direction (`parseGapsItems`) on a false-negative bug. Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule inlined twice; `extractFrontmatter` now routes through it, so "one fence parser" is enforced rather than asserted in a comment. Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and its "shared by both readers" comment corrected — it has one call site, and the two readers differ deliberately, each mirroring its own established sibling (`parseGapsItems` vs #2286). Documented at the divergence. Tests: `B2` asserted a name substring, so it passed while the row was mis-numbered and would have passed through outright loss; it now asserts positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim, B2c the survivors' file positions across skipped rows, B2d the gaps reader. All four fail-first against the reviewed head; 332/332 green with the fix. * fix(#3850): make status authoritative, and let the two gaps readers agree Round 4 review, all five findings. Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as closure regardless of `status:`, so `status: failed` + `resolution: "attempted retry, still failing"` vanished from the report — the silently-vanishing-item defect #3850 exists to close, reached by field combination instead of file status. Closure is now per key, because the two keys have different conventions and one rule cannot serve both: `gaps:` `status: resolved` only, byte-identical to the rule `parseGapsItems` applies to a `## Gaps` markdown section, so one authored entry cannot read closed in one reader and open in the other. `human_verification:` a bare `resolution:` still closes, since that is how verifier-written entries record it — but a readable `status:` that contradicts it wins. A single unified rule was the first draft and is wrong: it closes a frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which `parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring claims it mirrors that reader's fail-safe status handling. The contradiction guard is not a judgment call about YAML. It is the rule this codebase already applies to the same field pair: `validateResolution` (probe-core.cts) rejects a populated `resolution:` on a non-resolved status outright — "a populated payload is an authoring mistake ... Reject it so the mistake surfaces." A reporter cannot throw, so it surfaces the item. Minor 1. Direct unit tests for `frontmatterListEntries` and `flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that historically co-changes with `frontmatter.cts`. They were reachable only through `uat.cts`' readers before. Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted directly. Verified unreachable through content rather than assumed: both readers enter through `frontmatterRegion`, `extractFrontmatter`'s only extra argument gates a warning, and `normalizeParsedValue`'s `value.map` is 1:1. It is a drift alarm for a future edit to either parser, so the helper is exported for tests rather than left as the one unpinned branch. Minor 3. The vestigial `const skipResolved = true` and its dead conditional are gone. Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:` entry has no `test:` in its vocabulary — the template's entries carry truth/status/reason/artifacts/missing — so it was speculative support for a field the shape does not have, and it collided with the 1..N row numbers `parseHumanVerificationItems` assigns by array position. Not reading it makes the collision impossible; an offset would have rewritten an authored value, against `entryField`'s verbatim contract. Docs, changeset and the dispatcher docstring all stated the unconditional rule and are corrected — three prior rounds here were comment/code drift. Fail-first proven: restoring the universal rule reddens all three new unit tests and both rewritten properties. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
585a8b7f1b |
fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity * fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path * fix(#3747): scope doc evidence claim to skills discovery per adversarial review * chore(#3747): add changeset * chore(#3747): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
1fe85cd43e |
chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk Repo-wide sweep (ahead of adding lint rules for these exact bug classes) found both incident patterns still live and unfixed on `next`: - scripts/run-tests.cjs's sweepProtectSet walk stopped on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. win32 dirname('D:\') is a fixed point (length 3, never satisfies `> 1`... wait, it does satisfy length>1), so a selected file living outside runTempRoot (the common case) spins the walk forever on Windows. Extracted a pure, exported computeSweepProtectSet helper that terminates on dirname(cur) === cur instead, with in-process RuleTester-style coverage for both win32 and posix paths. - tests/run-tests-temp-root.test.cjs's own #4020 regression test set only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never reads TMPDIR on Windows (only TEMP, then TMP), so the redirect silently no-oped there — masked because Windows CI died in the dirname-walk hang above before ever reaching this test. - tests/config-schema.property.test.cjs's fallow config-set test had the same TMPDIR-only pattern, direct process.env assignment this time, restored in its own finally block. Origin: #4220 and its shared root cause #4020. * feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for either shape. - local/require-full-tmpdir-triad: flags a TMPDIR environment override (direct process.env.TMPDIR assignment, or a TMPDIR property in a spawn-like call's env: object literal) not accompanied by TEMP and TMP in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows. Registered on tests/**/*.cjs, matching the require-userprofile-with-home precedent. - local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning from dirname() with no fixed-point termination guard (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a no-op at the platform root, but the value differs by platform (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped length/equality bound never fires on Windows. Registered on BOTH tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in scripts/run-tests.cjs, not tests/. Both rules join the zero-escape-hatch discipline already established for this catalog (no bespoke comment marker; PROTECTED_RULES in tests/portability-rule-disable-ban.test.cjs independently bans eslint-disable of either). ADR-1703 and its two companion contributing docs get an amendment documenting the mechanism, code examples, and the repo-wide sweep (three live instances found and fixed in the prior commit; no others found). CI test-scope selection updated so an edit to either rule or to scripts/run-tests.cjs re-runs the right suites. * fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too checkWhile bailed out early unless node.test was a LogicalExpression, so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } -- was silently skipped and never reported. That is the EXACT minimal shape of the original #4020/#4220 bug, and it is literally the shape used by this rule's own shipped RuleTester fixtures (the "equality-only bound" invalid cases), which were failing (0 errors reported, 1 expected) until this fix -- confirmed by running RuleTester directly against both fixtures, not just via a passing test-runner exit code. The conjunct-collection helper already handled a non-LogicalExpression test correctly (it pushes a single node as the sole conjunct); only the early-return gate needed to stop requiring a compound && / || test. Verified: RuleTester run directly against both previously-broken fixtures plus two new sanity cases (a guarded single-condition loop stays valid; an unrelated single-condition loop stays silent), and a fresh `npx eslint .` across the whole repo remains clean (no other single-condition dirname-walk shape exists in the tree). * fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call isSpawnLikeCallee only recognized a MemberExpression callee (child_process.spawnSync(...)) or a bare identifier in ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare -- const { spawnSync } = require('child_process'); spawnSync(...) -- has an Identifier callee named "spawnSync", which matched neither branch, so the whole env-literal check was skipped. gsd-test caught this: both "invalid: child_process.spawnSync with TMPDIR-only env" cases in tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors reported, 1 expected). Widened the bare-identifier branch to also match any of the known ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same lightweight convention this repo's other eslint-rules/*.cjs use (e.g. no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow tracing. Verified: RuleTester run directly against all 11 cases in tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that were failing), all pass; a fresh npx eslint . and npm run lint:ci across the whole repo remain clean. * fix(#4244): correct a stale escape-hatch reference in a test comment The comment on the "length comparison against another expression's length" case referenced a "// allow-dirname-walk marker" that doesn't exist -- the rule has zero comment-based escape hatches by design (ADR-1703), and an earlier draft's marker mechanism was removed before this branch's first commit. Spec-axis review caught the stale reference. No behavior change; comment-only. * chore(#4244): backfill changeset PR number (pr:0 -> pr:4246) --------- Co-authored-by: sim <sim@local> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
91ed46882a |
feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives Adds the full behavioral (tests/quick-batch.test.cjs) and property-based (tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core primitives per the phase's 35-row test matrix — task-list parsing (inline + --file, with path-confinement/symlink-escape/non-regular-file rejection), collision-safe quick-id preallocation under withPlanningLock, BATCH.json schema/validation/resume, dependency-DAG + partitionByFileOverlap wave construction, and exactly-once STATE.md completion (including the STATE-row-written-but-manifest-not-yet-updated crash window). The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a not-yet-written src/quick-batch.cts) does not exist yet — every test in both files fails at the top-level require() before any assertion runs. Five fast-check properties cover collision-freedom under lock contention, resume idempotency, exactly-once STATE completion, wave totality, and DAG-respecting wave order, per the design doc's property-based-coverage requirement. * feat(#3675): implement quick-batch core primitives Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core primitives per the phase design lock — pure/state primitives and CLI-testable core operations only, no agent dispatch, no worktree creation, no user-facing command (Phase 4/#3676's job): - parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list parsing (>=2 items required) and a --file variant strictly confined to the planning workspace root via requireSafePath, rejecting non-regular-file targets. - allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id preallocation under withPlanningLock, checked against both on-disk .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json manifests (never on-disk-only, which would miss another in-flight batch that hasn't dispatched any real quick directory yet) — replicates cmdInitQuick's own grammar rather than delegating to it (that function's 2-second granularity is not batch-safe). - computeWaves: deterministic wave construction combining dependency-DAG layering with partitionByFileOverlap (#3674), called per DAG layer over path-separator-normalized planned_files — normalization happens at this module's boundary, never inside the Phase 2 helper. - loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated JSON, wrong types, missing fields, out-of-batch dependency references, dependency cycles, a worktree path absent from disk). - resumeBatch: skips complete items, never auto-retries failed items, propagates/reverses blocked status along the DAG to a fixed point, and detects a STATE.md row that already exists for a non-complete item (the "STATE written, BATCH.json not yet updated" crash window) — completing it without re-appending. Idempotent across repeated calls. - completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion — appendQuickTaskRow (unmodified) is called at most once per quick id, gated by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks Completed" table, since appendQuickTaskRow itself carries no idempotency. BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling of .planning/quick/ — never inside it, so scanQuickTasks never misreads a batch manifest as a broken quick task. * docs(#3675): register the new quick-batch module New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row plus the regenerated docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary entry matching the convention set by the sibling File Overlap Partitioner Module (#3674) entry it sits beside. NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically sorted single-entry insertion into families.cli_modules, matching the existing file's structure) rather than via `node scripts/gen-inventory-manifest.cjs --write` — this session's MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script path from Bash, with no available Memtrace tool to route through instead. The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs --check` to confirm this hand-edit is byte-identical to the generator's own output before merging. * fix(#3675): resolve lint findings in quick-batch primitives and tests Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color array, two unnecessary `as string[]` casts TS 5.5's inferred type predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch` import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a CONTEXT.md glossary illustration that looked like a real file reference. * feat(#3675): close acceptance-criteria gaps found in review Standards- and spec-axis review (plus a self-caught race) surfaced real gaps against issue #3675's own acceptance criteria and this repo's test conventions: - BATCH.json was missing options, base_revision, per-item wave, and per-item commit — the issue's AC explicitly lists all four as things the manifest must track. Added them: createBatch persists caller-supplied batchOptions/baseRevision verbatim and assigns each item its computed wave index; completeQuickItem now persists the commit onto the item, not just the STATE.md row. All four are backward-tolerant on load (an older/hand-built manifest without them still validates). - resumeBatch had no "incompatible base divergence" check at all, despite the AC and the ADR's own "Base divergence" section requiring one. Added an opt-in currentBaseRevision comparison that fails closed with a recoverable diagnostic on mismatch, and touches nothing on refusal. - resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the only durable write path in this module that wasn't lock-protected, a real lost-update race against a concurrent completeQuickItem or another resume. Now runs inside the same lock createBatch/ completeQuickItem use. - loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no size cap (security review, Low/informational); switched to the existing safeJsonParse (1MB cap) for defense-in-depth. - Parser (parseTaskList) had only example-based tests; CLAUDE.md requires a fast-check property test for parsers. Added one plus a companion reject-property for <2 items. - The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for direct limit-1/limit/limit+1 testing without needing 46k fixture dirs. - Issue AC explicitly asks for prompt-injection-payload test coverage, distinct from the existing shell-metacharacter test; added one. - Test row 9 (FIFO skip) silently returned instead of calling t.skip(), so an unsupported platform would report a pass rather than a documented skip; fixed to bind the test-context param and skip properly. - Extracted toWaveInput to remove a 2-site production duplication of the QuickBatchItem -> computeWaves reshape (Standards-axis smell). - Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/ is user-facing even though the compiled .cjs is gitignored). * fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason gsd-test caught this: switching loadBatch to safeJsonParse changed the parse- failure message shape ("... parse error — ...") without preserving the "not valid JSON" substring row 27's own test asserts on. Re-wrap safeJsonParse's error into the original diagnostic phrasing regardless of which of its three failure modes fired. * docs(#3675): backfill changeset pr number to 4190 * fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one CI caught this on windows-latest: row 9's platform-skip only caught mkfifo throwing (command not found). On this runner mkfifo resolves to something that exits 0 without creating a file (NTFS has no FIFO concept), so execution fell through to parseTaskListFromFile against a path that doesn't exist, producing an ENOENT stat error instead of the expected "not a regular file" rejection. Check the artifact actually exists before trusting a zero exit code, and skip with a documented reason either way. --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
77dcdda534 |
enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4) * fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4) * test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift src/core-utils.cts's new #4014 import block shifted every subsequent line by 6, moving generateSlugInternal's real closing brace from line 193 to 199. tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that line number to plant a synthetic violation immediately after the function's real body; the guard script itself locates the boundary dynamically via brace-matching and needed no change. * docs(#4014): document the unreadable-directory scope signal and add changeset * docs(#4014): backfill changeset PR number to #4163 * test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff --------- Co-authored-by: sim <sim@local> |
||
|
|
b0572c0108 |
feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage golden script) as a regression safety net ahead of extracting partitionStages into a standalone module. Also adds the new module's unit and property tests (test matrix rows 1-11) against its expected public API, which does not exist yet and is added in the next commit. * feat(#3674): extract file-overlap partitioner into a shared, generic module Moves partitionStages' greedy first-fit file-overlap algorithm into a new, dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap), generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own Plan[] shape onto the generic input and back — behavior-preserving, no dependency ordering, no path normalization, no filesystem access moved or added. Enables a future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive without pulling in orchestration internals. * docs(#3674): register the file-overlap-partitioner module bookkeeping New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row (regenerated via gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry matching the convention set by similarly-scoped leaf modules (text-lines.cts, plan-dependency-graph.cts, spec-section.cts). * fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct - docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write, produced no diff (manifest is keyed by content, not row order) - tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is: tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier. * fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction The `no two plans in the same stage share a modified file` property reconstructs which physical item produced each output id via `remaining.findIndex(r => r.id === id)`. Under duplicate ids (an explicitly-supported input shape for `partitionByFileOverlap`) that reconstruction can pick the wrong physical occurrence, producing a false-positive overlap failure (observed counterexample: p0(f1), p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but misread by id-order as [[p0,p208#1],...], which do overlap). Properties (a) determinism and (b) totality already exercise duplicate ids correctly and are left unchanged; only this property's generated items are now constrained to unique ids via `fc.uniqueArray`, where the reconstruction is unambiguous. --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
7960374d15 |
fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git Behavioral coverage: extract the trailers:key token from ship.md and prove a real git commit carrying that trailer reads back via %(trailers:key=token,valueonly). gate_status contains an underscore, which git's trailer machinery cannot tokenize, so the audit read was structurally empty. * fix(#3962): rename the TDD-Audit trailer token to gate-status Underscore is not a valid git trailer token character, so %(trailers:key=gate_status,...) could never match. Renames the token at the read, the documented aggregate write, and the section's prose/table header. Self-suppression semantics (#2431) unchanged. * test(#3962): compare outcome against the seam's literal, fix doc token Review follow-ups: the round-trip test compared outcome to 'EXITED' but the process seam emits 'exited'; and the trailer rename is propagated to docs/ship-pr-body-sections.md's read/write examples. * chore(#3962): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
91d5fdff6f |
chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse advisory output alongside the existing additionalContext prose: - gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName - gsd-context-monitor.js: severity ('warning'|'critical') - gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523 pattern) - gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its findings array already existed from #3523) - gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg, distinct from the existing force-add block leg's code additionalContext stays byte-identical in every hook (verified per-hook against the pristine HEAD version across a spread of payload shapes). Migrates all 20 assertion sites named in the issue off additionalContext.includes(...)/assert.match(...) substring-matching onto the new typed fields, per CONTRIBUTING.md's prohibition on raw text matching on test outputs. Closes #3546 * test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's own doc comment says is sized for plumbing reads (rev-parse/branch/log), not write-heavy add/commit spawns. That file already documents the exact same defect class from a prior incident (PR #3323) and exports GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites - commitN just wasn't using it. Observed failure: `git commit -m filler 9` timed out under normal bench load, unrelated to any of this PR's own diff (hooks/*.js + 5 other test files). Not a flake: root-caused to the timeout bound being sized for the wrong call class, per this repo's no-flakes rule. * chore(#3546): backfill changeset PR number (#4167) --------- Co-authored-by: sim <sim@local> |
||
|
|
ff0361071d |
feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane cursor-agent exposes --model (204 selectable models) but the cursor lane declared modelArg: null / modelConfigKey: null, so review.models.cursor was rejected as an unknown config key and the #1517 reviewer-instances escape hatch silently discarded a configured model at modelExpansion. Wire the lane the same way codex already is: inject {{model}} into args right after -p, set modelArg to --model, and declare modelConfigKey as review.models.cursor plus its config schema entry. An unconfigured lane still invokes byte-identically to today. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): add changeset for review.models.cursor Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3653): update co-change surfaces that assumed cursor has no model key gsd-test surfaced three surfaces still hardcoding "cursor declares no modelConfigKey", broken by wiring review.models.cursor: - tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797 invariant test listed cursor among lanes that must own no model key. - tests/settings-integrations.test.cjs: the #3651 keyless-lane test listed cursor as keyless, including a live config-set assertion that now correctly succeeds instead of failing (swapped to qwen). - gsd-core/workflows/settings-integrations.md: the settable-keys enumeration and two prose call-outs still named cursor as keyless. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): acknowledge deliberate growth of settings-integrations.md settings-integrations.md grew 4 bytes because it now enumerates review.models.cursor as a settable key alongside the other reviewer lanes, matching the modelConfigKey wired for cursor in this PR. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer The previous commit's trailer was separated from Co-Authored-By by a blank line, splitting it into an earlier, non-trailer paragraph — git's trailer parser only recognizes the last contiguous block. Restating it here immediately adjacent to Co-Authored-By so both parse as trailers. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): backfill changeset pr number pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
4dfc46bbe7 |
enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate * feat(#3348): add context-drift pre-check gate for plan-phase Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md effective last-changed time (git commit time, falling back to mtime for uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently reuses an upstream artifact that predates a decision added to CONTEXT.md after that artifact was derived from it. Deterministic, no model call. New `gsd_run verify context-drift <phase>` command, sibling to the existing verify.codebase-drift/verify.schema-drift gates in the drift capability. Warn-only by default (workflow.context_drift_precheck), with an opt-in workflow.context_drift_action: block escape hatch. Wired at plan:pre in plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions. * fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution * fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe The #3859 empty-diff guard decides whether a submodule bump would land using `--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules` config. The real `git commit -- <paths>` that follows was never given the same override, so under a bare `diff.ignoreSubmodules=all` repo config the two calculations disagree: driven on git 2.39.5 (Debian bookworm, the linux-node24 test-matrix image), the guard correctly stands aside but the scoped commit itself then silently fails (exit 1, no error text) for a gitlink bump it had just confirmed would be recorded, surfacing as commit_failed instead of committed:true. Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the probe and the commit it protects can never diverge. Harmless when no submodule path is involved (driven: identical outcome on an ordinary scoped file, with and without the flag). * fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal * fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a current-milestone enumeration). This PR's refactor pass lifted that block into a shared helper, resolvePhaseDirByToken, also used by the new cmdVerifyContextDrift — the guard tracks exemptions by enclosing function name, so the readdirSync now lives in an unexempted function and started firing. Move the exemption to resolvePhaseDirByToken (same written reason, now covering both callers) instead of migrating to listAllPhaseDirs, which would introduce two real behavior deltas here: it catches readdirSync failures internally (old code let them throw) and sorts results by phase number before matchPhaseDirs picks matches[0] (old code used raw, OS-dependent readdirSync order). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): satisfy lint:ci — slash form, capability registry regen - docs/features/context-drift-gate.md used the deprecated /gsd: colon form; docs are never passed through the install-time slash-form converters, so lint-docs-command-form requires the hyphen form. Regenerated docs/FEATURES.md from the corrected fragment. - Regenerated gsd-core/bin/lib/capability-registry.cjs after editing capabilities/drift/capability.json (lint:generated-sync). * fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit` subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for every scoped commit call, breaking 17 position-based assertions in the commit-files pathspec regression suite that read `a[0] === 'commit'` to find the commit invocation among recorded git calls. `execGit` already accepts an `env` option merged onto `process.env` before spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0` as a per-invocation config override functionally identical to `-c key=val`, expressed via env instead of argv. Passing that env alongside the existing commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged) fixes the real commit's effective diff.ignoreSubmodules to match the empty-diff guard's probe without moving anything in argv position 0. Scoped to exactly the canScope branch, matching the probe's own preconditions and leaving no behavior change for commits the probe never evaluated. No test file changes needed — the 17 previously-failing assertions test argv[0] against the array passed into execGit, which never changes. * fix(#3348): register verify-context-drift in the check subcommand router The drift capability's new plan:pre gate declares check.query "verify.context-drift", which normalizes to `check verify-context-drift`, but no such subcommand was routed — phase6-capstone-conformance's uniform-block-field test failed with "Unknown check subcommand" for every declared gate query. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot of the drift capability's config keys. #3348 legitimately adds two new keys (workflow.context_drift_precheck, workflow.context_drift_action) for its own plan:pre context-drift gate — extend the expected set (and clarify the assertion message) without weakening the test's exactness. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction #3348 (an earlier commit on this branch, e4b80ad81) extracted cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block into the shared resolvePhaseDirByToken helper (also used by the new cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's function-scoped exemption from cmdVerifySchemaDrift to resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer contains a line the guard's detectors match, so it needs no exemption. tests/phase-locator.test.cjs's E2 test still pinned the exemption to the old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to match the same "migrated call site's exemption must move, not duplicate" pattern the test already applies to cmdRoadmapAnalyze and cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the still-exempt list and add symmetric assertions that it no longer carries the exemption while resolvePhaseDirByToken now does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs 'always exits 0 (query command contract)' included the no-phase-arg case, which contradicts the file's own earlier 'errors with usage message on missing phase arg' test (that case legitimately exits 1 via the Usage error) — drop it from the always-exits-0 cases. 'degrades to mtime comparison outside a git repo' and '...in a repo with no commits' relied on real wall-clock ordering between two back-to-back writeFileSync calls to prove CONTEXT.md is newer than RESEARCH.md; on a fast filesystem both can land in the same mtime tick, producing a tie that computeContextDrift's strict `<` correctly treats as not-stale, so stale_artifacts comes back empty. Make both tests deterministic via explicit fs.utimesSync instead of relying on timing (CONTRIBUTING.md: never assert elapsed wall-clock time). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture The "all plan:pre when-keys false" fixture explicitly disables every known workflow.* plan:pre toggle, but didn't yet know about the new workflow.context_drift_precheck key (defaults to true), so the new drift context-drift gate stayed active and broke the empty-activeHooks assertion. Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3348): backfill changeset PR number (pr:0 -> 4147) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c0fd2e3f4c |
feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route Extends tests/host-integration.test.cjs with negotiateHostCapabilities maxConcurrency negotiation coverage (test matrix rows 1-15, including a fast-check property test) and a new #3673 dispatch-capacity CLI route describe block spawning the real gsd-tools.cjs (rows 16-25). Extends tests/host-integration-validator-parity.test.cjs with an all-19-descriptor maxConcurrency presence/validity sweep (row 26) and adds a hostile-input validator test to host-integration.test.cjs (row 27). The dispatch.maxConcurrency field does not exist yet, so these tests fail. * feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration Interface (ADR-1239 Phase 1), following the existing dispatch.isolation sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES floors, and a negotiateHostCapabilities branch that passes through a positive safe integer and fails closed to 1 otherwise (no engine-side reduction, per the design doc's explicit rejection of a min(host,engine) rule). capability-validator.cjs gains parity validation for the new field (optional, positive safe integer or the "undocumented" sentinel — mirroring isolation's "added after existing descriptors" treatment). gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling of `dispatch-isolation` with no side effects: live env (GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence. All 19 capabilities/*/capability.json descriptors now declare dispatch.maxConcurrency: claude carries the one cited value (20, per code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented" (not yet researched for this axis). * docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix Updates docs/reference/host-integration-interface.md's dispatch struct entry (also backfilling the previously-undocumented isolation/backgroundDispatch sub-fields found stale in the same table) and adds fail-closed/live-transport precedence prose for the new maxConcurrency field. Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in docs/reference/host-integration-capability-matrix.md — required for tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check, which asserts every declared dispatch sub-axis is documented there. Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist and example descriptor block to mention maxConcurrency (and, likewise backfilling a stale gap, isolation/backgroundDispatch). * fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check Exports isPositiveSafeInteger from src/host-integration.cts as the single source of truth for the dispatch.maxConcurrency positive-safe-integer contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity now both call it instead of independently reimplementing the same predicate. Also removes the __proto__/constructor/prototype reserved-name branch from capability-validator.cjs's maxConcurrency check — copy-pasted from the string-enum fields above it, but unreachable for a numeric field (the generic positive-safe-integer branch already rejects any string) and absent from maxDepth, the field the code's own comment claims to mirror. --------- Co-authored-by: sim <sim@local> |
||
|
|
b848b23861 |
feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md - Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's existing render + generic contribution dispatch pattern - Inject planner-targeted contribution fragments into the planner prompt, after AGENT_SKILLS_PLANNER, matching D-08 ordering - Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions - Record quick.md's byte-growth rationale in this commit trailer for every gsd-core-verbatim runtime Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does. * test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call < agent-skills placeholder < injection block < spawn ordering) assertions, all extracted via a brace-bounded slice anchored on the literal injection instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER} and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain brace pairs). Add a capability-registry.test.cjs describe block proving the registry-wide D-07 exclusion is meaningful: at least one plan:pre contribution exists, every plan:pre contribution has a non-empty into/fragment.inline, and the registry as a whole carries at least one non-planner-into contribution. Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's real content, and buildContract() still yields exactly 5 entries. Verified red-without-Task-1 by temporarily reverting quick.md to its pre-f30de9cc content and re-running these three suites (D-08 failed as expected), then restored via git checkout and re-confirmed green. * docs(#3778): note quick planning also renders plan:pre in the tutorial The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the plan:pre hook set. Since quick.md now dispatches the same hook set (f30de9cc), the sentence understated the capability's real reach. * feat(#3778): add changeset fragment * chore(#3778): reference the upstream issue in the changeset fragment The fragment was the only one of 81 in .changeset/ without a trailing (#NNNN) reference or a bold lead-in. serializeChangelog auto-appends only the pr: field, so the rendered CHANGELOG entry carried no link back to issue #3778. * test(#3778): scope the D-07 registry assertion to what it actually proves The registry-wide non-planner check was named "D-07 exclusion is meaningful", which overclaims: it proves only that `into` takes non-planner values somewhere in the registry, not that anything is excluded at plan:pre. Every plan:pre contribution is currently into: "planner", so the filter is a forward-looking safeguard there. Narrowing the assertion to plan:pre (as review suggested) would fail today. Asserting plan:pre is all-planner would be brittle — it would break the day a legitimate non-planner plan:pre contribution lands, which is exactly when the safeguard starts doing work. So the assertion is unchanged and only the name and comment are corrected. * chore(#3778): point the changeset fragment at the upstream PR The fragment carried pr: 3, the fork staging PR. changeset lint derives the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that would read as pr-field drift. Point it at open-gsd/gsd-core#3934. * test(#3778): require contributions in Quick revision prompts * test(loop-host): require Quick auxiliary registration * fix(#3778): preserve contributions in Quick plan revisions * fix(#3778): validate Quick as a planner contribution host * fix(#3778): tighten Quick contribution contract * test(#3778): drop unnecessary source-contract exemption * fix(#3778): require Quick planner target coverage * docs(#3778): describe targeted auxiliary coverage --------- Co-authored-by: davdittrich <davdittrich@gmail.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7ca5479f97 |
docs(#3672): amend ADR-1239 for quick-batch concurrency and single-writer contract (#4146)
Locks the architecture gate the maintainer required before any /gsd:quick-batch implementation PR (epic #3344, condition #2): canonical terminology, the foreground-coordinator single-writer invariant over BATCH.json/STATE.md/roadmap metadata, the dispatch.maxConcurrency sub-field (fail-closed to 1), the GSD_DISPATCH_MAX_CONCURRENCY live-capacity transport and its precedence over the descriptor value, the gsd-tools query dispatch-capacity contract, the dispatch.isolation interaction, and backpressure/recovery/deterministic-merge rules. Docs-only; closes its own Phase 0 sub-issue, not the epic. Co-authored-by: sim <sim@local> |
||
|
|
bdfc62889b |
fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape advancePlanCore derived the value FORMAT from the field NAME, so it handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the compound `Plan: N of M`, but not the hybrid of the two: the legacy field name carrying a compound value with no Total Plans sibling. `legacyTotal` is null so the legacy branch fell through, and the compound branch reads the `Plan` field through a `^Plan:`-anchored pattern that never matches `Current Plan:`. Both produced NaN against a file whose plan numbers are plainly readable. The shape is not exotic. An agent wrote it unprompted into a project's STATE.md, believing it was the parseable form, and every subsequent run in that project inherited the failure and worked around it by hand. Track the field name and the value shape separately (`planSourceField`, `planRawValue`) so write-back targets whichever field the value came from. The legacy pair still takes precedence when both fields exist, so a stray "of N" inside Current Plan cannot override an explicit Total Plans — covered by a new test. Also replace the caller's catch-all error. It reported "Cannot parse Current Plan or Total Plans" for ANY transition failure, and named no accepted shape, so a reader learned neither what failed nor what to write. It now distinguishes "no result" from "unreadable plan position" and lists all three shapes. The existing test asserted the literal "cannot parse"; it now asserts the message names the shapes, which is the property that makes it actionable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): keep zero-padding when advancing a compound plan value The compound write-back rewrote only the leading half of "N of M", so a padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic on its own, but a plan line that looks wrong is one the next writer tidies by hand, and hand-tidying this particular line is what produced the hybrid shape the previous commit had to teach the parser to read. Pad the incremented number to the width it was written with. padStart never truncates, so a value that outgrows its padding widens correctly: 09 of 12 advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances to 3 of 6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): pass a literal field name to the compound write-back The previous commit passed `planSourceField` — a variable — as the field-name argument to `stateReplaceField`, which trips the state-write-path drift guard's `unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care: a Title-Case literal cannot collide with a lowercase or snake_case frontmatter key, so it is safe whatever the content argument is, while a variable could hold anything and therefore requires its content to be demonstrably frontmatter-stripped first. The content argument here IS stripped — `body` is `stripFrontmatter(content)` — but the guard does a narrow backward scan rather than dataflow tracking, by design, and the nearest preceding assignment to `body` is another `stateReplaceField` result. Rather than baseline a bypass or ask a future reader to re-derive that the invariant holds, dispatch on the discriminator and pass the literal. Guard goes from 1 finding to 0; its own 32 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(3784): add changeset fragment for #3785 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring Triage published six acceptance criteria; two were only half-covered. AC1 asks that the hybrid write back to the SAME field with padding preserved. The existing hybrid test used an unpadded value and asserted only `result.data`, so it proved the parse but never the write. Now asserts the written content is `05 of 06` on the original field, and that no separate `Plan:` field appears as a side effect. AC6 asks that the shared field reader not be loosened. Reading the hybrid is the transition's job; `stateExtractField('Plan')` is line-anchored and has 13+ callers, so teaching it to match a name merely ENDING in "Plan" would be the wrong fix and would silently change what those callers read. This holds by construction here — the reader is untouched — but nothing locked it in. The new test fails if anyone later reaches for that shortcut. Also drops the changeset fragment written against the auto-closed PR number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(#3784): add changeset fragment for #3791 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(#3784): write the advanced plan back to the field it was read from Review findings 2-6 on #3791 were one defect seen from several angles: the read path learned the hybrid `Current Plan: N of M` shape, the write path did not follow it. - `bumpLeadingNumber` now owns the increment for all three parse branches. Only the leading digits belong to this transition; the padding width and everything after it (` of M`, and the `\r` of a CRLF file) are the author's text and are preserved. The legacy branch wrote `String(newPlan)`, which turned `2 of 99` into `3` and `04` into `5`. - `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm only ever looked for `Plan:`, so on a hybrid file the `## Current Position` section was never reached; combined with the body-level write being single-shot and bold-preferring, a file carrying the field at both sites advanced the header and left the section a plan behind. The parameter defaults to `Plan`, so the two callers that pass no plan are unchanged. - Tests: both-sites-advance (fails without the section arm), legacy write-back content assertions (the previous test read only `data` and so could not see the lossy write), hybrid boundary at limit-1 and limit+1, a CRLF fixture, and an fc property pinning the padding-width contract. Two characterization tests pinned `**Current Plan:** 02` advancing to `3`. That dropped padding is the defect #3784 reports, so the expectation is corrected to `03` rather than the fix being narrowed around it. * fix(#3784): drop the unreachable advance-plan error branch, sync the doc Findings 1 and 8 on #3791. The `!resultData` arm could not fire: the transform callback assigns `resultData` unconditionally, only runs once STATE.md is known to exist (the missing-file case returns "STATE.md not found" upstream), and every `advancePlanCore` return path sets `data`. It was a speculative second failure mode with a message no caller could receive, and the comment beside it claimed to distinguish two things that were never two. `!resultData` stays in the condition as a type guard, which is all it ever was. `docs/json-errors.md:142` quoted the old error literal verbatim and was the sole occurrence in the tree; it now quotes the emitted one. * chore(#3784): describe the write-back fix in the changeset * fix(#3784): anchor the plan grammar and widen the schema row to match Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit. B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to `['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs --check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not surface this row, so there is nothing stale to regenerate. B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as `4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE `Status: Phase complete — ready for verification` into the user's file. Both shapes are now anchored at the start and every number comes from a capture group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER` rather than letting `data` and the persisted string disagree. Nothing on this path calls `parseInt` on a raw field value any more. The grammar keeps a trailing remainder after the total, because `Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from requiring `of <total>` to follow the leading number immediately, not from forbidding a suffix. M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there were no leading digits, so `+2` reported `advanced: true` while writing the file untouched. M2 — both section arms use replacer functions. File-derived text was being spliced into a `String.replace` replacement string, where `$&` / `` $` `` / `$'` expand: a value of `04 of 06 $&` spliced part of the document into itself. `stateReplaceField` already used a function; these now agree with it. M3 — the section arm targets the name the SECTION carries, and the body write now writes both spellings, each with its own rendering. Keying off the header's name left the other name stale in both directions: a legacy header beside a `Current Plan:` section line, and a `**Plan:**` header beside one. * fix(#3784): derive the shape error from the schema, widen the test coverage Review round 3 on #3791: B3, m1, m2, m5 and the two test nits. B3 — the accepted-shape set had two owners: the parser branches and an English list hand-written beside them in `state.cts`. Nothing coupled them, so adding a branch left the message stale and removing one left it advertising a shape that errors, with no test able to see either. The message is now built from `STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the schema instead of restating the list. `Plan: N of M` is still spelled out explicitly because no schema row owns the body-only `Plan` field — `buildStateFrontmatter` never reads it into frontmatter, so it has no key to hang a row on. m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the branch under review. It now drives both compound spellings and ranges past 99 so the width transition is covered by the property rather than one example. A second property covers the legacy pair's own preservation contract. Both were mutation-checked: dropping the padStart turns 9 tests red. m2 — degenerate boundary fixtures around the threshold (`0 of 0` is phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored grammar must refuse, including Arabic-Indic digits. m5 — `docs/json-errors.md` described rather than quoted the message, since it is now schema-derived and a verbatim quote would be a third owner. Nits — the CRLF assertion could not see a `\n` at index 0; the `!/^Plan:/m` presence proxy is now an identity assertion on the whole `## Current Position` body. * fix(#3784): give the section plan write its own flag, and stop narrowing what parses Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2. B1 — the section fallback was guarded by `!mutated`, and `mutated` is FUNCTION-wide, already set by the phase/status/lastActivity arms that `advancePlanCore` always populates. A section spelling the field bold or as a pipe-table row therefore skipped its fallback because an UNRELATED field had been refreshed, and stayed a plan behind the header — the split-brain document this arm exists to prevent. The arm now tracks its own `planWritten`. Worth recording: the reviewer's fixture does not reproduce. The body-level status write lands on the section's own `Status:` when the document has no header `Status:`, so `mutated` is still false by the time the plan arm runs and the fallback fires. The discriminating shape needs a header `Status:` to absorb that write AND a bold section plan line. The mechanism was right; the example was not, and the regression test uses the shape that actually fails. B2 — `fallbackName` chose one name by ternary. In the legacy shape both values are populated, so it always chose `Current Plan` and a `**Plan:**` section line — which base did write — got nothing. Each name is now attempted independently with its own fallback. B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`, `Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond this transition's read/write out of scope, so that narrowing was not licensed. Both grammars now carry the same trailing tolerance. The prose defect stays closed by the START anchor, not by forbidding suffixes. Major 1 — the whole-body `Plan` write is scoped to documents that declare a `Plan` field, instead of firing unconditionally where `stateReplaceField`'s first match could be prose outside `## Current Position`. Major 2 — the error message names both `Plan` spellings the parser accepts; it previously omitted the sibling-paired form, which is the same message-disagrees-with-parser drift the derivation exists to close. B4 — the changeset claimed a guarantee B1 broke; it now describes what ships. Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the CR; the trailing group is belt-and-braces, not the primary defence). All three blocker regression tests verified red against the pre-fix source. * test(#3784): pin the hybrid shape against #3807's ambiguity refusal #4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this branch's last run, on the same function. The guard sits above the parse, so a refused document is never parsed and the shape #3784 adds cannot reach the mutation — but that is a property of source ordering, so assert it as behaviour instead. Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled, the ambiguous hybrid document advances its FIRST entry's `Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect, reached through #3784's shape. Both tests go red; both go green with the guard restored. The control pins the other direction: an unambiguous hybrid section still advances, and its zero-padding still survives. * fix(#3784): advance every spelling from its own text, refuse when they disagree Round 6 review. B1 and M1 are one defect, so they are one fix. `advancePlanCore` picked one field to parse from, computed `newPlan`, then wrote BOTH spellings from that field's numbers. Two symptoms: B1 With `Plan` as the parse source, `Current Plan` was re-stamped with the number just derived from `Plan`. `Current Plan: 7` beside `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing derived for that field, no error, no diagnostic. M1 With the legacy pair winning, the `Plan:` line was re-rendered from a bare `${newPlan} of ${totalPlans}` built out of the sibling field. `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`. The changeset's claim that padding and everything after it survive was true only for whichever field happened to be the parse source. Now: every spelling is advanced from its own raw text via `bumpLeadingNumber`, so each keeps its own padding, its own total and its own trailing annotation. Differing TOTALS are preserved, not reconciled — `Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`. Differing CURRENT numbers are refused, with `reason: "ambiguous_plan_position"` and both candidates named. Same posture as #3807's multi-`Phase:` guard one field over: name the conflict, let the caller resolve it, never pick. The guard sits immediately after the parse, BEFORE the phase-complete branch — guarding only the normal advance would let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase complete" into a document whose two spellings never agreed. A field present but unreadable (`Plan: TBD`) is left exactly as authored. Refusing the whole document because an unrelated line cannot be read would be a narrowing #3784 does not license; writing a derived number over it is the fabrication B1 was filed for. The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone. It existed only so the write path could ask which field the value came from, and the write path no longer asks. M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A revision of this PR added it; base refused it. It cannot be given the schema-row + forcing-test coupling the other shapes have, because `Plan` is body-only and `buildStateFrontmatter` never reads it into frontmatter, so there is no `current_*` key to hang a row on. Parser, the spelling in `advancePlanShapeError`, and the lockstep test move together — the invariant is the lockstep, not the length of the list. N1. The whitespace narrowing (`5phases` no longer parses where `parseInt` read 5) is documented in the changeset beside the other deliberate narrowings, rather than loosened. Loosening restores the half-parse this change exists to remove. Tests: eight new cases plus a property that crosses the two spellings with agreeing and disagreeing numbers — the review noted the existing properties never did. Fail-first proven: restoring the old write path reddens seven of the eight, both new property arms, and two pre-existing padding tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * fix(#3784): report Current Plan as updated only when it was written The write became conditional in the previous commit — a `Current Plan:` that is present but unreadable is left as authored — but the `updated` push stayed unconditional, so `transitionCore` reported a field it had not touched. `reconcileReportedFields` would have caught it against the persisted bytes at the `state.cts` caller, but `transitionCore`'s own `updated` is consumed directly (milestone-lock, the transition tests) and has to be true on its own. Covers the mirror of the unreadable-spelling case: `Current Plan: TBD` beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the legacy field is the one that cannot advance. Fail-first proven. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * test(#3784): account for the new refusal in the output({error}) census `tests/io.test.cjs`' A3 census asserts the exact population of `output({error})` call sites in `src/`, per module. The `ambiguous_plan_position` refusal added a 27th to `state.cts`, so the census went red at 26/65. Updated the way #3807 updated it when it added the ambiguous-POSITION error one line above: bump the count and name the addition inline, so the next person reads why the number is what it is. The alarm did its job — it is the only gate that noticed a new user-visible error path had been introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7c116b1c17 |
fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744)
* fix(#3697): warn when the Requirements line under-selects REQ-IDs
`cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting
on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is
correct for the canonical comma list the template ships, and silently wrong
for every other form:
`RANGE-01 … RANGE-05` -> the two ENDPOINTS only; the interior IDs are
never considered, yet `requirements_updated`
reports true with zero warnings
`RANGE-01…05` -> ZERO IDs; the whole line is inert
The silence is structural: the only cross-check, `ghostReqIds`, is itself
`citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it
by construction — and to `traceabilityWriteMisses` and `requirements_updated`
with it.
Warn on both paths. This does not add range support: the selected set is
unchanged, so no existing ledger write changes. The trigger is ID-SHAPED
EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a
range operator joining two IDs — with parenthetical citations and HTML
comments stripped before the scan, so the #2334/#2339 over-warning on
`None`, on the shipped `<!-- brackets optional -->` template comment, and on
annotated lines cannot return.
Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs
(10 cases: 4 defect, 2 canonical controls, 4 negative-space controls).
Fixes #3697
* fix(#3697): rework under-selection detection onto tokens, not a free-text scan
Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb.
That review refuted 5 of 9 claims; three were false-positive classes in exactly
the category #2334/#2339 had to REMOVE:
`RANGE-01, RANGE-02 - 3 points` the bare-hyphen alternative read
`RANGE-02 - 3` as a range
`REQ-01, REQ-02 — locked per ADR-7.` the trailing period kept `ADR-7.` out
of the anchored filter, so the
unanchored substring scan reported it
as unparsed
`REQ-01, REQ-02 (see (ADR-7), then ADR-8)`
nested parens left `ADR-8)` behind
Replaces the free-text substring scan + loose range regex with three narrow,
token-based rules (R1 range-shaped token, R2 pure range operator flanked by two
selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's
CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks
nothing — it now says "selected".
Side effect: the two false NEGATIVES the same review found are now covered —
`RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`.
NOT YET DONE (see the handoff prompt): regression tests for the four false
positives, the two new true positives, and the #3697-4 tightening the review's
CLAIM 8 asked for (it currently filters on the warning's phrasing rather than
asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests
green, and a 20-case standalone harness covering every case above.
* test(#3697): pin the v2 token-detector boundary end-to-end
Six new cases + two hardenings for the review findings against v1:
- #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) —
the operator set's `to|thru|through` arm was previously untested.
- #3697-5 (new): a tight range hidden inside balanced parentheses
(`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token,
and ticks exactly RANGE-01 — the paren shave must not hide it.
- #3697-4 gains the four false-positive classes a free-text detector
produced: numeric estimate (`- 3 points`), date annotation, em-dash
citation with trailing period (`— locked per ADR-7.`), and nested
parenthetical citations.
- #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty,
not that one phrase is absent — a re-worded over-warning cannot pass.
Negative control: against the merge-base with its lib rebuilt, all 6
defect tests fail and all 10 controls pass.
* fix(#3697): close round-2 review findings — annotation false positives
Round 2 of the adversarial review (against 822a72a04) refuted five
claims; this closes the false-positive class and the cheap misses:
- R1's bare-hyphen arm now demands a full ID on BOTH sides
(`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation
(`FY-2026-08`) and a sub-numbered ID, and warning on those is the
expensive class. Tight hyphen shorthand with a live selection is the
disclosed false negative; at zero selection R3 still catches it.
- R2 requires the endpoint pair to imply an INTERIOR (same prefix,
gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop
nothing, so an annotation hyphen between adjacent IDs stays silent.
- R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty
line citing its rationale, not unparsed residue.
- Token shave: quotes/backticks now shaved from alphanumeric tokens
(`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a
bracket-only shave so `(..)` surfaces its operator.
- 256-char token cap bounds the quadratic unanchored substring test.
- Warning text mentions range expansion only when a range rule fired.
Tests: 6 new cases (22 total). Negative control against the merge-base:
8 defect tests fail, 14 controls pass.
* fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers
Round 3 of the adversarial review (against 2eb92dd0e) refuted four
claims; this closes them:
- Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the
tokenizer before R1's `\s*` can see them and under-selected silently.
A glued-fragment rule warns when an operator is glued to a full ID
with an ID-shaped neighbour on the open side and the endpoint pair
implies an interior.
- Cross-prefix pairs around a separator no longer read as ranges:
`REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and
real ranges are same-prefix by nature. `impliesInterior` now returns
false on prefix mismatch and computes the gap with BigInt (parseInt
lost precision past 2^53).
- The token shave now removes markdown emphasis markers and curly
quotes, so `**None** (per ADR-7)` reaches the placeholder gate and
`**RANGE-02..RANGE-05**` reaches R1.
- The unanchored-substring cap rises to 2048 (a markdown-link range
with a long URL cleared 256); the anchored range regexes scan
linearly and drop their cap.
Tests: 6 new cases (28 total; 377/377 file-wide). Negative control
against the merge-base: 11 defect tests fail, 17 controls pass.
* fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule
`TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment
rule read it as `to` + `REQ-05`, warning on the canonical two-ID list
`REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only
(`..`+, ellipsis, dashes): a word operator glued to an ID is an ID,
not a range spelling.
Tests: word-operator-prefixed ID control (misparse channel silent; the
fixture's ghost-ID warning legitimately fires, so the whole-channel
assertion stays with the registered controls) and an underscore-wrapped
tight-range defect case. 30 targeted cases; 379/379 file-wide; negative
control: 12 defect tests fail on the merge-base, 18 controls pass.
* fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording
- The glued-fragment TRAILING arm takes the word operators back: an ID
must end in digits, so `REQ-01through` can never be an ID — the
round-4 TOREQ collision was leading-arm-only, and symbol-only on both
arms lost the `REQ-01through REQ-05` typo class.
- A trailing run of 2+ dots survives the punctuation shave: `REQ-01..`
is a glued range operator, not sentence punctuation, and the shave
was silently eating the `REQ-01.. REQ-05` form.
- The warning now says the line "could not be parsed as" a
comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list
— the selector just cannot parse decorated tokens — and a warning
that misstates the input teaches readers to distrust it.
Tests: two new trailing-glue defect cases (32 targeted; 381/381
file-wide). Negative control: 14 defect tests fail on the merge-base,
18 controls pass.
* test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset
CONTRIBUTING bans try/finally inside test bodies (it masks failures);
the approved shape is `t.after(() => cleanup(tmpDir))`. All seven
converted tests are this PR's own additions; the file's pre-existing
instances are untouched.
* chore(#3697): add changeset fragment for the Requirements-line under-selection warning
changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a
user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed
fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per
the house format, with the (#3697) backlink.
* refactor(#3697): extract the Requirements-line detector to a testable surface
Round-3 review Blocker 1 requires a fast-check property test over this
detector (`RULESET.TESTS.property-based-testing`: modules implementing
parsing contracts must include at least one), and Blocker 2 requires
limit-1/limit/limit+1 fixtures on its 2048-char token cap
(`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while
the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test
reaches it by spawning the CLI, and a property test cannot pay a subprocess
per generated case.
So the selector and the three detection rules move to module scope as
`analyzeRequirementsLine` (pure, exported) plus
`formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This
commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and
the pre-round suite passes against it unmodified (403/403).
Two things the move makes explicit rather than incidental. The selector and
the detector tokenize the SAME line DIFFERENTLY — the selector strips only
`[` and `]`, the detector also shaves quotes, emphasis and trailing sentence
punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected
while `ADR-7` is still nameable in a warning. They now sit adjacent with the
reason written down, so they cannot drift apart silently.
And the stale citations in the moved comment are corrected. It pointed at
src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had
drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is
`gsd-core/templates/roadmap.md:32`. Both are now anchored by content.
* fix(#3697): stop the warning claiming a misparse that did not happen
Round-3 review Major 3 and Minor 4. Both are the same defect: the warning
asserted more than the evidence supported.
MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05
deferred` warned "could not be parsed ... Range forms are not expanded;
rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05,
i.e. every ID written on the line. Nothing was dropped, and the pinned
control only stayed silent because its pair was ADJACENT (gap == 1), so the
control was passing by accident of the fixture rather than by the rule.
The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a
range and as an annotation separator are textually identical, and no
token-level rule separates them; staying quiet re-opens the exact silent
under-selection #3697 is about. Deciding the ambiguity by assertion in
either direction is wrong. So it is DISCLOSED: the warning now has two
channels, chosen by whether any ID-shaped token was actually left unselected
(`droppedIdShaped`).
* something was dropped (tight range, glued fragment, inert residue)
-> "could not be parsed as a comma-separated REQ-ID list", as before.
* nothing was dropped (only the spaced-operator rule fired)
-> "contains what reads as a range between two cited REQ-IDs", stating
both readings and saying explicitly that an annotation separator means
the line is already correct.
This retires the "could not be parsed" wording for the four #3697-1 spaced
cases too, and that is a deliberate expectation change rather than a fix
counted twice: those lines never failed to parse either. They still warn,
still name the selected IDs, and still assert the endpoint-only marking is
unchanged; #3697-1 now also asserts the misparse channel stays SILENT.
MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a
citation as requirement content it had failed to read. The trigger is
correct and stays: #3697's acceptance criterion asks for a warning "when it
selects zero IDs from a line that is non-empty and is not the `TBD`
placeholder", and inferring placeholder-ness from arbitrary prose is the
free-text heuristic this detector exists to avoid. What was wrong is the
wording, so the non-range arm now says "ID-shaped text that was not
selected" and names the escape the author actually has (`TBD` / `None`).
Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse,
must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A
(tracked in ADR-12)` — must warn, must not say "Unparsed text", must not
diagnose a range, must name the placeholder escape).
Reversion control: reverting the ambiguous channel fails #3697-9 (3 named
tests); reverting the R3 wording fails #3697-10 (2 named tests).
* fix(#3697): cap every token predicate, complete the dash set, cover the boundary
Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three
are about the detector's own predicates, so they land together.
BLOCKER 2 — the 2048-char budget had no boundary coverage.
`RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit /
limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 /
2048 / 2049 against BOTH predicate families the cap guards, and each asserts
its fixture's exact length before asserting behaviour, so a mis-built
fixture fails loudly rather than passing at the wrong size. Clause (d) of
that rule — an input pushed within reserve-distance of the limit — has no
referent here: this is a hard cap with no reserve constant beside it, and
the test comment says so rather than leaving the omission to be re-derived.
NIT 6 — the cap guarded only the unanchored ID-substring regex. The
anchored range regexes were left uncapped, justified by a comment asserting
they scan linearly. The finding is right that this is informational (they
are anchored; the input is a local ROADMAP.md), but an asserted property is
cheaper to enforce than to defend, so all three predicates now share one
`short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must
now decline to classify.
SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this
code fixes at author time over a domain that grows without it, so the round
owes a census of what the enumeration reaches.
reached: `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through
NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure
dash, U+2015 horizontal bar, U+2212 minus sign
consequence: a range spelled with any of those is SILENTLY under-selected
— #3697's own defect, in the code that exists to fix it
Those five close. They are the same operator at a different codepoint and
carry none of the ASCII hyphen's collision risk, because they are not the
REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a
U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm
beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is
held to is untouched, and #3697-12 pins that.
Still NOT reached, declined with reason rather than left unstated: `→`, `~`,
`..=`, `..<`, `until`, and `up to` (two tokens, so never one operator
token). Each is a symbol or word with an independent non-range use between
two REQ-IDs — the over-warning class #2334 cost three rounds.
Reversion control: reverting the uniform cap fails #3697-B2 (limit+1);
reverting the dash set fails #3697-11 (5 named tests).
* test(#3697): add the fast-check property coverage the parser rule requires
Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md)
requires modules implementing parsing contracts to carry at least one
fast-check property test asserting a domain invariant, and the round-2 diff
had zero occurrences of `fc.` across its +311 test lines. Five properties,
1,900 generated cases:
P1 soundness of silence (boundary containment) — for ANY canonical comma
list of well-formed REQ-IDs, the selected set EQUALS the written set
and nothing warns. This is the #2334 over-warning invariant and the
#3697 under-warning invariant asserted as one statement, over
generated IDs rather than hand-picked ones. It generalises #3697-4b:
a prefix beginning with a word operator (`TORANGE-05`) is an ID, and
P1 covers that class rather than the single example.
P2 completeness — a same-prefix pair with an interior between them,
separated by any of the nine spaced operators, ALWAYS warns.
P3 the #2334 invariant — an ADJACENT pair around a separator can drop
nothing, so it stays silent however it is annotated.
P4 totality + idempotency — total over arbitrary strings, deterministic,
and the formatter agrees with the analysis on whether there is
anything to say (a warn with no text, or text with no warn, is a
channel that can go silent or noisy on its own).
P5 containment — every selected ID is ID-shaped and appears verbatim in
the input.
Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and
P5 hold against the round-2 code as well as this one — they are regression
guards, not bug-finders, and their value is that the invariants are now
stated and generatively checked rather than implied by examples. P2 is the
one that would have failed before the dash enumeration was completed.
fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from
`fc.array(...).map(join)` with the alphabet pinned to the selector's own
`[A-Z0-9]` class.
These live in tests/phase.test.cjs rather than a new
`phase.property.test.cjs`: `lint-test-file-count` caps a production module
at 2 test files and phase.cts is already at its allowlisted entry, so a new
file would trade one gate for another.
* docs(#3697): document the ROADMAP Requirements-line grammar
Round-3 review Minor 5 — the change adds net-new user-visible warning output
for a grammar constraint documented nowhere under docs/. `type: Fixed` is
docs-exempt so this does not block, but a warning about a rule the reader
cannot look up is not actionable, and that is worth fixing whether or not a
gate demands it.
Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the
existing SUMMARY artifact-check advisory it is a sibling of: the supported
comma-list form, why ranges are deliberately not expanded, that `TBD` and
`None` are the entire placeholder vocabulary, and what each of the two
warning voices means — including that the range/annotation one may be
reporting a line that is already correct.
Existing file rather than a new one, deliberately: docs/ carries generated
indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity
or index gate this change has no reason to touch.
* fix(#3697): rule-scope the warning-channel discriminator
Self-found at the round's pre-push review, against the Major 3 fix two
commits back. That fix chose the channel from a LINE-GLOBAL question — "was
any ID-shaped token left unselected?" — while the rules that produce the
warning are not line-global. The two disagree as soon as the line carries an
ID-shaped token no rule fired on:
`RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)`
`(ADR-7)` survives the selector's bracket strip, so the global test called it
a drop and sent the line to the assertive channel — putting the false "could
not be parsed ... rewrite the line" claim back on a correct line. That is
review finding Major 3 returning through a side door, and it directly
contradicts #3697-4, which pins a parenthetical citation as NOT unparsed
residue.
The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the
ambiguous channel requires that R2 fired, that no other rule did, and that
every endpoint R2 fired on was actually selected. The last conjunct is not
redundant — the detector shaves brackets and the selector does not, so R2 can
fire on a `(RANGE-02)` that was never selected, and that IS a drop:
`RANGE-01 (RANGE-02) — RANGE-05` -> assertive, correctly
`droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the
channel can ask about the endpoints the rule fired on) and the
`rangeReadingOnly` verdict.
Reversion control: against the line-global rule, #3697-9b fails. #3697-9c
passes under both rules — there the dropped token IS the R2 endpoint, so the
two agree; it is a regression guard, not a bug-finder, and is recorded as
such rather than counted as a second control.
* fix(#3697): hold every dash to the strict range shape, not just ASCII
Self-found at the round's pre-push review, and it CORRECTS a claim made two
commits back. That commit widened the range-operator set by five Unicode
dashes and asserted they "carry none of the ASCII hyphen's collision risk,
because they are not the REQ-ID separator". That reasoning was wrong. The
collision is a property of the SHAPE — `PREFIX-\d+ <dash> \d+` is also a date
(`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not
care which dash sits in the operator slot, because the ID's own separator is
still ASCII either side of it. Measured:
RANGE-01 (target FY-2026-08) silent <- pinned by #3697-4
RANGE-01 (target FY-2026‐08) WARNED <- same line, U+2010
So the widening reintroduced the #2334 over-warning class on a date
annotation. It also exposed that the inconsistency PREDATES this PR: U+2013
and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash
forms of that same date annotation warned before round 3 ever ran.
One rule for every dash: a tight range spelled with any of the eight must
carry a FULL ID on both sides, exactly as the bare hyphen already had to.
`..`, `…` and the word operators stay loose — no date or sub-number reading
exists between two numbers, so the strict shape would cost them coverage for
nothing.
The cost is a false negative, and it is one the design already accepts:
`RANGE-01, RANGE-02-05` is silent today, deliberately, and now
`RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than
opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so
R3 catches it.
Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight
dashes), #3697-13b (full-ID tight range still warns for all eight),
#3697-13c (loose operators keep their numeric endpoint), #3697-13d (the
accepted false negative is symmetric, and the bare zero-selection line still
warns).
* fix(#3697): close the round's own pre-push review findings
An adversarial cross-AI review of this round refuted 4 of its 10 claims. All
four were real. Every fix below is to code THIS round introduced.
1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1).
`REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told
the author "the line is already correct and nothing needs to change" — while
`(REQ-02)` had been dropped by the selector, which does not strip
parentheses.
The channel choice is still right, and deliberately so: `(ADR-7)` and
`(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts
the false "could not be parsed" claim back on a line carrying a citation —
the misroute fixed two commits ago. No rule can adjudicate this; the author
can. So the voice stops asserting the line is correct (it now speaks about
the SEPARATOR, which is all it has evidence about), and BOTH voices gained a
factual clause naming ID-shaped text the selector skipped, with the reason
(brackets are not stripped) and no verdict attached.
2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2).
A 2049-char range token warned before this round and went silent after it:
the "uniform cap" commit bounded the predicate and, with it, the warning.
That is #3697's own defect, introduced by the fix for a nit.
The cap bounds the WORK, not the warning. An over-cap token carrying `-` is
now recorded as unclassified (a linear `includes`, never the unanchored
regex the cap exists to keep off it) and gets its own voice: "could not be
checked ... the REQ-ID selection on this line is unverified". Unclassified
is reported, never treated as clean.
3. THE CAP WAS NOT UNIFORM (review MISSED finding).
R2 capped the operator token but not its neighbours, so
`<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over
both endpoints unbounded. The glued rule had the same hole. Every
participant is capped now.
4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5).
It drew from a bare `fc.string()`, which over 500 samples produced max
length 10 and ZERO inputs containing a REQ-ID — the loop body never executed
an assertion. A containment property that never contains anything is a green
test measuring nothing. The generator now interleaves real IDs with noise
and the property ASSERTS it saw them (>50/500), so it can never silently go
vacuous again. The free-form coverage it was actually providing survives,
honestly labelled, as #3697-P6.
The same finding refuted this round's claim that P2 distinguishes pre-round
behaviour: every operator P2 uses was already in the pre-round operator set.
P2 is a regression guard, and its comment now says so.
Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review
MISSED finding) and is corrected with the code.
Tests: #3697-9d (soft voice names the skipped ID, never claims the line is
correct), #3697-9e (over-cap token reported as unclassified, still warns),
#3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the
boundary on the PREDICATE's verdict with `warn` asserted true at every length —
asserting `warn === false` at limit+1 was itself finding 2.
* docs(#3697): describe the third voice and the dash rule
Follow-on to the review-findings commit: that commit corrected the docs' claim
about the soft voice but left two things the code now does undescribed.
- There are THREE voices, not two. The over-cap voice ("could not be checked
... unverified") arrived with the fix for the review's CLAIM 2 and had no
entry.
- Dash spellings require a full ID on both sides, and `..` / `…` / the word
operators do not. That asymmetry is deliberate and load-bearing —
`PREFIX-<digits><dash><digits>` is date- and sub-number-shaped — so a reader
hitting `REQ-01-05` and getting silence has no way to find out why. The
accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still
reported) is stated rather than left to be discovered.
Documentation only; no behaviour change.
* fix(#3697): close the continuation review's findings
A continuation of the same adversarial reviewer, run against the reworked
round, refuted 6 of 7 claims. Four were real defects in this round's own work
and are fixed here; the other two are answered rather than changed, below.
1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A).
The previous commit's message said both voices gained it. Only the soft
return appended it. The assertive voice now carries it too — and, because
that voice already names range tokens and inert residue under its own
clauses, the note is filtered to what those did not already name. A warning
that says the same token twice is one readers learn to skim.
2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A).
It read "brackets and parentheses are not stripped". Square brackets ARE
stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so
only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md,
which had inherited the same error.
3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B).
`oversizedTokens` filtered on `includes('-')`, which misses an over-cap
OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined
to classify it once capped, and nothing reported it. That is the exact
regression the field was added to close, one input over. Any token past the
cap now counts — what it contains is irrelevant when we could not read it.
4. AND THEN OVER-REPORTED ONE (review MISSED finding).
With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the
uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line
is unverified" — a contradiction inside one warning. A token the selector
took was examined end to end, so it is excluded.
Two findings are answered, not changed:
CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and
deliberate: this round does not touch what gets MARKED, and the pattern is
anchored at both ends with no nested quantifier, so it is linear. The claim
that "all predicate paths are capped" was too broad; the DETECTOR's are.
CLAIM E — `REQ-01, REQ-02<dash>05` is silent for every dash. That is the
documented, deliberate cost of holding dashes to the strict shape, and it is
symmetric with ASCII, which behaved that way before this PR. The reviewer is
right that "without losing a range spelling that should be detected" was too
strong; a bare `REQ-02<dash>05` still warns.
Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim
true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a
selected over-cap ID is never called unverified). #3697-9f is rebuilt — its
first version used the SAME id twice, so R2 could not have fired even uncapped
and it proved nothing; it now uses endpoints with a gap and fails when the
neighbour cap is removed.
Reversion control: all four fixes fail a named test when reverted in isolation
(#3697-9h, #3697-9i, #3697-9g, #3697-9f).
* fix(#3697): scope the over-cap exemption to what could actually pair
A second continuation of the same reviewer, against the reworked round,
confirmed the two claims that matter most and refuted three. This closes the
one real defect; the other two are answered below.
CLAIM J / CLAIM K (one defect, found from both directions). The previous
commit exempted EVERY selector-accepted token from `oversizedTokens`, on the
reasoning that the selector is uncapped and anchored so it examined the whole
token. True of that token's SELECTION — and not the same as "no rule was
suppressed by it". Two over-cap valid IDs either side of `..` are both
selected, so both were exempted, and R2 is capped: a line that warned before
this round went silent.
That is the third appearance of one class in this round — the cap suppresses a
check, and the suppression is not reported. Each fix for it over-corrected in
the opposite direction, which is why the rule is now stated in terms of what
was actually suppressed rather than in terms of the token: an over-cap token is
exempt only when it was selected AND nothing beside it could have paired with
it into a range (no range operator, no glued fragment, no second over-cap
token). Everything else is unexaminable and says so.
Two findings are answered, not changed:
CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule
that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01`
silent: recognise `PREFIX-a<dash>b` only when another SELECTED id on the line
shares that prefix. That refutes this round's claim that the strict-dash
trade was FORCED, and the claim is withdrawn — it is a design choice. The
choice stands for this PR: the conservative rule is what ASCII already did
before #3697, adopting a new contextual heuristic unreviewed at the end of a
round is how the last three defects in this round were made, and #3697 asks
for a warning rather than better range inference. Named here so the
alternative is on the record rather than lost.
Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not
classified at all" and warns. Selection is not bounded; only range detection
is. Corrected.
Confirmed by the same pass, and worth recording because they are the PR's
load-bearing promises: a 20,000-input comparison of the pre-extraction selector
against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has
changed), and the uncapped selector regex was measured linear from 100k to 800k
characters.
Tests: #3697-9f now asserts the range case is reported rather than silent, and
#3697-9j pins the exemption's scope in both directions. Reversion control:
restoring the blanket exemption fails both.
* fix(#3697): warn on zero selection, as the acceptance criterion asks
`Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed
SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`,
the `placeholderLed` census comment, and the advice string the command emits
to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned,
because the citation supplied the ID-shaped residue R3 required, while bare
`Deferred` did not. The claim was written into three places and never
executed once.
This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length ===
0` while the raw capture is non-empty and not `TBD`".
R3b keys on the SELECTION being empty, never on what the prose means, so it
adds no free-text heuristic. It is deliberately not gated on ID-shaped
residue the way R3 is, and the negative space is what settles that: all
fifteen #2334/#2339 fixtures are held silent by non-zero selection or by
`placeholderLed`, and not one of them by the ID-shape gate — measured, not
argued. The gate was buying no negative space while costing the acceptance
criterion.
`tokens.length > 0` keeps an empty line and a comment-only line silent: the
tokenizer strips `<!-- ... -->` before splitting, so the shipped template's
own comment cannot reach the rule.
Selection behavior is unchanged. This warns; it never invents an ID.
Also extracts `warn` to a named const (round 3 review Minor 3) — this commit
adds a disjunct to exactly that predicate, and in the return literal a later
reordering would be a TDZ ReferenceError rather than a reader-visible error.
Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the
warning names the TBD/None escape), #3697-14b (five placeholder spellings
stay whole-channel silent), #3697-14c (comment-only line stays silent).
Fail-first controls: all six #3697-14 cases fail against the pre-fix tree;
-14b and -14c pass at both ends, which is correct — they pin silence the
widening must preserve.
* fix(#3697): name the REQ-ID a glued delimiter dropped
`RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with
`requirements_updated: true` — #3697's own half-success failure mode, reached
by one wrong delimiter, and silent before this rule. It is the issue's AC-1a
("a warning whenever the line contains ID-shaped content that the tokenizer
did NOT select") at the shape most likely to be typed by accident.
Round 4 review rated this Major rather than Blocker on the ground that the
case is indistinguishable from a parenthesised citation, since `(ADR-7)` also
shaves down to a bare ID. At the RAW token level it is distinguishable, and
that is what makes the rule shippable: `REQ-01;` is shaved of a trailing
DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and
requires the token to sit outside any parenthetical.
Measured before implementing: 0 false positives and 0 false negatives across
21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first
cut without the parenthetical test scored 3 false positives — every one of
them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that
test is the rule's boundary rather than an optimisation.
Adds the delimiter census the module did not have. The range-operator domain
was already censused; the comma-substitute domain was not. Swept 26
spellings: exactly two produce a silent under-selection, `; ` and `: `. Every
other spelling either selects both IDs or selects none and already warns. The
review hand-listed the semicolon; the colon is the sibling that sweep found,
and it fails identically.
`rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing
was dropped, and must not speak for a line where something demonstrably was.
Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick
exactly the unchanged selection), #3697-15b (three citation forms stay
whole-channel silent). Fail-first control: all four #3697-15 cases fail
against the previous commit's tree; -15b passes at both ends, pinning the
boundary the widening must not cross.
* fix(#3697): give the Requirements-line warning a stable machine kind
The warning's kind existed only in the prose of its message, so every consumer
and every test had to regex an English sentence — and rewording a message
silently un-asserted the tests that pinned it. Round 4 review Major 3.
The repo already had the settled seam for exactly these semantics.
`CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a
truncated scan, which is precisely this module's third voice; and
`WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same
reason. ADR-3473 Decision 3 ("failure is a value") points the same way.
`formatRequirementsLineWarning` now returns `{ code, message }` instead of a
bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a
legitimate value, but the success arm is no longer a naked string one field
away from the shape the ADR standardises on.
The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is
a documented `string[]` in `phase complete`'s JSON output, rendered by
execute-phase.md's "If has_warnings is true" step, so re-typing its elements
would be a breaking output-contract change for a shipped command. The code is
emitted as its own additive `requirements_line_warning` field, absent
entirely when the line is clean.
Vocabulary, exported so tests key on it rather than on string literals:
`req-line-misparse`, `req-line-range-reading`, `req-line-unverified`.
Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code;
message-content assertions stay where the user-visible wording is itself under
test. #3697-16 pins the code end-to-end through the CLI's JSON for four line
shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean
line emits no kind at all, because a field present on every run carries no
information. #3697-P4 holds kind-and-message-appear-together and
kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added
later cannot ship without one.
* test(#3697): pin the divergence against the second parser of the same line
CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when
sharing constants/arrays/parsers between parallel surfaces, add a parity
assertion test that fails if they diverge." Round 4 review Major 2.
`normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP
`**Requirements:**` value — its own docblock says callers "may pass the
roadmap value through verbatim" — and diverges on four axes. Measured, not
inferred:
line phase complete gap-checker
RANGE-01..RANGE-05 [] 5 IDs
None (per ADR-7) [] ["ADR-7"]
(REQ-02) [] ["REQ-02"]
REQ-01a [] ["REQ-01a"]
REQ-01, REQ-02 both both
This pins the divergence rather than removing it, which is the review's
second option and the correct one here: unifying the two would change what
`phase complete` MARKS, and "the ledger-writing set is byte-identical to base"
is the one invariant this PR holds fixed. Every axis is now asserted in BOTH
directions, so drift on either side fails here instead of widening silently.
The range axis is a DELIBERATE disagreement and is labelled as such — #3697
declines range expansion in terms ("I am not asking for range syntax to be
supported") while gap analysis adopted it under #1269.
The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a
declared-empty line to `phase complete`, which reads the lead token, and a
one-requirement line to gap-checker, which strips parentheses first so the
citation survives its ID-shape filter. That is a citation being reported as a
requirement.
#3697-17b states the cost concretely: one line, five requirements in scope to
gap analysis and zero to phase complete. This PR is what makes that
contradiction visible, by finally giving the silent side a voice.
* fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap
Two round 4 review minors, both about a warning saying something it cannot
support.
MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so
`FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`.
`REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape
silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no
warning at all — but whenever some OTHER rule fired on a line that also
carried a date annotation, the rider told the author to "check whether any of
it is a requirement" about a date. Not a false warning, since the line was
warning anyway; false CONTENT, in the #2334 voice, through the side door.
Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped`
stays a faithful record of what the selector skipped — it is documented as a
fact that never routes — while the user-facing clause declines to assert
requirement-ness about a shape the design already ruled unadjudicable.
#3697-18b is the other half, so the filter cannot become a silencer: a
genuinely dropped REQ-ID is still named.
MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent,
so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel
would have passed. Whole-channel silence is not available on that fixture (the
pre-existing ghost-ID warning legitimately fires on the unregistered
`TORANGE-05`), so the precise assertion is that no Requirements-line warning
of ANY kind was emitted. The machine code added earlier in this round is what
makes that statable; before it, "both channels" could only have meant a second
prose regex.
* docs(#3697): record the Requirements-line seam in CONTEXT.md
CLAUDE.md names the CONTEXT.md glossary as a PR gate, and
`get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25
co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a
named seam, three warning kinds, a bound, a rule taxonomy and a deliberate
cross-parser divergence, and recorded none of it. Round 4 review Major 4.
The precedent is explicit rather than inferred: the directly analogous seam
is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound
and its boundary obligation — and that entry is the one this module's third
voice was modelled on.
Eight predicates, in the machine-oriented section beside it:
.module the two exported functions and the code vocabulary
.selector-identity citedReqIds is byte-identical to base and is the
only thing reaching the ledger — a change to what
phase.complete MARKS is outside this contract
.rules R1 / R2 / R2' / R3 / R3b / R4 / over-cap
.kinds the three codes, and why they ride beside
warnings[] rather than inside it
.cap 2048, neighbours included, and the boundary rule
.placeholder the gate that actually holds the negative space
.census-domains both open domains with their NOT-reached members
.gap-checker-divergence the four axes, pinned not unified
The changeset type is `Fixed`, which exempts this PR from the docs/
co-change requirement — but the glossary gate is separate from that
exemption, and the 2048 cap in particular is a machine-canon-shaped fact
that until now existed only inside a source comment.
`docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync
confirms all six targets in sync.
* docs(#3697): document what the command now does, in one changeset sentence
DOCS. The grammar section predated this round's two new rules, so it
under-described the behaviour it exists to make lookup-able:
- The placeholder paragraph enumerated three words; the rule is a DEFAULT.
Any wording that selects no REQ-IDs warns, and the placeholders are matched
as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty
too. The comment-only line is called out, because "any other wording" would
otherwise read as covering the shipped template's own `<!-- ... -->`.
- The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is
the quietest way to lose a requirement on this line — `requirements_updated`
reads `true` either way — so it gets its own paragraph, with the
parenthetical exemption stated beside it.
- The machine kind is documented where a consumer would look for it, with the
instruction to key on the kind rather than the wording.
- The skipped-text note no longer implies it reports date shapes; it
deliberately does not, and silently omitting that left the doc promising the
behaviour this round removed.
CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is
`**<Bold user-visible change>** — <symptom-led explanation>.` and both
canonical examples are one sentence; this fragment ran three. Now one, and
covering what the round actually delivers rather than only the range shape it
started from.
* test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint
A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The
docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/`
references in exempt test files and fails when a new one appears, so that
comment turned four green gates red — `ci-docs-guard-registry` and the
registration lint — for a file that reads no documentation at all.
Caught by diffing the full suite's failing-name set against the same suite run
at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base
(install / config-home / shadowing tests under the sandbox HOME), and exactly
these 4 did not.
Rephrased rather than baselined. Adding the path to
DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely
starts READING a new docs path — the violation text asks the author to
re-confirm the exemption still holds. Nothing here reads documentation; the
guard matched prose. Baselining would have recorded a coupling that does not
exist and made the next reader wonder what phase.test.cjs does with
CLI-TOOLS.md. The comment still names where the contract is written, just
without planting a path string.
* fix(#3697): close four defects this round's own pre-push review drove
An adversarial cross-AI review of this round, run before the push, returned 6
CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree,
and every one was a shape the author had not probed — the rules were correct
across the probe set and wrong just outside it.
(1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through
the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired:
`ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does
not reach a BARE citation. The FP probe that scored this rule 0/0 only ever
tested the parenthesised form.
Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the
module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already
demands an agreeing prefix for the same reason. Cost, stated in the census: a
dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays
silent. Same trade the strict-dash rule takes — under-report a rare shape
rather than over-report a common one. Pinned as a declared blind spot by
(2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped
REQ-01 silently: the selector strips square brackets and R4's raw scanner did
not. The bracket spelling the shipped template recommends was the one shape the
rule could not see.
(3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal
requirement id — gap-checker's `parseRequirements` accepts it from
REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement
behind a rule meant only to hide dates. Narrowed to a four-digit year segment.
The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED,
and the removal is recorded in place: its premise was refuted, it was not
inconvenient.
(4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while
reading as empty to the author, so R3b fired with nothing on screen to explain
it. Zero-width and format characters are now stripped — stripped rather than
treated as delimiters, because splitting on one would fabricate two fragments
out of one ID.
Also corrects the documentation the same review found overstated: the line is
split on commas AND whitespace, and the ID shape is matched case-insensitively,
so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was
pre-existing selector behaviour; this round is the one that asserted the docs
were true of it.
Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is
stripped, not split on), -19c (three citation forms), -19d (both bracket
spellings), -19e (both halves of the rider boundary), -19f (the declared blind
spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0
failures; lint:ci clean.
* fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop
The pre-push review's continuation refuted six of seven follow-up claims. The
first one is the one that mattered: the invisible-character fix committed in
c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from
the detector wholesale made `REQ-01<ZWSP>, REQ-02` go SILENT — the selector
really does drop REQ-01, and the strip removed the only evidence of it. The
test written alongside asserted the tokens and the empty R4 result and never
asserted `warn`, so it DOCUMENTED the bug rather than catching it; that
omission was the reviewer's own MISSED finding.
An invisible is two different questions about one character, and the fix is to
stop conflating them: absence-of-content for the empty test, DECORATION on a
token for the drop rule. Neither is a reason to delete it from the line.
R4 is generalised accordingly, because the continuation drove four more shapes
a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`,
`REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus
`**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one
class: decoration on a token the selector then cannot take. One rule, not four
patches; patching them individually is how a list stays short and wrong.
PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving
them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was
never there and broke #3697-9d's channel routing with it. A parenthesis is this
rule's citation marker.
The rider stops adjudicating an undecidable shape. `API-2-01` is a legal
requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15`
are dates — no regex separates them, and both filters this round tried scored a
miss in each direction. It now NAMES the token and states the ambiguity, which
is the same thing the two warning voices already do about a range separator.
Filtering hides a real dropped requirement; reporting it bare asks the author
whether a date is a requirement; saying "this may equally be a date" does
neither.
The census and the docs are corrected to what the code does, including the part
that is NOT complete: the prefix gate does not stop a citation that SHARES a
selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level
separates that from a real drop. A prose heuristic on "see" is the free-text
detector this module exists to avoid, so the honest move is to say so.
CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a
20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is
identical to upstream/next's for every input, and marking is untouched.
513 tests in phase.test.cjs, 0 failures; lint:ci clean.
* fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line
Third pass of the round's own pre-push review, scoped to regression-hunting
rather than further polish. Three findings, all driven, all mine.
R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed
a dropped requirement: the previous cut treated any shaved decoration as
evidence, and emphasis is not evidence. Nothing separates that line from
`**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive
signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the
token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to
the skipped-text rider, which names the id without asserting a drop, exactly as
`(REQ-02)` is handled. That is the #2334 class caught one cut before shipping.
R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper
sitting behind a delimiter. Shaves to a stable point now.
The range OPERATOR lost its invisibles handling. `REQ-01 <ZWSP>..<ZWSP> REQ-05`
went silent, because the previous commit removed the invisible strip from BOTH
the tokenizer and R4 when only R4's was wrong. An invisible is two questions
about one character: for the classification rules it is noise and is stripped
from the token; for the drop rule it is the evidence and must survive on the
raw line. Stripping in both places hid a dropped id; stripping in neither hid a
range. The reviewer's MISSED finding named the missing control — regression
tests covered invisibles inside ids and not beside operators — and #3697-19i is
that control.
UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03`
reported nothing: a running-depth counter left the unclosed `(` open through
end-of-line, so every genuine drop after it inherited citation immunity. A
parenthesis confers that immunity only as part of a MATCHED span now — an
unmatched one is a typo, not a citation.
CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode:
`citedReqIds` identical to upstream/next, marking untouched, warnings appended.
522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero
head-only failures against a probe worktree at upstream/next.
* fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright
Fourth and final pass of the round's own pre-push review. Three findings, and
they shared one root cause, so this is a narrower rule rather than a longer list
of shapes.
TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single
whitespace token, so a matched parenthetical inside it conferred immunity on the
`REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans
are now deleted from the line outright — which states what is actually meant,
that for this rule a citation is not on the line — and an UNMATCHED paren is a
typo that confers nothing. That also retires the running-depth counter whose
previous bug was the mirror image: an unclosed `(` swallowing the rest of the
line.
DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was
reported as a dropped requirement. It is a citation with sentence punctuation.
The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be
touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify;
`**REQ-01**;` does not, because outside the styling that character is
punctuation. An invisible needs no adjacency test — nobody types one on
purpose, so anywhere in the token it is corruption rather than intent.
`**REQ-01**; REQ-02` therefore goes silent, and the test row asserting
otherwise is inverted rather than deleted quietly: it was added one commit ago
on the reasoning this pass refuted, and nothing distinguishes it from
`see **REQ-7**; next topic`.
Worth recording plainly: three successive cuts of this rule fired on a
citation, and each fix was a narrower definition of EVIDENCE, never a longer
list of shapes. The list-lengthening instinct is what produced the bug each
time.
The review's last MISSED finding named the missing control — the paren tests
all surrounded matched spans with whitespace, so none covered a span sharing a
token with an id outside it. #3697-19j carries both directions now.
526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero
head-only failures against a probe at upstream/next; the invariant that
`citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary
Unicode inputs.
* fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop
Two defects, both found by this round's own pre-publication body claim-audit.
1. A DEMONSTRATED drop was discarded by the unverified voice. On
`REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in
delimiterDroppedIds and the formatter then reported `req-line-unverified`,
whose message never mentions it — the one actionable finding masked by the
token beside it. The over-cap channel now excludes a line carrying an R4
hit, exactly as rangeReadingOnly already did and for the same reason: that
voice's whole claim is that nothing could be checked, and R4 has already
checked something. The assertive channel still carries the over-cap rider,
so nothing about the cap is traded away. Pinned by #3697-19l, which fails
against the pre-fix build and nothing else does.
2. Three shipped artifacts asserted behaviour the code does not have — the
same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules
predicate, the CLI tools reference, and the warning's own advice string all
listed markdown emphasis as an R4 trigger. It is not: styling is shaved
BEFORE the test and tolerated around an id, never a trigger on its own, so
`**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger
is exactly a glued `;`/`:` or an embedded invisible.
The census predicate was wrong in a second way. Its 26-spelling separator
sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and
therefore biased: one-sided attachment drops silently for every punctuation
outside the set — `/ | & + . > \` and the full-width and non-ASCII forms
`; , ؛` all measured silent. The domain is wide open and R4 covers two
characters of it. Said plainly in all three places rather than widened
here: every previous widening of this rule first fired on a citation, so it
is not done blind at the end of a round.
Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n
one-sided separators) so the documents and the code cannot drift apart again —
which is what the round-4 blocker asked for.
tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* fix(#3697): re-sweep the separator census properly, and say what it really found
The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `"
from a 26-spelling sweep. That conclusion was forced by how the sweep was
built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the
semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for
every other separator. Different members of the domain were tested in
different shapes, so no other answer was reachable. Caught by this round's
pre-publication claim-audit of the response comment, reading the census
comment against its own swept list.
Re-swept fully crossed and driven through the built artifact: 21 separators x
{bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select
both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All
34 are one shape — a separator glued to exactly one of the two ids, e.g.
`REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and
the `;`/`:` that R4 covers.
So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census
comment, the CONTEXT.md census-domains predicate and the CLI tools reference
all say. The set is deliberately not widened here: three successive cuts of
this rule fired on a citation, and a fourth at the end of a round with no
adversarial pass is how each of those got in.
Second false passage in the same block: styling-only decoration was described
as "left to the skipped-text rider, which names the id". A rider only exists
inside a message, and a message only exists once some rule sets `warn` — so on
a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent.
Describing it as handled reads as coverage. #3697-19m already pins the silence.
so the test matches the documented claim.
tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next
Rebased onto next @
|
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
1a358ce0fd |
feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
62b0d939b6 |
feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey Add an optional `timeoutConfigKey` field to the reviewer lane descriptor, resolved in `resolveLanePlan` at invocation time and falling back to the frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes declare `review.timeouts.<slug>` on both surfaces (the descriptor and their capability.json manifest), validated by capability-validator.cjs. For the antigravity lane, the native `agy --print-timeout` flag — previously a second hardcoded literal (`540s`) independent of the outer cap — is now derived from the same resolved outer timeout in `antigravityArgv`, preserving the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is declared data, the inner one is handler-owned). The antigravity default timeoutFloorMs stays at 600s per the maintainer's disposition; users raise it through the new config key instead. * docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper Address code-review findings on the timeoutConfigKey change: extract the inline timeout-resolution logic into a named, exported, directly-tested resolveTimeoutMs helper (matching the file's existing configString/ normalizeHost convention); document the new review.timeouts.* federated config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md, and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment. * fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner gsd-test caught two design mistakes in the prior commits: 1. SpawnPlan.argv is documented and tested as fully resolved by resolveLanePlan (model/effort/output/prompt already folded in) — leaving the antigravity '{{nativeTimeout}}' marker unresolved until the runner's antigravityArgv violated that contract and broke tests that read plan.argv directly (tests/antigravity-reviewer.test.cjs, tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}' is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself via the new nativeTimeoutToken() helper, exactly like the other four. antigravityArgv reverts to its pre-#3274 four-argument form. Also missed updating capabilities/antigravity/capability.json's invoke.args to match the descriptor, which broke the manifest/descriptor parity test. 2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three lanes with neither a model flag nor a host — may own no config key beyond their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes violated it. Fix: those three keep timeoutConfigKey: null and own no review.timeouts.* key, matching their existing modelConfigKey: null. The other 9 lanes are unaffected. * chore(#3274): backfill changeset PR number (pr:0 -> 4083) --------- Co-authored-by: sim <sim@local> |
||
|
|
6eea00b707 |
enhance(#3301): tell reviewers the plan ids and total count, grade coverage (#4084)
* test(#3301): add failing-first plan coverage manifest tests Failing-first regression tests for the plan-id manifest, the updated Review Instructions, and the mechanical per-reviewer coverage check, ahead of the review.md implementation. RED baseline before the fix lands. * test(#3301): raise allow-test-rule-refs unverified ceiling for new marker Adding tests/review-plan-coverage-manifest.test.cjs's source-text-is-the-product marker grows the unverified-exemption pool by one (282 -> 283), the same documented growth path scripts/lint-allow-test-rule-refs.cjs's own failure output names. Confirmed clean via 'npm run lint:allow-test-rule-refs' locally. * feat(#3301): tell reviewers the plan ids and total count, grade coverage build_prompt now derives a plan-id manifest from each *-PLAN.md filename (stripping the -PLAN.md suffix) and appends it, with the total plan count, to both gsd-review-instructions.md and gsd-review-prompt.md. The Review Instructions prose requires one heading-verbatim section per id before any cross-plan or overall-risk content. write_reviews grades each dispatched lane's real (non-stub, non-empty) review against that same manifest and records an optional plan_coverage: frontmatter block, present only when a lane is incomplete. The match escapes regex metacharacters in the id and excludes a preceding/trailing hyphen or word character as a boundary, closing the two traps named in the issue (a decimal phase like 12.6 satisfied by 12X6-01; a threat id like T-04-07 registering as coverage of plan 04-07). CodeRabbit is exempt, since it never receives the source-grounding prompt carrying the manifest. This closes the gap where a review that silently covers only some plans in a multi-plan phase is indistinguishable from one that covers all of them. * docs(#3301): add changeset fragment * test(#3301): use t.after() instead of try/finally for cleanup CONTRIBUTING.md bans try/finally inside test bodies. Code review caught this in the new coverage-manifest test file; switch every fixture-cleanup site to the approved t.after() pattern. * test: use t.after() instead of try/finally in #3300's build_prompt tests Pre-existing try/finally-for-cleanup pattern in this file (landed for #3300) violates CONTRIBUTING.md's explicit ban on try/finally inside test bodies. Surfaced incidentally while reviewing #3301's diff, which cites this file as its extraction-pattern precedent; fixed inline per the no-defer rule rather than deferred to a separate PR. * test(#3301): anchor coverage-check extraction on the fence line, not prose `.plans-manifest.md` also appears in write_reviews' own prose ahead of the ```bash fence, so indexOf found that occurrence first and the backward-walk-to-fence-open landed on the earlier, unrelated gate-check block instead. gsd-test caught this: coverage-check tests expecting a real verdict got null, because the wrong block ran and never writes .plan-coverage-<slug>.json. Anchor on the fence-only bash assignment line instead. Emitted-Drift-Ack-Growth: review.md — #3301 adds the plan-coverage manifest and mechanical coverage check to build_prompt/write_reviews. * fix(#3301): route id escaping through the canonical pattern seam ADR-3212 (epic #3212) consolidated ~44 hand-rolled regex-escape copies into one owner, src/pattern.cts's escapeRegex, specifically to stop this exact class of duplication. My coverage-check node -e script hand-rolled the identical metachar-escape regex — invisible to eslint-rules/no-adhoc-regex-escape.cjs only because it lives inside a workflow markdown file, not a .cts/.cjs source file the shape-matching guard scans. Require the compiled seam (gsd-core/bin/lib/pattern.cjs) instead, matching the established node -e-requires-a-compiled-lib idiom already used elsewhere in this workflow (code-review.md's code-review-flags.cjs/code-review-depth.cjs calls). Verified both named traps from the issue still resolve correctly under escapeRegex's RegExp.escape-backed implementation, which differs in escaped-text shape (hex-escapes hyphens/leading chars) but not match result. * test(#3301): run coverage-check block with cwd at the repo root The block's node -e now requires ./gsd-core/bin/lib/pattern.cjs, a path relative to the repo root (correct for production, which always runs from there). The test harness ran it with cwd at the fixture's own temp dir instead, so the require failed. Add an optional cwd param to runScript (default: root, unchanged for the plan-copy-block tests) and pass the real repo root for every coverage-check call site. Manually verified end-to-end before spending another remote run: the extracted block now produces the expected {complete:true} verdict. * docs(#3301): backfill changeset pr number (pr:0 -> pr:4084) --------- Co-authored-by: sim <sim@local> |
||
|
|
4d70b4dc43 |
fix(#4068): add --merge-async to c8 coverage-merge invocations (#4069)
* test(#4068): failing-first regression guard for c8 --merge-async flag Adds a config-invariant test asserting test:coverage:unit and test:coverage:report pass --merge-async to c8, guarding against the coverage-merge OOM (release.yml finalize dry-run, exit 134/SIGABRT) regressing a third time (prior stopgap: #199). Also commits the sourced research memo backing the diagnosis. * test(#4068): rename regression test to avoid lint-test-file-count collision coverage-merge-async-flag.test.cjs's effective prefix (coverage-merge- async-flag) matched the unrelated gsd-core/bin/lib/coverage.cjs module's 2-file cap under lint-test-file-count.cjs's startsWith bucketing, failing FAIL_EXCEEDS_LIMIT (confirmed via the RED gsd-test run on 7b6e6ca9e). Renamed to c8-merge-async-flag.test.cjs -- no colliding prefix. * fix(#4068): add --merge-async to c8 coverage-merge invocations The `test:coverage:unit` script OOM-crashed (exit 134, SIGABRT) in the release.yml finalize dry-run of 1.12.0: all 1785 unit tests pass, then c8's report/merge phase crashes ~76s later against the 6144 MB heap ceiling. Root cause (verified against this repo's pinned c8@11.0.0 source, node_modules/c8/lib/report.js): the default sync merge path, Report._getMergedProcessCov(), loads every raw V8 coverage file for the whole run into memory as one array before merging. This is a recurrence of #199 (same OOM class at ~466 tests, "fixed" by raising the heap ceiling) -- the suite outgrew 6144 MB as it grew to 1785 tests, and test.yml's own coverage-gate job already independently hit and stopgap-fixed the identical class once (test.yml:614-620, 4096->8192 MB). c8 ships the upstream fix for exactly this: --merge-async (c8 v7.14.0, already inside the pinned c8@^11.0.0 range -- no dependency bump), which switches to Report._getMergedProcessCovAsync(), reading and merging one raw file at a time instead of loading them all at once. The merge arithmetic (mergeProcessCovs()) and --all's zero-coverage-file inclusion are unchanged between the two paths, so the coverage gate's accuracy is unaffected. Verified empirically (not just by source inspection): against a 223 MB synthetic raw-coverage corpus derived from this repo's actual 235 gsd-core/bin/lib/**/*.cjs files (same order of magnitude as the 358 MB figure documented in test.yml), c8's real Report class crashes with the same "Ineffective mark-compacts near heap limit" signature via the sync path at a 300 MB heap ceiling, while the async path completes at a flat ~40-50 MB peak down to a 100 MB ceiling. Full repro steps and numbers: .gsd/bug/fix-4068-coverage-merge-oom/10-diagnosis.md. Only `test:coverage:unit` and `test:coverage:report` are changed. `test:coverage` (line 150) is a non-CI dev-convenience script, never invoked as a command by any workflow. `test:coverage:unit:raw` (line 153) already runs --reporter none, which skips the merge/report phase this flag affects, entirely -- adding it there would be a no-op. Fixes #4068 * fix(#4068): add changeset fragment for the coverage-merge OOM fix The changeset fragment was generated locally (npm run changeset) but never committed -- isolated code review caught the gap (no .changeset/*.md on the branch, confirmed via git diff --name-status). * chore(#4068): backfill changeset PR number pr:0 -> pr:4069 --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
472f585f7c |
fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates `milestone complete <version>` is a one-way door — ROADMAP.md and REQUIREMENTS.md archived, every phase directory in the milestone MOVED, STATE.md rewritten — and ran unconditionally on first invocation through every invocation path, including `query milestone.complete <version>`, whose `query` meta-prefix reads as a read-only namespace but performs no filtering (#167's invocation-compatibility shim + #3243's dotted-form normalization). The gate lives on the destructive command itself, not on the `query` prefix (the prefix is an intentional invocation mechanism, not a permission boundary — restricting it would break dozens of shipped workflow callers). Without --confirm and without --dry-run the command now refuses via error() before reading anything beyond its arg checks, so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run still previews with no confirmation needed and is now documented in the usage block (it was only documented for the sibling archive-quick). --force keeps its narrow meaning — bypassing the TRUNCATED-scope and unstarted-phase guards — and does not double as the mutation opt-in. --confirm follows the existing `phases clear --confirm` idiom in the same module. complete-milestone.md's two invocations pass --confirm (the workflow has gathered explicit user intent by that step). Existing tests get --confirm appended — pre-change behavior is exactly confirmed behavior — and a #3726 regression block covers: refusal + full-tree byte-identity on both invocation forms, --force not satisfying the gate, --dry-run still passing without confirmation, and --confirm proceeding. The refusal tests fail against pre-fix code (negative control run). Fixes #3726 * docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS Cross-AI review of the fix diff (codex, pre-create) caught three shipped doc sites still instructing the now-refused bare invocation: the CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's two guard-override instructions (`--force` alone now refuses without --confirm). Localized CLI-TOOLS copies already lag the English synopsis (no --force/--dry-run either) and follow the translation pipeline, not this fix. * chore(#3726): set changeset fragment pr to 3774 * test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack Two CI reds from the --confirm gate, both this branch's own misses: - tests/qa/scenarios/milestone-rollover.json invoked `milestone complete 1.0 --force` as a JSON arg-array fixture — a caller shape the test sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never enumerated. Adds --confirm; the scenario's boundary-crossing contract is otherwise untouched. - complete-milestone.md's +420-byte --confirm note trips the emitted-attribution growth ratchet. Acknowledged as a #3726 append to the existing complete-milestone.md entry in 3409-unreachable-guard-arms.json (two ack sources may never name the same path, per that fragment's own precedent). Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed. * docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm Review Major 1: the truncated-window and unstarted-phase guard paragraphs still told the reader to "Pass `--force` to override", which now refuses (--force alone does not satisfy the confirmation gate), while the flag table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair so the file no longer contradicts itself. * docs(#3726): synopsis renders --confirm and --dry-run as alternatives Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as "a dry run still needs --confirm", the opposite of AC 3. Render the pair as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage docblock, and let the flag rows carry the rule. * test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack Rebase onto next (26 commits) surfaced three tests the gate now refuses: the #3685 write-flag contract pair in tests/milestone.test.cjs and the `milestone complete` boundary fixture in tests/state-contract.test.cjs all invoke the command bare. Each now passes --confirm (a mutating run is exactly what they assert on). The +420 byte complete-milestone.md growth ack rode on 3409-unreachable-guard-arms.json, which #3078 swept from next as fully spent — hence the modify/delete conflict. Re-filed under a fresh fragment named for this issue, never resurrecting the swept one. * test(#3726): pin the present-but-falsy arm of the confirmation gate Review Minor 1: the boundary triple covered absent and present but not present-but-falsy. The gate is an exact-token match, so --confirm=false and --confirm=0 refuse today — pinned (canonical + query forms, whole .planning/ tree byte-identical) so a future `=`-aware or prefix-matching parser cannot silently turn --confirm=false into a confirmed run of an irreversible command. * test(#3726): drop --confirm from dry-run-only invocations Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run invocations that never needed it, so each stopped standing as incidental proof that a preview needs no confirmation. Reverted to the pre-PR form; the dedicated AC-3 test carries the explicit assertion. * docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate REQ-I18N-02 (docs/features/internationalized-documentation.md) requires translations to stay synchronized with the English source. The four localized CLI-TOOLS.md guides still advertised a bare `milestone complete <version>`, which now exits 1. Render the English synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and `[--archive-quick]` flags the translations had also fallen behind on. * test(#3726): drop --confirm from the remaining preview-only invocations Round 2 reverted the --confirm appends on --dry-run-only invocations in tests/milestone.test.cjs, but four more sat in two files the sweep missed: tests/milestone-archive.test.cjs (three) and tests/milestone-window-single-owner.test.cjs (one). Each is a preview run whose whole purpose is to document that a preview mutates nothing, so `--dry-run ... --confirm` contradicted the semantics the test exists to pin. Dropping the token restores each as incidental proof that a preview needs no confirmation; the dedicated AC-3 test keeps the explicit assertion. No assertion added, relaxed, or removed — the change is four tokens. * chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer #3954 (ADR-3942) moved emitted-drift acknowledgments out of tests/emitted-drift-acks/ and into git commit trailers, and the fragment directory no longer exists on next. The reason this PR's fragment carried moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit; the fragment file is removed rather than resurrected. Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift. * fix(#3726): name --confirm in the version-required refusal The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke the command without args and the error lists what is required) stopped at `version required for milestone complete (e.g., v1.0)` — one required argument short. Discovering --confirm took a second round trip through the gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test that also asserts the version-less invocation leaves .planning/ untouched. * test(#3726): pin the milestone complete docs against a silent regression The changeset is `type: Fixed`, which the docs-required lint exempts, so nothing in CI would notice a later edit that reinstated the bare-`--force` override prose or dropped `--confirm` from the synopsis. Four tests in tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md and its four localized mirrors; the `--confirm` flag row; both guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md, by guard name (a substring match on each instruction's `--force --confirm` text); and — as an identity ratchet over the milestone-complete sections — every `--force` sentence or clause that lacks `--confirm`, so a new bare instruction in its own sentence or clause fails whatever its wording. Named residual: a bare instruction spliced into the same clause as a compliant one coalesces with it and passes the ratchet; the by-name pins are what keep the four known instructions from losing the pairing that way. The file is registered in scripts/docs-guard-registry.cjs so the pin runs on the PR that changes those docs, not only after merge. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
370cfc6680 |
enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local> |
||
|
|
519ac23ebb |
fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first) * docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts; its exit-2 block IS the contract — a post-tool hook cannot prevent a commit) and gsd-session-state.sh is registered SessionStart (session orientation, not post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the three INVENTORY locales; the issue asked for a neighbouring-row scan, which is how the session-state row was found. All other rows in the four tables verify against the surface. * fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins Adversarial review found the same two wrong rows shipped in five more files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now covers all ten shipped tables. The parity parser unioned only the Kimi mirror list, silently exempting agent-isolation-guard (registered via the dynamic preToolEvent push): probes are now parsed too, with bare hook names resolved against hooks/ ground truth and dynamic event variables resolved to their canonical (non-Gemini) events; an exact-set pin replaces the loose size guard. allow-test-rule marker carries the issue ref; unverified-ceiling 280→281 (audited: the new marker is legitimate — the suite reads product docs whose text is the contract). * fix(#3839): register the hook-table parity suite in the docs-guard lane The new suite reads ten docs/ paths, so lint-docs-guard-registration requires it in the docs-guard registry — the first GREEN bench run caught the omission (the RED run's docs-guard failures were the same signal, previously misread as marker fallout). * chore(#3839): changeset fragment (pr number backfilled after PR creation) * chore(#3839): backfill changeset PR number (4041) --------- Co-authored-by: sim <sim@local> |
||
|
|
80de48c319 |
enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs
Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.
local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.
Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.
The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.
Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.
Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.
Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.
Verification runs on the remote runner.
Refs #3914
* docs(#3914): the epic added three guards, it did not remove one
The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.
An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.
Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.
Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.
Refs #3914
* fix(#3914): make the successor actually supersede before retiring the predecessor
An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:
const exit = 'exit';
process[exit](1);
n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.
This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.
The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.
The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.
The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the two exit rules are complementary — keep both
Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.
I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:
function f(exit) { process[exit](1); } 0 vs 1
let exit='exit'; exit='exit'; process[exit]() 0 vs 1
const { exit } = ...; process[exit](1) 0 vs 1
plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.
The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.
So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.
What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.
The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.
Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the exemption term was a no-op — the net is +4
Fourth correction to this ledger, and a fourth error of the same kind.
Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in
|
||
|
|
ac7587287b |
fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true #3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue was reopened rather than noted. The section now states three things: every field is single-valued, the section is overwritten rather than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line appended in good faith is silently ignored rather than winning. Progress history belongs in ## Performance Metrics, two headings down, and the text now points there. The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before being written rather than inferred from the issue title: stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase") -> "1 of 5 (First)" The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags im with NO g, so String.match returns the first hit. Writing "first wins" without running it would have repeated the exact error I had to retract twice in this epic already. A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines and asserts it resolves independently - negative space, because a reader returning the first line of the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a last-match variant returns "9 of 9 (Appended later)" and T1 reds. No assertion checks that the document contains a sentence. That is what local/no-source-grep exists to stop, and it would pin wording that is allowed to improve. The point of the test is that if that regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly becoming a lie with nothing to notice. Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four locales. The constraint is structural, not stylistic - confirmed by running that comparison after the edit. The four locale copies are translated rather than left stale. They are not gate-enforced for prose, so "nothing fails" was available and is not the same as correct: leaving four documents asserting something the English one now contradicts is a correctness problem. Code spans and the anchor link stay untranslated - they name real tokens. The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been false the fix would have belonged in the generator, and a hand edit would have been silently reverted. One real gate failure fixed inline rather than reported: the new test's comments referenced docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and lint-docs-guard-registration failed lint:ci correctly. Registered. Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope and stays unclaimed. Closes #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): the rule I documented was false — replace it with the measured one An isolated review returned two blockers. Both mine, and the first is the worse kind: I wrote a falsifiable rule into a reference page and got it wrong. 1. "A duplicate resolves to the FIRST occurrence" is FALSE. stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence beats document order. Measured against the built reader, all intra-section: Phase: A (plain) / **Phase:** B (bold, later) -> B LATER WINS Phase: A (indented) / Phase: B (plain, later) -> B LATER WINS Phase: A (plain) / | Phase | T (table) | -> A first wins My original verification tested plain-versus-plain, saw first-wins, and generalized to all forms. Measuring one case and claiming the general rule is the same error I have had to retract twice already in this epic. It is also worse than silence. The sentence told authors an appended line is safely ignored; a bold line appended "for emphasis" silently overrides the original. Someone trusting the doc would have corrupted their own state file. And #3812 never asked for a resolution rule - it asked for single-valued, overwrite-not-append, and where history goes. The rule was my unrequested addition. Replaced with the measured truth: resolution is by FORM (bold anywhere, then plain at line-start, then table row), and only WITHIN the winning form does the first occurrence win. Both consequences stated plainly - a higher-ranked form wins regardless of position, and an indented `Phase:` is invisible to the plain form. All five claims in the new paragraph verified by execution before being written, including the two I had wrong. 2. The tests tested the wrong case and passed for the wrong reason. T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section case, which #2956 already fixed by scoping. #3812 says verbatim that #2956 "fixed the inter-section case and never addressed intra-section duplication", so the case the new prose describes was untested, and the fixtures passed because of section scoping rather than field resolution. They also called bare stateExtractField rather than the production chain, T2 could not discriminate first from last at all, and no fixture mixed forms - which is precisely why the false claim survived to review. Rewritten as four rows, all intra-section, all through the real stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first wins within a form), plain-then-bold (the bold LATER value wins - the row whose absence let the false claim ship), indented-then-plain (indented invisible), and sibling-field independence. Each proven to fail against a reader that disagrees. 3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the anchor their own heading generates. ja-JP/ko-KR kept the English heading, so theirs already resolved. 4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section designed to grow" with a bare demonstrative whose nearest referent read as Current Position - saying the opposite of the point. Rewritten so the clause attaches unambiguously to `## Performance Metrics`. 5. Cross-locale drift, flagged by the implementing agent rather than by me: after fixing EN, the four locales still stated the OLD false rule. Four documents asserting something measured to be wrong is worse than four saying nothing. All four now carry a faithful translation of the corrected paragraph, with code spans, each file's own anchor, and the ja/ko referent fix preserved. Verified: all five claims executed against the built reader; every rewritten test row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to date, so the edits stay outside the generated marker regions; locale heading parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): second false rule on the same page — scope the ranking to the section A second isolated review found a second false falsifiable claim, and the failure mode is the same one twice in a row: attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT Both times the claim covered a wider surface than what was actually executed. The fix each time was not a better sentence, it was executing the surface the sentence describes. BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false. ## Current Position / Phase: 1 of 5 + ## Archive / **Phase:** 88 bare stateExtractField(whole doc) -> "88 (other section)" PRODUCTION (slice then extract) -> "1 of 5 (in section)" #2956's section slice means production never hands another section to the matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not seen. The ranking is real but scoped: it applies WITHIN `## Current Position`. I verified against the bare function and wrote a claim about the system. Every existing test placed its bold line inside the section, which is exactly why nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive` and asserts production returns the in-section plain value, with the unscoped reader asserted to DISAGREE so the row proves the scoping rather than assuming it. BLOCKER — the changeset still shipped the ORIGINAL retracted claim. I corrected the page and left the release note saying "resolves to the first occurrence ... a second entry added in good faith is silently ignored". The note contradicted the page it announces, and the release note is what most people actually read. Rewritten to the corrected rule. MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the file", which is the vacuous direction; the surprising case, and the one the very next clause illustrates with an APPENDED bold line, is "even if it comes LAST". All four locales reproduced the inversion faithfully, so it was an EN-source defect rather than translation drift. Two sharp edges now named, both measured: a bold `**Phase:**` followed only by trailing spaces resolves to an EMPTY STRING and does not fall through to a valid plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form ranking has to say where the ranking bites. Also fixed: all five files labelled the link `## Performance Metrics` while the heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only the label's level was wrong. Every clause in the final paragraph re-verified through the PRODUCTION chain (stateCurrentPositionSlice -> stateExtractField), clause by clause, before being written: bold in another section does not win; bold in frontmatter does not win; bold appended last does win; first wins within one form; trailing-space bold yields empty. All four locales carry the same corrected rule. gen-state-md-docs --check reports all 6 targets up to date; heading counts unchanged at 20/20 across all five files, so locale heading-parity is untouched; build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3812): backfill changeset pr number Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ab69b9ce56 |
enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from |
||
|
|
dd4f179672 |
feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
12f9d1d9a0 |
enhance(#3913): docs, and the guards come down (#3994)
ADR-3889 terminal phase. Generated docs/reference/exit-codes.md from the exit-code declaration with a --check drift arm; deleted the inert soft-error-exit-zero oracle; promoted untyped-success from SMELL to VIOLATION so it can fail a build; pruned all 5 smell-baseline entries. Fixed inline: two mis-scoped oracles (routing-validity, value-hygiene), a second source behind the band table, unescaped declaration strings reaching Markdown, and a pre-existing Windows 8.3 short-name path-comparison defect. Guard ledger corrected from a claimed net -4 to a measured net -1. Closes #3913 |
||
|
|
b351c83e03 |
docs(adr): ADR-3646 — per-task external-tracker content-resolution seam (#3991)
Resolves the four conditions on the approved-feature verdict for #3646: new execute:task granularity tier below wave, hard-halt enforced via a code-side resolver seam (Lens B) rather than prose dispatch (since #3647's dispatch-reliability defect is still open), registration/validation requirements for loop-hook-dispatch.md + capability-validator.cjs, and autonomous-mode behavior via a new non-gate kind. Closes #3969 Co-authored-by: sim <sim@local> |
||
|
|
3a6c0412a9 |
enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) Extends ADR-1703's portability rule catalog with a second production-runtime rule: it flags an exact-case read of a Windows case-varying environment variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any receiver that is not process.env itself, matched via an env-shaped-receiver check to avoid colliding with ordinary `.path`-named properties elsewhere in the tree. Exports the seam's private `_envGet` as `envGet` so the rule's remediation message names a real helper, and fixes the one pre-existing violation the tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read). Closes #3624 * fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case Review findings from the code-review + isolated-adversarial passes: - extractStaticName only matched non-computed Identifier keys, so a destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8 acceptance case) silently evaded the rule. Widened to accept a Literal key regardless of computed, which is safe for MemberExpression too (its non-computed property is always an Identifier by grammar). - Added the missing RuleTester valid case for "a case-insensitive accessor call" (envGet(env, 'PATH')) from the issue's Done-when checklist. * docs: backfill changeset PR number for #3624 (PR #3976) --------- Co-authored-by: sim <sim@local> |