eb336e9f777352d0e5682eec244d7ccdbda7a7d5
5635 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eb336e9f77 |
fix(#4081): decode git C-quoted paths in codebase-drift --name-status parser (#4307)
* test(#4081): failing-first regression for quotepath C-quoted paths in codebase-drift * fix(#4081): decode git C-quoted paths in codebase-drift --name-status parse * test(#4081): set drift_threshold 1 so decoded-path test triggers action_required * chore(#4081): add changeset fragment * chore(#4081): fix changeset fragment formatting * chore(#4081): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
9ee6d54cc3 |
fix(#4306): extend fault-injection fd-swallow fix across the whole suite (#4308)
* fix(#4306): forward real bytes through io.test.cjs's fault-injection mocks The bug #1008 fault-injection tests mock fs.writeSync scoped only by file descriptor. On their "success" arms (the retry-after-EAGAIN/EINTR call, and the short-write simulation) they fabricated a return byte count without ever calling the real writeSync -- the bytes went into a local array and nowhere else. node:test's process-isolation runner (default on Node >= 22) reads each test file's own stdout to parse its child-to-parent result protocol. If the runner's own reporter write for an adjacent test lands on fd 1 while one of these mocks is installed, that write was silently swallowed instead of reaching the real pipe -- observed in CI as "Unable to deserialize cloned data" (a corrupted/truncated byte stream on the parent's read side), not a thrown exception. Every "success" arm now forwards the real bytes to orig()/restore() instead of fabricating a return value, so anything else sharing the fd during the mocked window still gets its bytes delivered for real. writeAllSync (the only production caller reaching this mock) always passes a Buffer, so the forwarded calls use the buffer-form fs.writeSync overload unambiguously. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4306): extend fault-injection fd-swallow fix across the whole suite The originally-fixed instance (tests/io.test.cjs) was one occurrence of a copy-pasted defect: mocked fs.writeSync arms fabricated a return byte count without ever forwarding the call to the real fs.writeSync, silently discarding bytes. Under node:test's process-isolated runner, the parent reads the child's real stdout to parse v8-serialized report frames interleaved with plain output (confirmed against node's own lib/internal/test_runner/runner.js and a matching upstream issue, nodejs/node#64061) — a swallowed write on that fd corrupts the parent's parse ("Unable to deserialize cloned data"). Adds a shared, safe capture helper to tests/helpers.cjs, captureFdSync(fd, fn): it always forwards every write to the real fs.writeSync first, then records only the observed fd's bytes, sliced by the real return count (not the requested length), decoded once via Buffer.concat so a short write can't split a multi-byte codepoint across two decodes. 17 test files migrate their local copy of the unsafe mock to this shared helper. tests/worktree-base-ref.test.cjs keeps a narrower in-place fix instead (it needs to record every fd a write touched, which the shared helper doesn't expose). tests/io.test.cjs gets two follow-up correctness fixes on top of the already-committed forwarding fix: the EAGAIN/EINTR/short-write arms now derive their recorded chunk from the real return count everywhere (including the string-form overload), and the short-write test no longer forces a Buffer-shaped truncation call onto a string-form write that could land on the same fd. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4499933807 |
fix(#3802): resolve the heredoc body before validating the commit subject (#3816)
* fix(#3802): resolve the heredoc body before validating the commit subject With hooks.community: true, gsd-validate-commit.sh blocked EVERY heredoc-form commit with CONVENTIONAL_COMMITS_VIOLATION regardless of the message, including Claude Code's own documented idiom: git commit -m "$(cat <<'EOF' feat(auth): add login flow EOF )" Reproduced before changing anything: conforming heredoc -> exit 2; plain -m "feat(auth): add login flow" -> exit 0. Root cause is the extraction regex `-m[[:space:]]+"([^"]+)"`. Bash `[^"]` matches newlines, so the capture ran from the quote after -m to the FINAL quote at `)"`, swallowing the whole span. `head -1` then returned the literal `$(cat <<'EOF'` as the subject, which can never satisfy Conventional Commits. Fixed by not answering a regex bug with another regex. hooks/lib/git-cmd.js already exists because "a naive regex misses all three" invocation forms, and extractBranchArgument is the established precedent for pulling an argument off a git command line. extractCommitSubject joins it on the same tokenizeShellLike seam — which, checked first, already returns the entire heredoc span as ONE token, leaving only "resolve the body to its first line" as new logic. Because the walk starts at the subcommand, `git -C <path> commit` and env-prefixed invocations now extract correctly too — forms the raw string scan never handled. Deliberately unchanged, and pinned as such: a glued `-mfeat: x` and `--message=...` still yield no message, exactly as the regex left them. The fix stays scoped to the reported defect rather than widening on a true observation. Two things I got wrong and corrected by measuring rather than reasoning: - I expected `git commit -m ""` to be blocked. Checked against the ORIGINAL hook: allowed before, allowed now, identical. The scanner drops the empty token so it takes the null path. My expectation was wrong, not the code. - That exposed a false comment I had just written, claiming the exit-status split prevents silently allowing `-m ""`. It does not. The split IS load-bearing, but for a heredoc whose body's first line is blank, which resolves to an empty subject and is correctly blocked. The comment now names the real case and records that `-m ""` is not it. Tests at both layers: 9 unit rows on extractCommitSubject beside its sibling in tests/worktree-safety.test.cjs, and 5 behavioral rows piping real PreToolUse payloads through the hook in tests/hooks-opt-in.test.cjs. Replacing firstLineOfMessageArg with a plain first-line return reds 8 of them across both files. (A first mutation attempt silently no-opped and reported green — the mutated body is echoed in the transcript for the run that counted.) Out of scope, per the issue: the hooks.commit_types config surface, split off by the maintainer as #3811 and explicitly sequenced after this. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): confine the fix to heredoc resolution, closing four regressions Codex review of the first attempt. It was right, and the finding is one my own rules already name: a true observation is not a licence to widen the diff. The first attempt replaced the shell's `-m` extraction with a token walk. That looked like the better abstraction — this module exists precisely because a naive regex misses invocation forms — but selecting WHICH argument is the message was never the defect, and changing it regressed four forms that upstream allowed, plus opened a bypass: - `git commit -- -m WIP` -- introduces pathspecs; `-m` is a path - `git commit --amend && echo -m WIP` a later command's flag became the message - `git commit -m "" --allow-empty-message` the shared scanner drops empty tokens, so the next flag became the message - `git commit -m WIP` unquoted argument - `-m "WIP notes <<EOF\nfix: smuggled subject"` was ALLOWED — the opener was recognised unanchored, so validation skipped past the real, non-conforming subject. An enforcement bypass, not a misclassification. Now confined to the actual defect. The shell's `-m` capture is restored byte for byte, and only the subject-from-message step is delegated, to a PURE STRING helper `resolveCommitSubject()` that never tokenizes. Verified as a differential against the upstream hook run inside the real tree: the only behaviours that change are the two intended heredoc rows (2 -> 0); all four forms above read identical, and the bypass case blocks. That differential also corrected my own control. An earlier comparison ran the upstream hook from a scratch directory, where its `lib/` could not resolve `../../gsd-core/bin/lib/token-scanner.cjs`, so the classifier failed open and reported exit 0 for everything. That made a real regression look pre-existing. Re-run inside the tree, `<<-"TAG"` (a double-quoted tag nested in the double-quoted argument) is genuinely pre-existing — the capture truncates — and is now recorded as a known limitation rather than silently "fixed". Also fixed from the review: - `<<-` strips leading TABS from body lines; returning the raw line blocked a conforming message. - a non-identifier tag such as `END-MSG` is a valid bash word and was rejected. - an immediately-following terminator is an EMPTY message, not a subject. - a node/library failure now falls back to the previous `head -1` instead of skipping validation, so a broken extractor degrades to old behaviour rather than becoming a new silent-allow path. Tests strengthened per the review: the opener-spelling rows now assert BOTH directions per spelling, since "conforming passes" alone would also pass if the resolver returned an empty subject for a spelling it failed to parse. Added differential rows pinning the five previously-allowed forms, and a row for the bypass. Dropped two rows whose comments claimed the raw scan could not handle `-C`/env-prefix invocations — it could; the claim was wrong. Replacing resolveCommitSubject with a plain first-line return reds 9 rows across both files. (Mutant body echoed in the transcript; an earlier mutation attempt on this branch silently no-opped and reported green.) Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): keep the installed hook runtime-neutral `hooks/lib/git-cmd.js` ships into every runtime, including hermes and qwen, where tests/install.test.cjs enforces that no Claude reference leaks into the installed tree. My JSDoc named the idiom after the runtime that documents it. Reworded to describe the SHAPE rather than the vendor; the runtime is still named in the changeset, which feeds CHANGELOG.md where such references are allowed, and in the tests, which are not installed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#3802): backfill changeset pr number The fragment shipped with the documented `pr: 0` placeholder, which the changeset lint treats as always-silent, because the number does not exist until the PR is opened. Backfilled to 3816 now that it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): close the truncated-capture hole, add the required test artifacts Review round 1. Major 3 was the one that mattered, and it disproved a claim I had stated in falsifiable form — the PR body said only two behaviours change; the differential found five. Major 3 — an embedded `"` truncates the `-m` capture, so the resolver received a PREFIX of the real subject and the length gate measured the wrong string. Before this fix the whole form was blocked outright, so the gate was unreachable; the fix opened the path and then mismeasured it. A new enforcement hole, so it is CLOSED here rather than declared. Closed precisely rather than bluntly. A first attempt refused to resolve any body with no terminator, which also blocked commits whose SUBJECT was intact and whose quote sat further down the body — a false positive of its own. Truncation is only fatal to the line it lands IN, and a captured line is complete exactly when another line follows it, because the capture kept its newline. So an unterminated body whose subject line is followed by more text stays measurable; only a subject line running to the end of a truncated capture falls back to the opener, which fails the format gate exactly as this form did before the fix. Major 1 — fast-check property rows for the new parser, via the shared seeded setup helper rather than requiring fast-check directly, per repo convention: totality (a security property here, since an exception on this path fails OPEN), idempotency, and that the result is always a single line drawn from the input — the third catches a resolver that concatenated or trimmed while satisfying the first two. Major 2 — the 72-char gate is now exercised at {71, 72, 73} on the RESOLVED heredoc subject, with the fixture length asserted so a mis-built fixture cannot silently pass. 92 chars did not show which side of `> 72` the code sits on. Minor 1 — leading blank body lines are skipped, as git's cleanup=whitespace does. A conforming commit written that way was still blocked, which is the same defect class #3802 reports. Nit 1 — a backslash-escaped delimiter (`<<\EOF`) is now the same delimiter rather than failing closed on a delimiter that includes the backslash. Nit 5 — changeset trimmed from 2,208 chars of design note to the user-visible change. Mutation discipline, including a correction to my own: dropping the truncation guard reds the unit rows, and the pre-review naive shape reds the hook-level row too. My first mutant did NOT distinguish the hook row — removing the guard made an empty slice and blocked for an unrelated reason, so the row passed and looked proven. Only mutating to the actual pre-review shape showed it discriminates. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): measure the subject as git does — strip trailing whitespace, split CRLF git's cleanup=whitespace strips whitespace at BOTH ends of a line; the resolver handled only the leading direction, so a 72-char subject with trailing spaces measured 75 and stayed blocked — the defect class #3802 reports, surviving one round further (review of #3816, Major 2). The resolved subject now drops trailing spaces and tabs; the plain non- heredoc path is untouched, keeping the fix confined to heredoc resolution. The length-gate boundary rows gain dirty fixtures: 72+3 trailing spaces passes, 73+1 stays blocked on LENGTH. split('\n') left \r on every body line, so on CRLF input the delimiter never matched: the truncation guard was inert, an empty message resolved to 'EOF\r', and a real 72-char subject measured 73. Split on /\r?\n/ (Minor 3). The three property tests never reached the parser — the pinned-seed fc.string corpus contained no newline and no opener, so every property reduced to f(s) === s (Major 1). The generator now constructs heredoc- shaped input (all opener spellings, <<- tabs, optional terminator, CRLF) and each property asserts a floor on inputs its corpus actually resolved. All new rows proved failing-first against the pre-fix resolver. Also records the unquoted-delimiter expansion limit as one JSDoc sentence (Informational 5). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): close two recognition bypasses, pin the dquoted-delimiter limit Codex whole-PR review found two enforcement bypasses in the resolver: - The opener's path prefix was \S*, which accepted `id;/bin/cat` — the resolver then validated the heredoc BODY while bash runs `id` first and git's real subject is id's OUTPUT. The prefix is now a path-character class; any shell metacharacter fails recognition and the form falls back to the opener line and the format gate. - The blank-line skip used JavaScript trim(), whose Unicode whitespace class skips lines git KEEPS: a NBSP first body line resolved to the SECOND line while git's real subject is the NBSP line (verified against git stripspace — the c2a0 bytes survive). Blank is now git's ASCII space/tab only; a Unicode-blank line is returned and fails the format gate, the same fail-closed direction git takes. Both proven failing-first at resolver AND hook level. Also: the <<"TAG" spelling is recorded as a documented limit — the -m capture stops at the delimiter's own quote so the caller can never deliver it (fail closed; widening the capture would change every embedded-quote case) — with a hook-level row pinning the limit; and the derivation property no longer accepts '' unconditionally, only for heredoc-shaped input, so a conditional constant-'' regression can't satisfy the corpus floor unnoticed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): recognition whitespace is ASCII, and '' answers to the generator Codex round 2: the opener's \s accepted Unicode whitespace bash does not split on — $(<NBSP>/bin/cat was recognized here while bash reads <NBSP>/bin/cat as the executable NAME, so recognition claimed a substitution that does not run cat. Every whitespace position in the recognition is now [ \t], the same ASCII rule as the blank-line skip, proven failing-first. The derivation property's ''-acceptance now consults GENERATION-TIME metadata: the heredoc generator records whether it built an empty message (terminator reachable, all scanned lines ASCII-blank, <<- tab stripping accounted for), and '' is accepted exactly then — a resolver conditionally degrading to '' on non-empty heredocs now fails, closing the residual round-1 permissiveness without re-deriving resolver logic. The changeset no longer overstates the opener spellings: it names the capture-deliverable set and the documented <<"EOF" limit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): nothing after the terminator escapes measurement Round-3 BLOCKER: everything after the heredoc terminator was silently discarded, so `-m "$(cat <<'EOF'\nfeat: ok\nEOF\n) <200 a's>"` — one 200+ char real subject once bash substitutes — measured 8 chars and dodged COMMIT_SUBJECT_TOO_LONG, a hole the base did not have. The canonical idiom's tail is exactly one closing-paren line; any other tail now falls back to the opener line and the format gate, the pre-fix behaviour for the whole form. Proven failing-first at resolver and hook level, including the glued-text and second-substitution variants. Also from round 3: `cat<<'EOF'` (no space) is legal bash and now resolves — the token before << is still literally cat; the env-prefixed and option-terminated spellings join the JSDoc KNOWN LIMIT list instead (fail closed, modelling bash prefix words is cost with no reported user); the changeset states the embedded-quote truncation limit for the message body, not just the <<"EOF" spelling; the dquoted unit and hook rows now cross-reference each other; and the fast-check setup helper's docstring no longer claims property-file exclusivity. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): glued text outside the closing quote must not shrink the measurement Codex on the round-3 guard: bash concatenates -m "$(…)"suffix into ONE argument, but the capture holds only the quoted part — so the resolver measured the heredoc body (8 chars) for a 200+ char real subject, a net-new length-gate bypass the base did not have (base measured the opener and blocked). When the closing quote is followed by anything but whitespace or end-of-command, the hook now skips the resolver and keeps the pre-fix first-line subject: the heredoc form fails the format gate exactly as on base, and the plain single-line form keeps base behavior unchanged — both pinned as differential rows, the glued-suffix row proven failing-first against the unguarded script. The property generator's ''-oracle now models the post-terminator guard it previously predated: expectEmpty requires the FIRST reachable terminator to be followed by the one canonical closing-paren line, so a resolver regressing to '' on a non-canonical tail (e.g. a body line that doubles as an early terminator) fails the derivation property instead of being blessed by stale metadata. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: retrigger CI — the previous wave never started (Actions queue stall) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): only resolve a heredoc whose body bash does not rewrite Round-4 review found two net-new enforcement bypasses: commands the base hook blocked (exit 2) that this branch allowed (exit 0). Both reproduced as a base-vs-head differential against the real hook, not inferred. The predicate "may I resolve this?" was computed from the resolver's input string alone, while two of its determinants live outside that string: 1. WHICH -m quote arm produced the input. Inside -m '...' bash performs no command substitution, so $(cat <<'EOF' is literal text and git's real subject is the opener line. The resolver ran on both arms, so all four delimiter spellings went 2 -> 0 on the sq arm — reachable by the ordinary slip of typing ' for ". The hook now records MSG_QUOTE and gates the resolver on dq; sq keeps head -1, exact base parity. 2. WHETHER the delimiter suppresses expansion. Only <<'D', <<"D" and <<\D do; a bare <<D is expanded by bash before git sees it. Resolving the literal dodged the format gate (feat: $UNSET_VAR reaches git as feat:) and the length gate (feat: ${LONG} reaches it at any length). The opener regex now separates the backslash-quoted and bare alternatives and refuses the bare one — the same fail-closed rule the metacharacter, truncation and post-terminator guards already follow. A test row asserted exit 0 for a bare-delimiter body, so the suite defended the second bypass and the fix could not land without editing a test that read as intentional. That row and its two unit counterparts now assert the block, per RULESET.TESTS.delete-bad-tests. Two unrelated rows used <<-EOF to exercise tab stripping; they move to <<-'EOF' so each tests what it names. Scoping the adjacency guard to the matched arm — required by the fix above — also removes a spurious block (round-4 Minor 1): a double-quoted heredoc whose body mentioned a glued single-quoted token tripped the sq arm. The JSDoc claimed <<"EOF" was unreachable through the caller and that the bare-delimiter gap was pre-existing. Round 4 disproved both; both corrected here, along with the matching changeset sentence. Verified: 7 bypass commands now block at head (was allow), the #3802 fix and plain-form parity are unchanged across 8 control commands, hooks-opt-in 44/44, worktree-safety 401/401, property-test non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): resolve only where the captured text is provably git's subject Codex review of the full PR found two more inputs where the validated text is not the subject git receives, both net-new bypasses (base 2 -> head 0), plus one escalation of round-4 Minor 2. All reproduced here against the real hook and confirmed against real commits before fixing. BLOCKER — the matched -m need not be git's message. The capture is a search over the whole command and the double-quoted arm runs first, so it could select a -m that is not the subject at all. git concatenates multiple -m values and takes the FIRST as the subject, so git commit -m 'WIP first' -m "$(cat <<'EOF' … )" commits the subject `WIP first` while the hook validated the heredoc. Same for an unquoted earlier -m, for a heredoc after `--` (a pathspec, not a message), and for one belonging to a later `&& echo`. The mis-selection is pre-existing; resolving it is what made it a bypass. The hook now resolves only when nothing before the matched -m could have been an earlier message, an end-of-options marker, or another command. BLOCKER — cleanup mode is part of the predicate. The resolver skips leading blank lines and strips trailing whitespace because git's DEFAULT cleanup=whitespace does. Under --cleanup=verbatim git does neither, so a 72-char subject plus three trailing spaces is committed at 75 bytes while the hook measured 72 — COMMIT_SUBJECT_TOO_LONG dodged. This one hides from `git log --pretty=%s`, which strips trailing whitespace in its own output; the raw commit object shows 75 vs 72. Any named mode other than whitespace, in either the --cleanup= or -c commit.cleanup= form, now refuses to resolve. MAJOR — recognition trusted any path ending in /cat, so a planted `../evil/cat` printing `WIP injected` had its heredoc body validated while git's real subject was `WIP injected`. Only a bare `cat` or an absolute path is recognised now. A bare `cat` shadowed on PATH is a documented residual and is not fixable from a string — nor a meaningful boundary, since planting an executable already allows running git directly. The changeset and the JSDoc both asserted that a `"` anywhere in the message blocks. Measured false: a `"` on a later body line resolves fine, because the subject completes before the truncation point; only a `"` in the subject line blocks. The changeset also listed <<"EOF" as covered when it measures 2/2. Both rewritten to claim only what is measured, and the residual false positives are now named. Verified: 4 + 2 + 3 new bypass commands now block, with non-vacuity controls proving the default path still resolves; all round-4 maintainer blockers stay closed; the #3802 fix and plain-form parity unchanged across 7 controls; hooks-opt-in 47/47, worktree-safety 402/402, property non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): scope the cleanup-mode guard to the command outside the message The guard scanned the whole $CMD for `--cleanup=` / `commit.cleanup=`, and the heredoc BODY sits verbatim inside $CMD, so any conforming message that merely MENTIONED the token was refused, fell back to the opener line, and was blocked with CONVENTIONAL_COMMITS_VIOLATION. These are ordinary English in this repository, whose own hooks and docs discuss cleanup modes constantly. Reproduced against the real hook: `fix: document commit.cleanup=strip behavior` blocked, the same message without the token allowed (review of #3816, round 5 — BLOCKER). Scoping to $MSG_PREFIX alone, as prescribed, would have reopened the round-4 length-gate bypass the guard exists for: git accepts the flag on EITHER side of -m, and `git commit -m "<heredoc>" --cleanup=verbatim` is caught today only because the scan is command-wide. Measured, not assumed. The scan now covers MSG_PREFIX + MSG_SUFFIX — the whole command minus the one span that is message text — joined with a space so a token cannot be forged across the seam. Swept the guard class rather than the reported instance. The adjacency guard does not share the defect: an in-body `-m "foo"bar` is refused by the already-documented embedded-quote capture limit (any `"` in the subject line truncates the capture), and an in-body `-m ` without quotes resolves and is allowed. Deliberately untouched. Both directions pinned failing-first: the three false-positive rows red against the unscoped guard, and the trailing-flag row reds against prefix-only scoping. Each mutation was echoed back to prove it landed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XogDtuuuGQEfsWaLSaZCLB * fix(#3802): read commit options the way bash hands them to git Round 6 reported the adjacency guard scanning all of $CMD for a glued `-m "..."`, so a glued -m belonging to a chained-after command refused a heredoc that was never truncated. Glue is a property of the ONE character following the matched span, so that character is now the whole window. Separators and redirections are excluded because bash does not concatenate across them: in `-m "msg"&& echo hi` the argument ends at the quote, so there is no truncated capture to defend against. An independent full-PR pass then found three accept-direction defects this PR had introduced in earlier rounds, each measured against a real commit by reading the raw commit object — `git log --pretty=%s` strips the trailing whitespace that makes the length wrong and hides it: --cle=verbatim git accepts any unambiguous prefix of a long option, so the mode was set by a token that is not the literal --cleanup. 75-byte subject recorded, 72 measured. -am 'WIP first' git reads this as -a -m, so the real subject is `WIP first` and the heredoc is only the second message. The scan looked for a standalone -m. --clean""up=verbatim bash removes quotes before git sees the -""m argument, so a spliced spelling is the same option and matched no literal. The two option-name scans now read their window with quote characters removed, which is what bash does to it, and the cleanup class covers git's abbreviations. The adjacency test deliberately keeps the raw text: it asks about a literal character position, not an option name. Narrowing the cleanup window to git's own command segment was tried and reverted. `;`, `&` and `|` end a command only outside quotes, and this is a substring scan, not a parse: an unconditional trim cut the window short on `--author "a&b"`, and a quote-aware trim still cut it on `--author a\&b`. Each hid a real trailing --cleanup=verbatim and accepted a 75-byte subject. The resulting false positive — a --cleanup carried by a chained command refuses the commit — is documented and pinned instead. Refusing a commit git would take is recoverable; accepting an over-long subject is not. Sixteen rows in tests/hooks-opt-in.test.cjs. Seven mutations, including both reverted narrowings, so no dead end can be reintroduced silently. * fix(#3802): close six accept-direction bypasses in the resolve guards Round 7's FIRST-MESSAGE GUARD Major does not reproduce. Measured against the real hook in a complete tree at the reviewed head: the classifier gate runs before any guard, so `git add -A && git commit …` (git->add stops on a non-commit subcommand) and `cd dir && git commit …` (the first executable is not git) exit 0 without a guard being evaluated. The control is the proof — a subject the bare form blocks with CONVENTIONAL_COMMITS_VIOLATION exits 0 in both chained forms, so the hook never validated them and cannot be over-blocking them. The guard is unchanged; scoping this scan to $MSG_PREFIX alone is what reopened the round-4 trailing-flag bypass. The class was real, though, one shape further out: `FOO=bar; git commit …` IS classified and then refused, because assignment detection is prefix-anchored and the tokenizer does not split operators. Pinned as a counterexample and disclosed rather than generalised away; narrowing it means changing isGitSubcommand, the shared git-commit detector every gating hook uses, and it fails closed. Six accept-direction bypasses are fixed. Each let the hook resolve and ALLOW a commit whose real subject the rules refuse; the three that turn on git's recorded subject were confirmed against the RAW COMMIT OBJECT, since `git log --pretty=%s` strips trailing whitespace and hid two of them: --cleanup=whitespace -m <72+spaces> --cleanup=verbatim git kept 75 bytes -mWIP -m <heredoc> git recorded `WIP` --mes=WIP -m <heredoc> git recorded `WIP` -\m WIP -m <heredoc> git recorded `WIP` git commit --amend --no-edit \n echo -m <heredoc> echo's argument read --squash=HEAD -m <heredoc> `squash! …` Causes: one BASH_REMATCH inspected only the FIRST cleanup directive while git applies the last, so multiplicity now refuses rather than guesses at an argument order a substring scan cannot recover; the option scan required a trailing space or `=`, missing attached values and long-option abbreviations; dequoting removed quotes but not the syntactic backslashes bash also removes; the separator scan omitted newline; and --squash/--fixup have git compose the subject, so the supplied message is not the subject at all. Every fix widens refusal, the direction this file documents as recoverable. The multiplicity count first broke the hook outright: the script runs under `set -euo pipefail` and grep exits 1 when it matches nothing, which is the common case, so every ordinary commit died at exit 1 with no verdict. Guarded, and only caught because the probe runs the real hook rather than the scan. Five new rows, all five proven red against the pre-fix hook, each carrying a non-vacuity assertion that the canonical single-`-m` heredoc still resolves. Changeset corrected on three counts: "all fail-closed" was wrong (persistent commit.cleanup fails OPEN, as do the -C/-c/-F/-t message sources), "global options are all walked through" was too broad, and the chained-before claim now states what is measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9 * fix(#3802): stop the separator and glue classes matching a literal backslash Round 8's Major, with two corrections to its account. `;`, `&` and `|` are metacharacters inside `[[ ]]`, so an inline bracket class must escape each one. POSIX bracket expressions have no escape mechanism of their own, so on bash 3.2 -- the system /bin/bash on macOS, already a supported target here per the `declare -A` ban in tests/install.test.cjs -- those backslashes reach the regex engine and add a literal `\` to the class. bash 4+ consumes them, which is why this is invisible on a modern bash. The hazard is specific to bracket expressions: `\(` outside one is made literal correctly on every version, and the subject validator and the `-m` capture classes were checked and are unaffected. The prescribed fix is not taken, because it does not parse. Inline `[;&|]` is a bash SYNTAX ERROR on 3.2 and on 5.3 alike -- the backslashes exist to get the metacharacters past the `[[ ]]` parser, so removing them leaves an unparseable script. Each class is held in a variable and expanded unquoted on the right of `=~` instead, which is a plain regex on both versions. One root cause, consequences in BOTH directions. The reported half is the separator scan over-blocking. The half not reported is the accept direction, and it is the more serious: the glue class is NEGATED, so on bash 3.2 a backslash-glued suffix fell inside the exclusion and the hook RESOLVED a heredoc it should have declined -- measured exit 0 on 3.2 against the unfixed hook, exit 2 everywhere else, with a letter-glued control refused in all four cells. The reported repro is not actually fixed by this, and the changeset says so. A `\`-newline line continuation carries a literal newline, which the round-7 separator guard refuses on every bash, so that shape stays blocked with or without this change. Narrowing the newline guard is not attempted: telling a continuation from a separator by substring scan is the class that was tried twice in earlier rounds and reverted both times, and an escaped backslash sitting immediately before a real newline is indistinguishable from a continuation. Disclosed as a known fail-closed limit instead. Every new row runs under each bash on the machine. Against the unfixed hook both bash 3.2 rows go red while all four bash 5.3 rows stay green -- written the ordinary way these rows would run under PATH bash, pass against the broken hook, and prove nothing. Two non-vacuity controls per interpreter prove the validator is reached rather than passing everything. All 8 rows of the established differential harness are byte-identical before and after on both versions: no regression, no new refusal. * fix(#3802): remove the $ of a dollar-quote from the option-name scans Independent round-8 review, accept direction. The option-name windows are dequoted so they match "the command as bash hands it to git" -- round 6 removed quote characters, round 7 removed syntactic backslashes. Both passes missed that bash has two further quoting forms whose introducer is a `$`: `$'...'` and `$"..."`. Removing the quote characters alone left that `$` stranded INSIDE the option name, so `-$"m"` dequoted to `-$m` and matched no literal, while bash passed a real `-m` to git. Measured on bash 3.2.57 and 5.3.15 against a real repository: the hook allowed git commit --allow-empty -$"m" WIP -m "$(cat <<'EOF' fix: a perfectly ordinary conforming subject EOF )" with exit 0, and `git cat-file -p HEAD` recorded the subject `WIP`. The comparison that establishes this is HEAD-internal, not a differential: the same command spelled `-m WIP` is refused (exit 2). The merge-base refuses EVERY heredoc form, including a perfectly conforming one, so its exit 2 on this input says nothing about whether any guard fired -- it is the absence of the feature, not a working check. The same miss covered `$'m'`, spliced `--message`, `--cleanup`, `--squash` and `--fixup`. An option NAME finished by a command substitution -- `--clean$(printf up)=verbatim` -- is a different problem and gets its own guard: bash runs a program to complete the name, so the argv git receives is not derivable from this string at all, and resolution is refused rather than guessed. The guard is scoped to the NAME: the class is a `-`-leading token whose characters up to the substitution contain no `=`. A substitution supplying a VALUE -- the ordinary `--author="$(git config user.name)"`, spaced or glued, in either window -- is untouched and still resolves, pinned in both directions. It is a SHAPE, not a segmentation of the command line; segmenting was tried twice in earlier rounds and reverted both times, and that reasoning stands. Both new rows fail against the unfixed tree with their own assertions, proven in a complete worktree at the previous head rather than a hook copied out of its tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): recognise a canonical cat, not any absolute path ending in /cat Independent round-8 review, accept direction. Round 4 restricted heredoc-opener recognition to an absolute path, after a relative `./cat` was measured being trusted to echo its stdin. It stopped at "absolute", so any absolute path ENDING in `/cat` was still trusted -- the same claim the round-4 reasoning had rejected one spelling earlier. Measured on bash 3.2.57 and 5.3.15 against a real commit: with an executable at `/.../fake-cat/cat` printing `WIP injected`, the hook validated the conforming heredoc body and allowed the commit (exit 0) while `git cat-file -p HEAD` recorded the subject `WIP injected`. The same command through `./cat` was already refused, which is the control that shows this is the round-4 class one spelling out rather than a new one. Recognition is now the canonical system locations -- bare `cat`, `/bin/cat`, `/usr/bin/cat` -- which is the only identity claim a string can support. `/usr/local/bin` is deliberately excluded: it is user-writable on ordinary machines, which is the plantable case this guard exists for. Anything else falls back to the opener line and the format gate: fail closed, exactly the pre-fix behaviour for the form. The pre-existing residual is unchanged and still documented: a bare `cat` shadowed earlier on PATH is indistinguishable here, and is not a meaningful boundary -- anyone able to plant an executable on PATH can run `git commit` directly. This hook stays an authoring guard, not a security control. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): an option name carrying a shell expansion is unresolvable Independent review, round 9, accept direction. Four more spellings, and a change of strategy that is the actual point of this commit. Rounds 6, 7 and 8 each tried to EMULATE what bash does to an argument before git sees it -- round 6 removed quote characters, round 7 syntactic backslashes, round 8 the `$` that introduces a dollar-quote -- and each round review found another transform that had been missed. Round 9 found four more. All measured on bash 3.2.57 and 5.3.15 against a real repository, each with the plain spelling of the same command as its control (refused, exit 2) and `git cat-file -p HEAD` for the subject git actually recorded: -$'\155' WIP hook 0, real subject `WIP` ANSI-C octal -> m -$'\x6d' WIP hook 0, real subject `WIP` ANSI-C hex -> m -`printf m` WIP hook 0, real subject `WIP` backtick substitution x= … -${x}m WIP hook 0, real subject `WIP` parameter expansion -? WIP hook 0, real subject `WIP` pathname expansion and the same class through the cleanup guard, where git recorded a 75-character subject the length gate had measured as 72: --cle$'\141'nup=verbatim, --clean`printf up`=verbatim, --cle?nup=verbatim The last two settle it. An option name finished by a PARAMETER expansion depends on a variable's value at run time; one finished by a PATHNAME expansion depends on the contents of the working directory. Neither is derivable from the command string at any level of effort, so emulation cannot be completed -- not "has not been completed yet". A fifth patch in that direction would have the same shape as the previous four. The rule is therefore no longer "normalise it and match the literal". It is: an option NAME carrying a shell expansion or quoting construct is UNRESOLVABLE, and unresolvable refuses. One rule covers every spelling above and every spelling nobody has thought of yet, in the fail-closed direction. The dequoting passes are kept rather than replaced: they still normalise the deterministic removals, so the guards RECOGNISE `--clean""up=` and `-\m` as the options they are instead of merely refusing them, which keeps the existing rows meaningful. Scope is unchanged and still pinned in both directions: the class is a `-`-leading token whose characters up to the construct contain no `=`, so a construct supplying a VALUE -- `--author="$(git config user.name)"`, the backtick spelling, `--date="${NOW}"`, a glob character inside an author string, a pathspec after `--` -- still resolves. Nine such forms are asserted to pass beside the seven that must refuse. The class is bracket-only and holds no backslash, per round 8: a POSIX bracket expression has no escape mechanism, and a backslash written inside one becomes a literal member on bash 3.2. The new rows fail against the previous head with their own assertion message, in a complete worktree with the lib built, not a copied hook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * docs(#3802): disclose and pin the two spellings the round-9 class over-blocks A scoped review of the round-9 class asked one question -- does it refuse a conforming heredoc commit that the previous head accepted -- and found two spellings that it does. Both measured on bash 3.2.57 and 5.3.15, previous head 518d97b64 exit 0, current head exit 2: git commit -S$SIGNING_KEY -m <conforming heredoc> git commit -m <conforming heredoc> -- -*.txt Disclosed and pinned rather than narrowed, for two reasons. Narrowing is not available cheaply. Dropping the bare `$` member reopens `-$xm`: with `xm=m` bash hands git a real `-m`, which is the parameter expansion bypass the round-9 commit exists to close. Skipping tokens after `--` means deciding where git's options end from a substring scan, which is the class this file has already reverted twice for opening accept-direction holes -- a `--` inside a quoted value (`--author "a -- b"`) would truncate the window and hide a real trailing directive. And the limits are narrower than they look, because in both cases the spelling a developer actually reaches for still resolves: -S "$KEY" and --gpg-sign="$KEY" resolve '-*.txt', "-*.txt", ':(exclude)-*.txt' resolve The pathspec one is worth stating precisely: a glob only reaches git AS a pathspec when it is quoted, because an unquoted one is expanded by the shell before git is executed. So the refused spelling is not passing a glob to git at all, and the spellings that do are unaffected. Refusing a commit git would take is the recoverable direction; accepting a non-conforming subject is not. That is the trade this file already makes everywhere else, and it is made explicitly here. Nine rows pin the working spellings beside the three that refuse, so a later narrowing cannot silently drop the cases that must keep working. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): join backslash-newline continuations before the resolve guards Round 9's Major, with a correction to its diagnosis. The cited bracket classes at :223 and :260 no longer exist -- round 8 moved both into SEP_CLASS and GLUE_CLASS, and a lone backslash before -m resolves (exit 0) at the reviewed head on both bash 3.2.57 and 5.3.15. What refuses the repro is the NEWLINE a `\`-continuation carries: round 7's separator guard reads any newline in a window as a command boundary, and `git commit \` newline ` -m "$(cat <<'EOF' …` was refused for that reason. Round 8 disclosed it as a fail-closed limit; round 9 calls the idiom common and the limit a Major, and it is fixed here. It was left as a limit because "is this newline a continuation" looked like the segmentation question this file has reverted twice. It is not: bash's rule is local and character-level. A newline preceded by an ODD run of backslashes is a continuation and bash removes both; an EVEN run (`\\` then newline) is a literal backslash followed by a real newline, which IS a separator. Both scan windows are joined that way immediately after they are cut from the command and before any dequote copy is derived, in three bash-3.2-safe parameter expansions: every `\\` pair is parked on \x01, any backslash-newline that remains is a lone one and is removed, then the pairs are restored. Measured on both bashes, both directions: git commit \<nl> -m <heredoc> 2 -> 0 the fix git commit \\<nl> -m <heredoc> 2 -> 2 literal \ + real separator git commit<nl> -m <heredoc> 2 -> 2 bare newline -m <heredoc>\<nl>suffix 2 -> 2 bash glues it; the glue guard sees it glued git commit … \<nl> --allow-empty<nl>echo -m … 2 -> 2 the REAL newline still separates The prescribed `[\;&|]` is not taken: a backslash written inside a bracket expression becomes a literal member on bash 3.2, which is the round-8 defect from the other side. Rows run under each bash on the machine. The fix row fails against the previous head in a complete worktree with the lib built; the four control rows were measured against that same head and were already refused, so they pin existing behaviour rather than the change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c27da5e993 |
fix(#4079): forbid ScheduleWakeup wake-tool escape hatch at background-wait sites (#4299)
* test(#4079): guard every background-wait site against ScheduleWakeup wake-tool escape hatch Content-contract tests over plan-phase.md, stall-detection-helpers.md, and manager.md: every wait-instruction site must forbid ScheduleWakeup (the host /loop tool whose partial-args call surfaces the red validation error in #4079), while the sanctioned wait mechanisms stay byte-preserved. * fix(#4079): forbid ScheduleWakeup wake-tool escape hatch at every background-wait site The orchestrator could literalize 'I'll wait' by calling the host's ScheduleWakeup tool (/loop pacing surface) with partial arguments while a background subagent was in flight, surfacing the red validation error '`prompt` is required when `stop` is not true.' GSD never sanctioned a wake call; now every wait-instruction site says so explicitly: the two synchronous stop-and-wait ORCHESTRATOR RULEs in plan-phase.md, the stall-detection-helpers.md fragment (loaded before every gsd_stall_watch wait), and both manager.md background-dispatch rules. The sanctioned wait mechanisms (blocking Agent() return, gsd_stall_watch polling, dashboard loop) are unchanged. The regression test uses the shared readFileNormalized helper (code-review finding). Emitted-Drift-Ack-Growth: plan-phase.md — #4079 guard sentence at the two stop-and-wait rules (+396 B, within the phase6 shrink-only line) Emitted-Drift-Ack-Growth: manager.md — #4079 guard sentence at both background-dispatch rules * chore(#4079): add changeset fragment (pr placeholder, backfilled after PR open) * chore(#4079): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
02ad0b91f3 |
fix(#4078): phase.complete next-phase cascade reads dash-grammar checkbox rows (#4301)
* test(#4078): phase.complete mixed-grammar roadmap picks lowest outstanding phase, not positional-last * fix(#4078): accept dash-grammar checkbox rows in phase.complete next-phase cascade Stage 2 (roadmap identity scan) and stage 3 (#2028 lowest-outstanding override) required a colon separator after the phase number, while the canonical phase lookup has accepted the bullet-house dash grammar (- [ ] **Phase N — Name**, #2199) for years. On a mixed-grammar roadmap the only parseable row above N was a later phase.add-ingested colon-form phase - positionally last - and it won the numeric-minimum vote it should never have been alone in: completing Phase 1 of 18 selected Phase 18 and skipped phases 2-17 (#4078). The checkbox branches now accept the #2199 separator class (em/en-dash, hyphen, colon); heading branches stay colon-only, mirroring findRoadmapPhaseInContent exactly. * test(#4078): align regression fixtures with slug name + checked-box semantics * fix(#4078): drop unnecessary type assertion flagged by eslint * chore(#4078): add changeset fragment * chore(#4078): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
30c40e0fe2 |
fix(#4302): bound restatement detector's deferral window structurally, not by a flat char count (#4303)
* test(#4302): add failing regression test for deferral-window paragraph leak RED: hasNearbyDeferralMarker's 500-char trailing window has no structural boundary, so a compact citation-free restatement immediately followed by an unrelated section that cites tdd.md borrows that neighbor's citation and evades restatesCycleStructurally(). Reproduces the exact shape confirmed via mutation testing against the real agents/gsd-executor.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4302): bound the deferral window at the next heading, not a flat char count hasNearbyDeferralMarker's 500-char trailing window had no structural boundary, so a compact citation-free restatement immediately followed by an unrelated section that cites tdd.md borrowed that neighbor's citation and evaded restatesCycleStructurally() (confirmed on the real gsd-executor.md, whose next section after the cycle pointer has its own independent citation 82 chars past the REFACTOR anchor). First attempt bounded at the next blank line instead, and was rejected after it broke a real case caught by direct execution before committing: execute-plan.md's FIRST RED/GREEN/REFACTOR occurrence is a numbered list's own intro sentence, separated by a genuine blank line from the list item that actually carries the citation — a blank line is not reliably "still the same statement" once list structure is involved. Bounds at the next markdown heading (`\n#`) instead, capped at 500 chars as before when no heading appears. Verified against both real files, the new regression fixture, and every pre-existing #4268 fixture/property test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4302): close a narrower same-section evasion found by orthogonal review The heading-only fix left a narrower gap: a citation-free restatement followed by a blank line and an unrelated PROSE paragraph (no heading) that cites tdd.md for an unrelated reason still evaded detection — confirmed by the reviewer via execution. Generalizes the boundary rule: any blank line whose following line is NOT a markdown list continuation is now also a boundary (in addition to the existing heading boundary), reconciling with the earlier-rejected flat blank-line bound by adding the list-continuation exception that execute-plan.md's real shape (a numbered list's intro sentence, then a blank line, then the list item carrying the citation) needs. Verified against both real files, all three restatement fixtures (heading-bounded, prose-bounded, list-continuation negative-space), every pre-existing #4268 fixture, and the fast-check property test (200 runs). Also restores an honest, updated limitation disclosure describing the one narrower residual case not resolved by this design (a citation reachable via exactly one list-item hop from an unrelated list) — matching this repo's disclosed-not-silently-deferred practice from the original #4268 PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
04ac8723b9 |
fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure (#4288)
* test(#4024): pin quantitative-criteria trap shapes for verify plan-structure Rows 1-3 and 20 of the #4024 test matrix reproduce the issue's shapes (exact grep -c counts, bulk all-N observed-failing claims) and are expected to FAIL against unmodified next: nothing judges these shapes today. Corrected-arm rows pin that each rule is silent on its own fix. * fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure Add scanQuantitativeCriteria, the third plan-discipline scanner in the cmdVerifyPlanStructure family (#429, #968). It judges criteria text in <acceptance_criteria>/<automated>/<verify> blocks against a six-rule ban list of shapes proven to be traps at HEAD: exact grep -c counts (R1), bulk all-N observed-failing claims (R2), unquoted $VAR in command position (R3), fallible git swallowed by a non-final pipeline stage (R4, warn), wc output compared by string equality (R5), and relative HEAD~N git anchors (R6; bare git diff warns). Legitimate exit: <!-- plan-criteria-allow: R# - reason -->. Pure text scan, fail open. * test(#4024): bind node:test before hook locally below the fold-point * fix(#4024): R3 command-position anchor tolerates list bullets and inline-code backticks * test(#4024): bind VERIFY_CJS locally in the unit block instead of relying on fold scope * fix(#4024): R6 argument span ends at inline-code backtick or redirection * fix(#4024): satisfy no-adhoc-markdown lint on the R4 stage-boundary regex * chore(#4024): add changeset fragment * chore(#4024): backfill PR number in changeset fragment * fix(#4024): escape backticks in regex literals so drift-lint tokenizers keep function attribution --------- Co-authored-by: sim <sim@local> |
||
|
|
1ec4b38bd3 |
test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch (#4300)
* test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch Epic #4272 Phase 5's own checklist named this deliverable ("the same class of coverage loop-walk.cjs gives the loop") separately from #4268. Adds tests/qa/tdd-walk.cjs, extracting and REALLY EXECUTING (via a real `bash -c` subprocess against a real temp fixture project) the shipped bash resolution snippets from both TDD dispatch backends — never reimplementing or grep-simulating the predicate. Proves, by execution rather than text-shape assertion: the CLI predicate and both backends agree for a type: tdd plan and a plain plan; the worktree backend's fail-closed guard genuinely halts (non-zero exit, FATAL stderr) on a missing plan file; and the tdd.md embed ternary's condition tracks the real resolved value (#3800). This is exactly the class of proof #4264/#4265 (unassigned/divergent predicate) and #4268 (static-shape checks can't see backend divergence) could not provide. Extraction uses indexOf/slice on fenced-code markers only, never a backtracking regex over whole-file text (per the #4228 incident this repo's tests already document). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4298): scrub ambient env, tighten fail-closed assertion, fix comment Standards+Spec review found: (1) executeBackendScript spread raw process.env unfiltered into the spawned bash subprocess, unlike tests/helpers.cjs's runGsdTools, which deliberately scrubs SESSION_IDENTITY_ENV_KEYS + config-location env vars before spawning (#2665) — an ambient developer/CI override could silently change what phase.tdd-applicable resolves to in a way a gsd-test bench container won't reproduce; (2) the row-5 fail-closed test asserted only `stderr.includes('FATAL')`, which would also pass if the file's unrelated ISOLATION fail-closed guard fired instead of the TDD one; (3) a docstring called the worktree backend's first fenced block a "shim preamble" when it's actually the whole ISOLATION-resolution block. Fixes: spread the exported TEST_ENV_BASE (every scrub-listed key set to '') before the two intentional RUNTIME_DIR/GSD_TEST_MODE overrides; assert the exact TDD-applicability FATAL text; correct the docstring. Re-verified by direct execution against real fixtures — all three precedence-tier cases and the fail-closed case behave identically to before the fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
747a3730d4 |
fix(#4268): harden tdd-single-statement.test.cjs against reworded restatements and backend divergence (#4297)
* test(#4268): harden reworded-restatement and backend-predicate-divergence detection tests/tdd-single-statement.test.cjs's restatesCycle() keyed on the exact literal `commit: `test({phase}-{plan})`` substring, so a reworded restatement of the RED/GREEN/REFACTOR procedure shipped green. Adds restatesCycleStructurally(), a structural (span + list-marker) detector that stays linear-scan (per the #4228 catastrophic-backtracking incident this must not reintroduce) and is proven, empirically, to flag a paraphrased multi-step fixture while not flagging the real compact citations in execute-plan.md and gsd-executor.md (#4267's legitimate pointers). tests/tdd-backend-wiring.test.cjs never compared the two dispatch backends' `gsd_run query phase.tdd-applicable` calls against each other, so a one-word divergence between them (e.g. a changed --pick flag in only one backend) shipped green. Adds a byte-identity assertion on the command-substitution content (normalized for the two backends' differing variable-name prefixes), proven to have teeth via a RED-first mutation check before asserting it against the real files. The third gap in #4268 (nothing proves TDD_APPLICABLE has a real definition) was already covered by this file's existing assertTddApplicableIsComputed (epic #4272 Phase 2, #4266) — verified by inspection, no new test needed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4268): redesign restatement detector around deferral, not length Standards+Spec review of the prior commit proved by execution that a compact, no-list-marker restatement (under the 200-char span threshold) sails past the span/list-marker-only signal. Redesigns the primary check: the actual invariant is deferral, not length — a legitimate RED/GREEN/ REFACTOR mention always names tdd.md as the authority nearby, a restatement never does. Flags when no tdd.md/canonical reference appears within a 500-char trailing window past the cycle mention, regardless of length or list-marker shape; keeps span>200 and three-distinct-list-marker-lines as secondary defense-in-depth OR-conditions. Verified independently against both real files (execute-plan.md span=10, gsd-executor.md span=14, both with a nearby deferral marker at +228/+82 chars) — no false positive, and the reviewer's exact gap class (a 189-char no-citation paraphrase) is now flagged. Also fixes: boundary coverage at the span threshold (199/200/201, isolated via a factored-out measureCycleSpan() helper), a fast-check property test proving the fix holds for arbitrary filler text, and a fragile line-match in tdd-backend-wiring.test.cjs that happened to work only because a FATAL echo message containing the same substring came later in document order than the real assignment line. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5214ad5802 |
fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication (#4295)
* fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication execute-plan.md and gsd-executor.md each cited "Red-Green-Refactor Cycle" for three facts (commit-scope contract, fail-fast rule, error handling), but only the commit-scope contract lives there. Fail-fast is in tdd.md's "Fail-Fast Rules" subsection (under "Gate Enforcement Rules") and error handling is in tdd.md's "Error Handling" section — cite each correctly. gsd-executor.md's "Plan-Level TDD Gate Enforcement" section also fully restated the gate-sequence rules tdd.md's "Gate Enforcement Rules" already owns (and covers more thoroughly, including the actual git-log validation script). Collapse it to a short pointer, matching the treatment already used by the cycle-steps pointer immediately above it. Adds tests/tdd-reference-correctness.test.cjs asserting the pointer text cites the correct section names, that those sections actually carry the guidance, and that the old gate-sequence restatement is gone from gsd-executor.md. Closes #4267 Closes #4269 * docs(#4267): add changeset for tdd.md pointer correctness fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4267): acknowledge execute-plan.md growth execute-plan.md grew 95 bytes (39766 -> 39861) from the corrected three-section citation in the #3990/#4267 cycle-steps pointer. Emitted-Drift-Ack-Growth: execute-plan.md — net +95 bytes from citing the "Fail-Fast Rules" and "Error Handling" sections by name instead of a single mis-scoped section (#4267). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4267): update PROSE_ALLOWLIST line numbers shifted by the pointer-citation edit This branch's edits to agents/gsd-executor.md and gsd-core/workflows/execute-plan.md shifted line numbers, leaving tests/no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST pointing at stale lines. Update both entries to their new correct lines (811 and 419 respectively) without changing the underlying prose. * fix(#4267): restore INVALID_RED citation, fix allowlist line shift after #3770 rebase The rebase onto next picked up #3770's already-merged fail-fast update to gsd-executor.md's plan-level gate section, which this branch's own commit collapses into a pointer. The conflict resolution kept the pointer but dropped the literal "INVALID_RED" term that tests/tdd-red-evidence.test.cjs requires gsd-executor.md to name — restored it. Also updates no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST line number for gsd-executor.md, shifted again by the rebase. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4267): fix non-matching regression-guard regex in tdd-reference-correctness The fail-fast regression guard asserted gsd-executor.md no longer contains "If a test passes unexpectedly during the RED phase" — but the actual old prose (removed by this branch's pointer-collapse) read "If a test passes unexpectedly during RED, STOP". The regex never matched the real old text, so the assertion would have passed even against the unmodified pre-change file. Caught by an isolated orthogonal review pass. Fixed to match the actual removed wording, and confirmed (via a direct grep) it is genuinely absent from the current file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4267): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2e056488d9 |
fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED) Five-case matrix: decline on unsummarized plans (regression), fire on fully-summarized phase, fail-open on unresolvable phase dir, idempotent decline, normal advance untouched. * fix(#4067): derive advance-plan phase-complete from disk, not the plan counter The phase-complete branch of state.advance-plan was decided purely by STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale counter carried into a newly planned phase, or a counter raced by wave-parallel executors, let 'Phase complete — ready for verification' land while sibling plans were still executing. cmdStateAdvancePlan now re-decides that branch from disk before the write: every plan in the Current Position phase's directory must have a SUMMARY.md (scanPhasePlans single owner, the same source state.update-progress recalculates from). Outstanding plans decline the entire write byte-identically (idempotent, concurrency-safe, counter stays display-only); an unavailable disk answer fails open to the counter-derived decision. * fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs #3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed (current-milestone) lookup first so an archived milestone's stale dir cannot shadow the live one; unscoped retry when the window cannot answer. Also restore the transform's undefined-data error semantics and extract scanOutstanding. * chore(#4067): add changeset fragment * chore(#4067): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
d29b50d696 |
fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md * fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding * fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch. * chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names * fix(#4051): review fixes — em-dash description style, split audit-fix route * chore(#4051): sync skill mirrors of execute-phase/phase descriptions * chore(#4051): add changeset (pr backfill to follow) * chore(#4051): backfill PR 4289 in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
580059251a |
fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing Red: init.progress/init.resume/init.new-project payloads carry no partial-init discriminator, and progress.md/resume-project.md/ new-project.md route an interrupted bootstrap (.planning/PROJECT.md + config.json only) to Route F / STATE reconstruction / a hard error. * fix(#4040): route partially-created .planning to initialization recovery A bootstrap interrupted after .planning/PROJECT.md (but before REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways: progress.md read it as between-milestones (Route F) or 'no planning structure', resume-project.md offered STATE.md reconstruction, and new-project.md errored 'already initialized' — a routing loop with no recovery exit. Add a shared buildInitCompletenessFields discriminator (planning_exists / requirements_exists / milestones_exists / init_incomplete) to the init.progress, init.resume and init.new-project payloads, and branch on init_incomplete in progress.md, resume-project.md and new-project.md BEFORE the legacy branches. MILESTONES.md presence excludes the archival between-milestones state, so Route F and the STATE-reconstruction path keep working. Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes. Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch. Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring. * chore(#4040): add changeset fragment * chore(#4040): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
b1b7cabfb5 |
docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry Adds one `type: "eos"` entry for a Qoder host integration and regenerates docs/registries/eos-registry.md. Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop). The integration depends on @opengsd/gsd-core, negotiates the ADR-1239 host-integration handshake, and projects GSD's agents, skills, and hook scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the China edition), merging GSD's lifecycle hooks into settings.json. Every axis is sourced from Qoder's own docs per the never-infer rule. `dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a frontmatter-declared, per-agent-definition property, and GSD's two isolation negotiation models both assume a per-dispatch injection point Qoder does not expose. Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in favor of the EoS path. Closes #4123 * docs(#4123): backfill changeset pr field |
||
|
|
75ee7b0214 |
enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate One query verb computes TDD-applicability for a plan (CLI flag, plan type: tdd frontmatter, a task's tdd="true" attribute, or the workflow.tdd_mode config default), mirroring phase.mvp-mode's precedence-cascade shape. Foundation for epic #4272 Phase 2, which wires both dispatch backends to consume it instead of restating the predicate independently. Also fixes workflow.tdd_mode, workflow.research, and workflow.nyquist_validation, which never reached cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone because loadConfig() never populates config.workflow — a dead accessor found while wiring this verb's own config read, fixed inline per the no-defer rule rather than left alongside it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4273): document phase.tdd-applicable's FEATURES.md entry Add a docs/features/ fragment for the new phase.tdd-applicable query verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left untouched: it documents /gsd-* slash commands only, and the sibling verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI reference entry anywhere in docs/ either -- only inline prose mentions in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md precedent to extend. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): stop whitelisting capability-owned config keys centrally workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are each already owned by their own first-party capability's federated config schema (the tdd/research/nyquist capabilities declare them under their own capability.json `config`), resolved via isCapabilityConfigKey. Adding them to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as the prior commit in this branch did (mirroring workflow.mvp_mode, which genuinely is central-only), declares the same key in two places at once. That collision breaks capability-loader.cts's loadRegistry composition: gsd-test caught this as 84-85 unrelated failures across capability-cli/capability-command-dispatch/capability-lifecycle test files, every one showing "unknown capability: <id>" for a freshly-installed third-party capability that should have resolved fine. Verified directly (not asserted): reverting only this file, keeping the config-loader.cts tdd_mode/research/nyquist_validation flattening and the init.cts call-site fixes from the prior commit, and re-running the exact capability install + capability set repro from tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the failure with the whitelist entries present and clears it without them. loadConfig() still surfaces all three flattened values correctly with no central whitelist entry (confirmed directly against the compiled module) — the whitelist additions were never required for the #4273 fix to work; they were an incorrect over-application of the mvp_mode precedent to keys that aren't central. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4273): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f4bf449296 |
fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat cmdAuditUat admits `human_needed` OR `gaps_found`, but parseVerificationItems had a body only for the first and returned an empty array for the second — standing on a comment deferring to `plan-phase --gaps`, a different command audit-uat never reaches. Since cmdAuditUat pushes a file into `results` only when `items.length > 0`, a `gaps_found` report did not under-report: it vanished, taking its phase's `by_phase` row with it, so a clean-looking total gave the reader no cue anything was skipped. Eligibility now has one owner (the caller) and parseVerificationItems reports what the file says. The closed-entry filter could not be built on extractFrontmatter: its array-item parser keeps only each `- ` entry's FIRST line and has no notion of nested key/value objects, so an entry's `status:`/ `resolution:` siblings never reach its output and a closed entry is indistinguishable from an open one downstream. Rather than grow a competing object-list parser — or change extractFrontmatter, whose blast radius is every frontmatter consumer in the repo — this reads the raw segment BEFORE the flattening, via the existing anchored sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps` machinery that already parses exactly this `- `-opened, indentation- continued shape. The human_needed path is byte-for-byte unchanged: same reader, same display names, same numbering, no resolved-entry filtering — pinned by a test and verified by identical CLI output on base and head. parseGapsItems keeps its narrower `status: resolved` rule so no *-UAT.md behaviour moves. Closes #3850 * chore(#3850): backfill changeset pr number for #3879 * fix(#3850): one parse per entry, one fence parser, one resolved-entry rule Adversarial review on #3879: B1, B2, M3, m5, m8 and n9. B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell 5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found` report vanished from the audit exactly as it did before this fix — this issue's own symptom, on a platform the repo already has a named defect class for. `extractFrontmatter`'s BOM+fence logic is now factored out as `frontmatterRegion` and shared. One fence parser, not two. B2 — the resolved-entry skip paired two DIFFERENT parsers by array index: `parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A block sequence written at its key's indent — ordinary, legal YAML — makes them disagree about entry count, and from the first disagreement every index names a different entry, so an OPEN entry inherits a CLOSED one's resolution and is silently dropped. That is the defect this PR exists to fix, reintroduced inside the fix. Display name and sibling fields now come from ONE parse of the raw slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as `parseYamlRegion` does, so the string is byte-identical to what `extractFrontmatter` produced. The flattened array remains the #2286 GATE, but is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment. M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both readers use, rather than two copies differing only in `result`. m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an acceptance criterion #3850 does not contain: the issue has no AC section, and its suggested fix (2) states the skip unconditionally, naming a file with 14 of 16 entries resolved. That file is `human_needed`, so the asymmetry left the reporter's own scenario over-reporting by 14. m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and says the column-0 boundary rule is now a cross-module contract. n9 — the vestigial bare block is gone and its body de-indented. Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF fixture (M4 — it survived by accident, now pinned) and the unified skip rule. Fail-first verified by running the new tests against the pre-fix build: the BOM, nested-sequence and unified-skip cases are red there. * fix(#3850): read the entries as objects, not as re-parsed display text Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1 (#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml: `parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry now flattens to `test: A, resolution: R` rather than to its first line. The original mechanism existed ONLY to work around that lossy first-line flattening — it sliced the raw frontmatter segment and re-parsed each entry by hand so a `resolution:` sibling was visible at all. With a real parser upstream that workaround is obsolete, so it is deleted rather than repaired: `sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the `splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are all gone. `frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` — the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence, same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping one step before the display flattening. `flattenObjectListItem` is exposed alongside it so a caller deriving a display name produces the byte-identical string `extractFrontmatter` would have. That collapses the review's blockers into properties of the parse rather than things this fix has to get right: - B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI. - B2 (index pairing) — there is no second reader. Display name and sibling fields come from one object. - M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers. - M4 (CRLF) — js-yaml's, not ours; verified through the CLI. Also confirmed on the rebased base, per review: #3850 still reproduces on `next` after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found` fixture), so this PR is still doing work #3707 did not do. Nothing was dropped as redundant. One behaviour note: `entryField` returns a present value verbatim and treats only whitespace-only as absent. Trimming would rewrite an author's `truth:` on its way to becoming the display name. * fix(#3850): keep every frontmatter list entry at its own row Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result to objects, and filtering COMPACTS: `parseHumanVerificationItems` then numbered the survivors by their position in the compacted array. On a list mixing object and non-object entries the non-object rows disappeared outright and the rest were renumbered — #3850's own vanishing-row defect, reached through entry SHAPE instead of file STATUS. Base never had it: it walked the display array, so every row surfaced at its own position. Renamed to `frontmatterListEntries` and it no longer filters (the name now matches what it returns). Deciding what a non-object entry MEANS is a caller's judgement; dropping it is nobody's. Both readers now walk the DISPLAY array — one element per row, the array #2286 already gates on — and consult the parsed array only for "does this entry carry a closure field?". `parsedEntriesFor` owns that pairing and checks the two lengths agree before trusting an index; all-null is the correct degradation, since over-reporting a closed row is recoverable and closing the wrong one is not. Names stay byte-identical to base for every entry shape, including a nested sequence (`[nested]`, not `["nested"]`). Same class closed in the gaps reader: a non-object `gaps:` entry surfaced nothing at all and now surfaces as `unknown`, which is this module's documented fail-safe direction (`parseGapsItems`) on a false-negative bug. Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule inlined twice; `extractFrontmatter` now routes through it, so "one fence parser" is enforced rather than asserted in a comment. Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and its "shared by both readers" comment corrected — it has one call site, and the two readers differ deliberately, each mirroring its own established sibling (`parseGapsItems` vs #2286). Documented at the divergence. Tests: `B2` asserted a name substring, so it passed while the row was mis-numbered and would have passed through outright loss; it now asserts positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim, B2c the survivors' file positions across skipped rows, B2d the gaps reader. All four fail-first against the reviewed head; 332/332 green with the fix. * fix(#3850): make status authoritative, and let the two gaps readers agree Round 4 review, all five findings. Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as closure regardless of `status:`, so `status: failed` + `resolution: "attempted retry, still failing"` vanished from the report — the silently-vanishing-item defect #3850 exists to close, reached by field combination instead of file status. Closure is now per key, because the two keys have different conventions and one rule cannot serve both: `gaps:` `status: resolved` only, byte-identical to the rule `parseGapsItems` applies to a `## Gaps` markdown section, so one authored entry cannot read closed in one reader and open in the other. `human_verification:` a bare `resolution:` still closes, since that is how verifier-written entries record it — but a readable `status:` that contradicts it wins. A single unified rule was the first draft and is wrong: it closes a frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which `parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring claims it mirrors that reader's fail-safe status handling. The contradiction guard is not a judgment call about YAML. It is the rule this codebase already applies to the same field pair: `validateResolution` (probe-core.cts) rejects a populated `resolution:` on a non-resolved status outright — "a populated payload is an authoring mistake ... Reject it so the mistake surfaces." A reporter cannot throw, so it surfaces the item. Minor 1. Direct unit tests for `frontmatterListEntries` and `flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that historically co-changes with `frontmatter.cts`. They were reachable only through `uat.cts`' readers before. Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted directly. Verified unreachable through content rather than assumed: both readers enter through `frontmatterRegion`, `extractFrontmatter`'s only extra argument gates a warning, and `normalizeParsedValue`'s `value.map` is 1:1. It is a drift alarm for a future edit to either parser, so the helper is exported for tests rather than left as the one unpinned branch. Minor 3. The vestigial `const skipResolved = true` and its dead conditional are gone. Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:` entry has no `test:` in its vocabulary — the template's entries carry truth/status/reason/artifacts/missing — so it was speculative support for a field the shape does not have, and it collided with the 1..N row numbers `parseHumanVerificationItems` assigns by array position. Not reading it makes the collision impossible; an offset would have rewritten an authored value, against `entryField`'s verbatim contract. Docs, changeset and the dispatcher docstring all stated the unconditional rule and are corrected — three prior rounds here were comment/code drift. Fail-first proven: restoring the universal rule reddens all three new unit tests and both rewritten properties. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
585a8b7f1b |
fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity * fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path * fix(#3747): scope doc evidence claim to skills discovery per adversarial review * chore(#3747): add changeset * chore(#3747): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
75bad7aedd |
test(#3936): tighten quick researcher regression coverage (#4169)
* test(#3936): tighten quick researcher regression coverage * test(#3936): restore adjacent quick dispatch coverage Assert the default researcher model and bind the executor persona check to its Agent payload. * test(#3936): make parse-list assertion wrap-safe Bound the researcher-model check to the full parse paragraph so formatting-only line wraps do not fail the regression test. * test(#3936): isolate Windows model defaults --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7eefad6f92 |
docs(#4260): record npm-audit retry/backoff enhancement as out-of-scope (#4271)
Co-authored-by: sim <sim@local> |
||
|
|
18e5cfff8a |
fix(#4250, #4260): distinguish a timed-out npm audit from a JSON parse failure, retry with backoff (#4251)
* fix(#4250): distinguish a timed-out npm audit from a JSON parse failure npm-audit-baseline.cjs's runPackageLockAudit, and the near-identical auditProductionVulns helper in npm-integrity-gate.test.cjs, both grabbed e.stdout whenever an npm audit child process exited non-zero -- without checking whether the process was actually killed by its 180s timeout. A timeout-killed process's stdout is truncated mid-write, not complete JSON, so JSON.parse threw a misleading "Unexpected end of JSON input" instead of naming npm's registry timeout as the real cause. Root-caused live during a CI investigation: npm's own status page reported degraded service, and the registry's bulk-advisories endpoint was returning 503/hanging, causing npm audit to sit until the timeout fired. Adds a shared isTimeoutKill(error) predicate (checks execFileSync's documented killed/signal fields) and checks it first in both catch blocks, throwing a clear, actionable error before ever reaching JSON.parse. The pre-existing "non-zero exit with complete JSON" recovery path is unchanged and still covered by regression tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4250): add changeset for npm-audit timeout fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4250): share the timeout-kill error message and cover auditProductionVulns Two independent review passes (standards + spec) on the first commit found real gaps: the timeout-kill error message was duplicated verbatim between runPackageLockAudit and the near-identical auditProductionVulns helper in tests/npm-integrity-gate.test.cjs (this repo's own Generative Fix Divergence anti-pattern -- shared logic across parallel surfaces with no parity check), and auditProductionVulns picked up the same production fix with zero test coverage of its own. Extracts buildTimeoutKillError(cwd), used by both callers so the message cannot independently drift. Gives auditProductionVulns the same injectable execFileSyncImpl seam runPackageLockAudit already had, and adds the matching regression tests (timeout-kill throws the clear error; the pre-existing non-zero-exit-with-complete-JSON path still recovers correctly). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4250): backfill changeset PR number to #4251 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * diag(#4250): surface captured stderr in the timeout-kill error The killed child process's stderr is buffered in-memory by execFileSync and attached to the thrown error, but nothing surfaced it -- the timeout message named the timeout but discarded the one piece of data that could show WHY npm was still running when it fired (DNS stall, TLS handshake stall, a registry-side retry loop, all look identical without it). buildTimeoutKillError now takes the killed error and includes its stderr (or an explicit 'no stderr was captured' note) in the message. This is a diagnostic improvement for the next CI occurrence, not a behavior fix -- local reproduction has directly ruled out npm version (installed the exact CI-bundled 11.17.0 and ran it against this repo: 0.49s, clean), general npm registry reachability (0.4-1.4s locally, repeatedly), and npm ci speed (2m, succeeded) as explanations for the 180s CI hangs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4260): bounded retry with backoff for npm audit calls, finish the extraction The audit backend has real, independent latency variance from the rest of the npm registry -- measured (see #4260): a bulk-advisories POST took 43.41s vs 0.20s for a plain registry fetch on the same host, and the same endpoint returned no response at all (000) twice in the same window, while status.npmjs.org reported fully operational throughout. Against that, runPackageLockAudit and its near-duplicate auditProductionVulns each made exactly one attempt with no retry -- any single bad moment failed a REQUIRED CI gate on a transport hiccup, not a real advisory. Replaces the single 180s attempt with runNpmAuditWithRetry: up to 3 attempts at 60s each (comfortably above the worst measured working latency) with exponential backoff between them. Only a confirmed timeout-kill is retried; a genuine non-timeout failure still fails immediately, and exhausting all attempts still fails the gate -- per #4260's own caveat, silently disarming a required security check on a transport error is worse than occasionally re-running CI. Also finishes the extraction #4260 flagged as stopped halfway: auditProductionVulns (tests/npm-integrity-gate.test.cjs) duplicated runPackageLockAudit's entire candidate loop, recovery branch, and timeout classification, differing only in npm args and precondition check. It is now a thin wrapper delegating to the newly-exported runInstalledTreeAudit, which shares runNpmAuditWithRetry with runPackageLockAudit -- one implementation instead of two that could independently drift. buildTimeoutKillError now reports attempt count and still surfaces captured stderr from the last kill. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4260): update changeset for retry/backoff scope Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4260): budget for two sequential retry-audit calls, close coverage gaps Two review passes on the retry/backoff commit found real gaps: - TEST_TIMEOUT_MS budgeted only one retry-audit call's worst case (210s), but checkTreeAgainstBaseline makes two sequential calls (HEAD tree via auditProductionVulns, baseline tree via runPackageLockAudit) -- combined worst case is ~372s. If both genuinely exhausted retries, node:test's own timeout would fire first and mask buildTimeoutKillError's clear message, undercutting #4250's own fix in that edge case. Recomputed using the same backoff formula the production code uses, so it can't independently drift. - buildTimeoutKillError's default-attempts(1) singular-phrasing branch had zero direct test coverage (nothing calls it with a single attempt anymore) -- a real mutation-testing risk. Added direct tests for both phrasing branches plus the no-error-object case. - runInstalledTreeAudit's null-guard skip paths (missing package.json, missing node_modules) had no tests, unlike runPackageLockAudit's matching paths. Added for parity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
97ce61dee2 |
fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally * fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD * chore(#3990): changeset for the single-statement TDD cycle * chore(#3990): backfill changeset pr number * fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap * test(#3990): allowlist pin tracks the rebased line * fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError * fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with empty stdout; the empty string survived the recovery path and surfaced as 'SyntaxError: Unexpected end of JSON input', hiding the captured error. The recovery path now requires non-empty stdout, and an empty result throws with the captured stdout/stderr/message so the actual error is on the record. Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only stops masking it. --------- Co-authored-by: sim <sim@local> |
||
|
|
590edec7a7 |
fix(#3956): require positive evidence for verify artifacts/key-links pass (#4004)
* fix(#3956): require positive evidence for verify artifacts/key-links pass An all-string or path-less must_haves.artifacts / key_links block is item-by-item skipped, leaving zero checked results, yet the pass verdict was computed as `passed === results.length` (0 === 0), so all_passed / all_verified read true with status valid and exit 0: a silent false GREEN over zero acceptance evidence. Add a positive-evidence floor (results.length > 0) to both verdicts, mirroring the no-vacuous-pass rule at src/uat-predicate.cts. A well-formed block, the fully-empty-block error, the parser's string tolerance, and key-links pending (#1202) semantics are all unchanged. Governing: ADR-3473 section 8 / 37C (absence, emptiness and failure must not encode as success) and Decision 3 (failure is a value). * chore(#3956): add changeset for verify vacuous-pass fix * test(#3956): add mixed-block coverage and correct the key-links vacuous-pass comment Addresses review on #4004: - Correct the cmdVerifyKeyLinks positive-evidence-floor comment: only bare-string items are continue-skipped; a from:-less object is NOT skipped (it falls through to a verified:false hard failure), so it was never part of the vacuous-pass surface. The prior comment overclaimed symmetry with the artifacts side. - Add a mixed-block regression test per verb (one bare-string prose bullet + one well-formed entry): the string is skipped, results.length === 1 > 0, and the verdict follows the single real entry — pinning that the floor does not over-reject a partial block. - Tighten the changeset wording to match (all-bare-string, not "no path:/from: key"). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a788afb120 |
fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives (#4021)
* fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives stateReplaceField's bold and plain patterns used `\s*` for the label-to-value gap, which matches the newline after an empty field; `(.*)` then captured the following line and the rebuild discarded it -- silent STATE.md data loss on any `state update` against an empty body field (Status:, Stopped at:, Paused at:), with exit 0 and no warning. Confine the gap to same-line whitespace (`[ \t]*`), mirroring the already-correct read side (stateExtractField, src/state-document.cts:404/:409), and pin the label-value separator to a single space when the label line had none, so an empty field yields `**Status:** value` rather than a glued `**Status:**value`. Non-empty and pipe-table replacements are byte-identical to prior behaviour. ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; this aligns the writer to it. Regression test fails before / passes after and covers bold and plain shapes, LF and CRLF, the non-empty byte-identity guard, and an end-to-end transitionCore characterization at the consumer (ADR-3180 Decision 4(c)). * chore(#4010): add changeset for the stateReplaceField empty-field fix * test(#4010): add boundary and property coverage; scope the changeset's unchanged claim Addresses review on #4021: - Add boundary tests for the shapes the example tests missed: an empty field at end-of-document (no following line, bold + plain), two consecutive empty fields (only the target is filled, the other empty field's line survives), and an empty new value on an empty field (joinFieldReplacement synthesizes no dangling separator and the following line is preserved). - Add a fast-check property over the bold/plain branches and joinFieldReplacement: for any field name, any values (empty fields included), and any new value, replacing one field changes only its own line and never the total line count — the invariant #4010 violated, now guarded directly. - Scope the changeset's "unchanged" claim to ordinary space/tab separators (an exotic vertical-tab/form-feed separator, which no GSD template emits, now normalises to a single space). * test(#4010): pin glued-separator non-empty field, scope joinFieldReplacement JSDoc Round-3 review carried forward a Minor finding: joinFieldReplacement's JSDoc still claimed non-empty replacements are unconditionally "byte-identical to prior behaviour", but a non-empty field written with no label-to-value separator (**Status:**value) gains a single inserted space under the narrowed [ \t]* gap. Round 2 scoped only the changeset prose; the source JSDoc was left making the false unconditional claim. - Scope the JSDoc's byte-identity claim to ordinary space/tab separators and name the no-separator normalization as the one intentional exception. - Add a test pinning the glued-separator case (**Status:**Planning): exactly one space inserted, following line survives, not byte-identical. Emitted .cjs is gitignored (class-1), so no emitted-drift-ack applies. build:lib clean; 74/74 state-document tests pass. Claude-Session: https://claude.ai/code/session_01Mzmut6aeqZ1APfUBAkBZTR --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
02c6955162 |
chore(#4241): add merge_group trigger to test.yml (#4242)
* chore(ci): add merge_group trigger to test.yml (#4241) GitHub's merge queue fires the `merge_group` event for the temporary merge-group commit it creates when a PR is added to the queue - not `pull_request` or `push`. Without this trigger, `required-tests` (the registered "Required tests" branch-protection check) never schedules for a queued PR, permanently stalling the queue on a check that never runs. This is workflow-side prerequisite wiring only; enabling the merge queue itself is a separate manual branch-protection step. * fix(#4241): pin AUDIT_BASELINE_REF for merge_group events too Code review on this branch caught that AUDIT_BASELINE_REF's ternary only branched on pull_request/push, so a merge_group run silently fell through to '' -- scripts/npm-audit-baseline.cjs's resolveBaselineRef() documents its origin/next live-tip fallback as unreachable from CI specifically because AUDIT_BASELINE_REF is "always set by test.yml". Reopens the exact race #4196 fixed, but only for merge-queue runs. Extends all three AUDIT_BASELINE_REF pins (test, test-inert, test-full) to also branch on merge_group, using github.event.merge_group.base_sha (confirmed against GitHub's own webhook payload schema: "the SHA of the merge group's parent commit") -- the base tip the temporary merge-group commit was built against. Adds a regression test asserting every AUDIT_BASELINE_REF pin branches on merge_group with the correct field. --------- Co-authored-by: sim <sim@local> |
||
|
|
1fe85cd43e |
chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk Repo-wide sweep (ahead of adding lint rules for these exact bug classes) found both incident patterns still live and unfixed on `next`: - scripts/run-tests.cjs's sweepProtectSet walk stopped on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. win32 dirname('D:\') is a fixed point (length 3, never satisfies `> 1`... wait, it does satisfy length>1), so a selected file living outside runTempRoot (the common case) spins the walk forever on Windows. Extracted a pure, exported computeSweepProtectSet helper that terminates on dirname(cur) === cur instead, with in-process RuleTester-style coverage for both win32 and posix paths. - tests/run-tests-temp-root.test.cjs's own #4020 regression test set only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never reads TMPDIR on Windows (only TEMP, then TMP), so the redirect silently no-oped there — masked because Windows CI died in the dirname-walk hang above before ever reaching this test. - tests/config-schema.property.test.cjs's fallow config-set test had the same TMPDIR-only pattern, direct process.env assignment this time, restored in its own finally block. Origin: #4220 and its shared root cause #4020. * feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for either shape. - local/require-full-tmpdir-triad: flags a TMPDIR environment override (direct process.env.TMPDIR assignment, or a TMPDIR property in a spawn-like call's env: object literal) not accompanied by TEMP and TMP in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows. Registered on tests/**/*.cjs, matching the require-userprofile-with-home precedent. - local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning from dirname() with no fixed-point termination guard (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a no-op at the platform root, but the value differs by platform (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped length/equality bound never fires on Windows. Registered on BOTH tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in scripts/run-tests.cjs, not tests/. Both rules join the zero-escape-hatch discipline already established for this catalog (no bespoke comment marker; PROTECTED_RULES in tests/portability-rule-disable-ban.test.cjs independently bans eslint-disable of either). ADR-1703 and its two companion contributing docs get an amendment documenting the mechanism, code examples, and the repo-wide sweep (three live instances found and fixed in the prior commit; no others found). CI test-scope selection updated so an edit to either rule or to scripts/run-tests.cjs re-runs the right suites. * fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too checkWhile bailed out early unless node.test was a LogicalExpression, so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } -- was silently skipped and never reported. That is the EXACT minimal shape of the original #4020/#4220 bug, and it is literally the shape used by this rule's own shipped RuleTester fixtures (the "equality-only bound" invalid cases), which were failing (0 errors reported, 1 expected) until this fix -- confirmed by running RuleTester directly against both fixtures, not just via a passing test-runner exit code. The conjunct-collection helper already handled a non-LogicalExpression test correctly (it pushes a single node as the sole conjunct); only the early-return gate needed to stop requiring a compound && / || test. Verified: RuleTester run directly against both previously-broken fixtures plus two new sanity cases (a guarded single-condition loop stays valid; an unrelated single-condition loop stays silent), and a fresh `npx eslint .` across the whole repo remains clean (no other single-condition dirname-walk shape exists in the tree). * fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call isSpawnLikeCallee only recognized a MemberExpression callee (child_process.spawnSync(...)) or a bare identifier in ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare -- const { spawnSync } = require('child_process'); spawnSync(...) -- has an Identifier callee named "spawnSync", which matched neither branch, so the whole env-literal check was skipped. gsd-test caught this: both "invalid: child_process.spawnSync with TMPDIR-only env" cases in tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors reported, 1 expected). Widened the bare-identifier branch to also match any of the known ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same lightweight convention this repo's other eslint-rules/*.cjs use (e.g. no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow tracing. Verified: RuleTester run directly against all 11 cases in tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that were failing), all pass; a fresh npx eslint . and npm run lint:ci across the whole repo remain clean. * fix(#4244): correct a stale escape-hatch reference in a test comment The comment on the "length comparison against another expression's length" case referenced a "// allow-dirname-walk marker" that doesn't exist -- the rule has zero comment-based escape hatches by design (ADR-1703), and an earlier draft's marker mechanism was removed before this branch's first commit. Spec-axis review caught the stale reference. No behavior change; comment-only. * chore(#4244): backfill changeset PR number (pr:0 -> pr:4246) --------- Co-authored-by: sim <sim@local> |
||
|
|
456136659d |
fix(#4220): terminate the Windows temp-sweep ancestor walk (and two bugs it unmasked) (#4245)
* fix(#4220): terminate the temp-sweep ancestor walk with a fixed-point check scripts/run-tests.cjs's sweepProtectSet block walked each selected test file's ancestor directories, stopping on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. path.posix.dirname('/') === '/' (length 1) correctly stops, but path.win32.dirname('C:\\') === 'C:\\' (length 3) never satisfies the length check, so the walk spun forever on Windows whenever a selected file lived outside runTempRoot (the common case). This has hung every Windows CI shard since #4207. Extract the walk into a pure, exported computeSweepProtectSet(selected, runTempRoot, dirnameImpl) helper and replace the length sentinel with a fixed-point check (stop when dirnameImpl(cur) === cur), which terminates correctly on POSIX, Windows drive roots, and UNC roots alike with no platform branch. * fix(#4220): repoint TEMP/TMP alongside TMPDIR in run-tests-temp-root test child env Node's os.tmpdir() on Windows never reads TMPDIR, only TEMP/TMP. The test's runNode child-process env override only set TMPDIR, so on a real Windows runner nested inside a run-tests invocation the child inherited the outer process's already-repointed TEMP/TMP and its mkdtempSync(os.tmpdir()) landed under the outer run's temp root instead of the test's intended `outer` directory. This was masked on gsd-test's benches and locally because Windows CI always died in the #4220 infinite loop before reaching this test. * fix(#4220): stop the ancestor walk from protecting the filesystem root itself computeSweepProtectSet added `cur` to the protect set before checking whether dirname(cur) === cur, so on the terminating iteration it protected the filesystem root (posix `/`, and analogously a win32 drive root) instead of stopping before adding it. Caught by the existing posix-parity regression assertion (`!protectSet.has('/')`) on the linux-node24 gsd-test bench. Reorder to compute the parent and check the fixed point before adding. * fix(#4220): backfill changeset pr number to 4245 --------- Co-authored-by: sim <sim@local> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8f013983e5 |
fix(#3751): provision the claude CLI in CI and cover agents/ in the plugin-validate fixture (#4229)
* test(#3751): the validation fixture must cover agents/ and CI must provision the CLI * fix(#3751): cover agents/ in the plugin-validate fixture and provision the claude CLI in CI * fix(#3751): wire the strict live-config guard into the plugin-validate job * fix(#3751): job-level strict-guard env, where the guard derivation reads it * chore(#3751): changeset for the CI-provisioned plugin-validate gate * chore(#3751): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
3639ab0431 |
fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated * fix(#3968): measure commit claims at all three surfaces Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain * fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin * fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument * fix(#3968): hold the gsd-executor size cap with a compact ledger contract Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) * fix(#3968): allowlist pin and HALT regex track the final prose * chore(#3968): changeset for measured commit claims * chore(#3968): backfill changeset pr number * fix(#3968): quote the BASE expansion (SC2086) * ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap) --------- Co-authored-by: sim <sim@local> |
||
|
|
d5f8191f66 |
fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical * fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick * fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run) Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction * chore(#3730): backfill changeset pr number * fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch --------- Co-authored-by: sim <sim@local> |
||
|
|
114dfcb739 |
fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones (#4214)
* fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones The #3588 gate failed on ANY advisory in the production tree, regardless of whether the PR/push actually introduced it. Because npm's advisory database updates continuously and independently of repo state, a commit could pass this gate at merge time and fail it minutes later on the identical tree -- proven on PR #4188/dce40eeb6, which passed on all 3 OSes at 15:57-16:27 and failed the same assertion at 16:19-16:30 on the unchanged commit, purely because GHSA-jqff-g426-hqxp was disclosed for fast-uri in the interim. scripts/npm-audit-baseline.cjs diffs the head tree's vulnerable-package set against a resolved baseline (the PR's target branch, or the prior commit on a direct push) and blocks only newly-introduced advisories. When no baseline can be resolved, falls back to the original zero-tolerance behavior -- fail-closed, never silently weaker. * fix(#4196): pin the npm-audit baseline instead of using a drift-prone ref Two orthogonal reviews found the same class of bug this repo already fixed once for a different gate (see GSD_EMITTED_BASE's own incident comment in test.yml): origin/<branch> is live under fetch-depth: 0 and can advance mid-run, so resolveBaselineRef()'s fallback to origin/${GITHUB_BASE_REF} could silently disagree with the tree ci-rebase-check.cjs actually merged. Wire AUDIT_BASELINE_REF from the workflow to github.event.pull_request.base.sha / github.event.before, the same pinned values GSD_EMITTED_BASE already relies on. Also: HEAD~1 assumed exactly one commit per push, which this repo's allow_rebase_merge:true setting can violate (a rebase-merged PR lands as several discrete commits in one push) -- github.event.before is git's own record of the correct pre-push state, not an assumed offset. HEAD~1 remains as a documented last-resort fallback for out-of-band invocations (e.g. gsd-test) that don't set any of the above, alongside a new local-branch fallback for gsd-test's local `next` (not origin/next) sandbox shape. --------- Co-authored-by: sim <sim@local> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
7d6d788b51 |
fix(#4020): bound the test run's temp footprint with a swept run-scoped root (#4207)
* test(#4020): the runner must bound and sweep a run-scoped temp root * fix(#4020): bound the run's temp footprint with a swept run-scoped root * test(#4020): isolate the env-mutating rows in child processes * fix(#4020): gate root removal on ownership so nested runners spare the outer root * test(#4020): pass the probe file via --files, the runner's explicit-file flag * test(#4020): resolve the probe by basename, as --files matching requires * test(#4020): assert root survival, not content survival, in the nested-row * chore(#4020): changeset for the run-scoped temp root * chore(#4020): backfill changeset pr number * fix(#4020): the sweep spares ancestors of the runner's own selected files * fix(#4020): TMPDIR precedence — an operator redirect beats inherited TEMP/TMP * fix(#4020): only the root's owner sweeps — a nested runner spares live sibling fixtures --------- Co-authored-by: sim <sim@local> |
||
|
|
858bb89769 |
ci(#4196): exempt dependabot[bot] from issue-link, title, and unsolicited-PR gates (#4203)
Dependabot has no mechanism to link a PR it opens to a repo issue -- its alerts live in the Security tab, not as issues -- so require-issue-link, pr-title-validator, and auto-close-unsolicited-prs all rejected its PRs by design (confirmed live on #4193: auto-closed for "no pre-approved issue", then flagged again by the title gate on reopen). Exempt by authenticated author login (github.event.pull_request.user.login / context.payload.pull_request.user.login), which GitHub attributes and a crafted title or branch name cannot forge -- scoped narrowly to dependabot[bot] only, no other author gets this treatment. Co-authored-by: sim <sim@local> |
||
|
|
91ed46882a |
feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives Adds the full behavioral (tests/quick-batch.test.cjs) and property-based (tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core primitives per the phase's 35-row test matrix — task-list parsing (inline + --file, with path-confinement/symlink-escape/non-regular-file rejection), collision-safe quick-id preallocation under withPlanningLock, BATCH.json schema/validation/resume, dependency-DAG + partitionByFileOverlap wave construction, and exactly-once STATE.md completion (including the STATE-row-written-but-manifest-not-yet-updated crash window). The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a not-yet-written src/quick-batch.cts) does not exist yet — every test in both files fails at the top-level require() before any assertion runs. Five fast-check properties cover collision-freedom under lock contention, resume idempotency, exactly-once STATE completion, wave totality, and DAG-respecting wave order, per the design doc's property-based-coverage requirement. * feat(#3675): implement quick-batch core primitives Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core primitives per the phase design lock — pure/state primitives and CLI-testable core operations only, no agent dispatch, no worktree creation, no user-facing command (Phase 4/#3676's job): - parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list parsing (>=2 items required) and a --file variant strictly confined to the planning workspace root via requireSafePath, rejecting non-regular-file targets. - allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id preallocation under withPlanningLock, checked against both on-disk .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json manifests (never on-disk-only, which would miss another in-flight batch that hasn't dispatched any real quick directory yet) — replicates cmdInitQuick's own grammar rather than delegating to it (that function's 2-second granularity is not batch-safe). - computeWaves: deterministic wave construction combining dependency-DAG layering with partitionByFileOverlap (#3674), called per DAG layer over path-separator-normalized planned_files — normalization happens at this module's boundary, never inside the Phase 2 helper. - loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated JSON, wrong types, missing fields, out-of-batch dependency references, dependency cycles, a worktree path absent from disk). - resumeBatch: skips complete items, never auto-retries failed items, propagates/reverses blocked status along the DAG to a fixed point, and detects a STATE.md row that already exists for a non-complete item (the "STATE written, BATCH.json not yet updated" crash window) — completing it without re-appending. Idempotent across repeated calls. - completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion — appendQuickTaskRow (unmodified) is called at most once per quick id, gated by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks Completed" table, since appendQuickTaskRow itself carries no idempotency. BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling of .planning/quick/ — never inside it, so scanQuickTasks never misreads a batch manifest as a broken quick task. * docs(#3675): register the new quick-batch module New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row plus the regenerated docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary entry matching the convention set by the sibling File Overlap Partitioner Module (#3674) entry it sits beside. NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically sorted single-entry insertion into families.cli_modules, matching the existing file's structure) rather than via `node scripts/gen-inventory-manifest.cjs --write` — this session's MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script path from Bash, with no available Memtrace tool to route through instead. The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs --check` to confirm this hand-edit is byte-identical to the generator's own output before merging. * fix(#3675): resolve lint findings in quick-batch primitives and tests Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color array, two unnecessary `as string[]` casts TS 5.5's inferred type predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch` import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a CONTEXT.md glossary illustration that looked like a real file reference. * feat(#3675): close acceptance-criteria gaps found in review Standards- and spec-axis review (plus a self-caught race) surfaced real gaps against issue #3675's own acceptance criteria and this repo's test conventions: - BATCH.json was missing options, base_revision, per-item wave, and per-item commit — the issue's AC explicitly lists all four as things the manifest must track. Added them: createBatch persists caller-supplied batchOptions/baseRevision verbatim and assigns each item its computed wave index; completeQuickItem now persists the commit onto the item, not just the STATE.md row. All four are backward-tolerant on load (an older/hand-built manifest without them still validates). - resumeBatch had no "incompatible base divergence" check at all, despite the AC and the ADR's own "Base divergence" section requiring one. Added an opt-in currentBaseRevision comparison that fails closed with a recoverable diagnostic on mismatch, and touches nothing on refusal. - resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the only durable write path in this module that wasn't lock-protected, a real lost-update race against a concurrent completeQuickItem or another resume. Now runs inside the same lock createBatch/ completeQuickItem use. - loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no size cap (security review, Low/informational); switched to the existing safeJsonParse (1MB cap) for defense-in-depth. - Parser (parseTaskList) had only example-based tests; CLAUDE.md requires a fast-check property test for parsers. Added one plus a companion reject-property for <2 items. - The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for direct limit-1/limit/limit+1 testing without needing 46k fixture dirs. - Issue AC explicitly asks for prompt-injection-payload test coverage, distinct from the existing shell-metacharacter test; added one. - Test row 9 (FIFO skip) silently returned instead of calling t.skip(), so an unsupported platform would report a pass rather than a documented skip; fixed to bind the test-context param and skip properly. - Extracted toWaveInput to remove a 2-site production duplication of the QuickBatchItem -> computeWaves reshape (Standards-axis smell). - Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/ is user-facing even though the compiled .cjs is gitignored). * fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason gsd-test caught this: switching loadBatch to safeJsonParse changed the parse- failure message shape ("... parse error — ...") without preserving the "not valid JSON" substring row 27's own test asserts on. Re-wrap safeJsonParse's error into the original diagnostic phrasing regardless of which of its three failure modes fired. * docs(#3675): backfill changeset pr number to 4190 * fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one CI caught this on windows-latest: row 9's platform-skip only caught mkfifo throwing (command not found). On this runner mkfifo resolves to something that exits 0 without creating a file (NTFS has no FIFO concept), so execution fell through to parseTaskListFromFile against a path that doesn't exist, producing an ENOENT stat error instead of the expected "not a regular file" rejection. Check the artifact actually exists before trusting a zero exit code, and skip with a documented reason either way. --------- Co-authored-by: sim <sim@local> |
||
|
|
1c4a00244f |
ci(#4196): auto-merge Dependabot patch/minor bumps once required checks pass (#4200)
Dependabot already opens a correct fix PR within minutes of a new advisory (e.g. #4193 for GHSA-jqff-g426-hqxp), but nothing merged it -- it sat until a human noticed `next` had gone red and an unrelated PR tripped over the same npm-audit gate. Auto-approve + auto-merge closes that gap for patch/minor bumps; major bumps still need a human. Co-authored-by: sim <sim@local> |
||
|
|
0598a2cf2c |
chore(deps): bump the npm_and_yarn group across 1 directory with 2 updates (#4193)
Bumps the npm_and_yarn group with 2 updates in the / directory: [browserslist](https://github.com/browserslist/browserslist) and [fast-uri](https://github.com/fastify/fast-uri). Updates `browserslist` from 4.28.2 to 4.28.8 - [Release notes](https://github.com/browserslist/browserslist/releases) - [Changelog](https://github.com/browserslist/browserslist/blob/main/CHANGELOG.md) - [Commits](https://github.com/browserslist/browserslist/compare/4.28.2...4.28.8) Updates `fast-uri` from 3.1.5 to 3.1.7 - [Release notes](https://github.com/fastify/fast-uri/releases) - [Commits](https://github.com/fastify/fast-uri/compare/v3.1.5...v3.1.7) --- updated-dependencies: - dependency-name: browserslist dependency-version: 4.28.8 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: fast-uri dependency-version: 3.1.7 dependency-type: indirect dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7c52344284 |
fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope * fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself * test(#4003): align shape assertions with the implemented gate text * fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next) * fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps * test(#4003): the gate pin tracks the anchored scope grep * fix(#4003): trim the gate rationale to hold the 93400 margin ceiling * test(#4003): the RED-grep pin tracks the milestone-bounded invocation * chore(#4003): changeset for the anchored resume-gate scope * chore(#4003): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
dce40eeb6e |
fix(#4002): rewrite zcode command @-refs to the zcode runtime home (#4188)
* test(#4002): zcode commands must rewrite at-refs to the zcode home * fix(#4002): add the missing zcode case to the runtime rewrite engine * chore(#4002): add ZCode to the bug-report runtime dropdown and drop the changeset * fix(#4002): attribute zcode command and skill ripples to the rewrite engine * fix(#4002): attribute zcode nested-skill ripples to the rewrite engine * chore(#4002): backfill changeset pr number * fix: bump qs past GHSA-x5fp-wj9c-mxmx (transitive, advisory reddened next) --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
6fdac3947b |
fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell on the watermark (#4184)
* test(#3996): antigravity diagnostic must carry stderr and gate the stall tell * fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell * fix(#3996): drop the stall token from the session-started sentence * fix(#3996): decide session-started from watermark growth, type the failure mode * chore(#3996): changeset for antigravity diagnostic stderr carry * chore(#3996): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
77dcdda534 |
enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4) * fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4) * test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift src/core-utils.cts's new #4014 import block shifted every subsequent line by 6, moving generateSlugInternal's real closing brace from line 193 to 199. tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that line number to plant a synthetic violation immediately after the function's real body; the guard script itself locates the boundary dynamically via brace-matching and needed no change. * docs(#4014): document the unreadable-directory scope signal and add changeset * docs(#4014): backfill changeset PR number to #4163 * test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff --------- Co-authored-by: sim <sim@local> |
||
|
|
2131fe13f3 |
enhance(#3464): exec() detection widening, citation-debt cleanup — Phase 8 (#4171)
* feat(#3464): widen no-source-grep to detect regex.exec() on tracked text Adds an execCall kind alongside the existing regexTest detection -- regex.exec(tracked) was invisible to the rule while regex.test(tracked) was already caught, despite both reading a source-derived string through a regex. Measured: 4 previously-invisible sites across 2 files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3464): migrate 4 sites newly flagged by the exec() widening docs-hooks-table-parity.test.cjs's three regex-extraction loops are site-scoped marked (source-text-is-the-product) -- the dynamic preToolEvent/postToolEvent dialect branching they mirror is explicitly documented as not statically parseable, so a literal-pattern mirror is the practical minimum-cost check. no-bare-gsd-tools-command-position.test.cjs's readRouterVerbs() now requires HOST_COMMAND_ROUTERS directly instead of regex-walking gsd-tools.cjs's source text -- the same accessor three other suites already use. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3464): pay down 6 grandfathered uncited allow-test-rule markers Two were genuinely load-bearing (suppressing a real detected violation) and just needed a citation added -- phase6-capstone-conformance.test.cjs, runtime-name-policy.test.cjs, both now (#3464). Four were dead-weight file-header markers suppressing nothing -- each file's real effective sites are covered by separate, already-cited markers elsewhere in the same file. Deleted outright rather than cited, per Phase 1's own precedent (remove non-load-bearing markers instead of grandfathering them forever) -- codex-config.test.cjs (two copies), gsd-check-update-worker-platform-gate.test.cjs, orphaned-hooks.test.cjs, settings-jsonc.test.cjs. allowlist.json: 134 -> 128 entries. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3464): re-baseline effective-exemption ceiling to 84 The exec() widening's 3 newly-marked sites are now suppressed and counted; ceiling rises 81 -> 84, the exact measured high-water mark. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3464): correct citation and restore a wrongly-deleted marker Two review corrections, both found by the orthogonal review pass: - docs-hooks-table-parity.test.cjs's 3 new exec() markers cited #3464 (mechanically "the phase that widened the rule") when the file's own established, correct reference is #3839 (the issue this whole test exists to enforce, already cited in its file header) -- fixed to match. - gsd-check-update-worker-platform-gate.test.cjs's deleted file-header marker was NOT dead weight: its codeOnly() helper wraps readFileSync and is called inline as an assert argument, a genuine source-grep pattern on real .cjs/.js source that the rule cannot currently see (helper-function indirection is a distinct blind spot from anything Phase 7/8 measured) -- CONTRIBUTING.md is explicit that "unverified" is not the same as "vestigial." Restored, site-scoped this time (directly above codeOnly(), not as an inert file-header comment) and cited (#3103, the issue the file's own docstring already references). codex-config.test.cjs's two deletions and orphaned-hooks.test.cjs's / settings-jsonc.test.cjs's deletions were independently re-verified and stand: their flagged lines read generated .toml/.json OUTPUT, not source, or have no residual pattern at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9b77320580 |
fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh (#4092)
* fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh gsd-node-runner.sh was registered in MANAGED_HOOKS but shipped without a gsd-hook-version header, so gsd-check-update-worker.js always classified it as 'definitely stale' (a missing header is indistinguishable from a pre-version-tracking file). Every install on an otherwise up-to-date version showed a permanent, unclearable '⚠ stale hooks — run /gsd-update' warning naming this one file. Root cause: the build-hooks.js comment claimed the file is 'not a registered hook' and 'staged verbatim — no templating', but it IS in MANAGED_HOOKS (managed-hooks-registry.cjs:34) and install.js already stamps {{GSD_VERSION}} into every .sh hook unconditionally, gsd-node-runner.sh included. The comment contradicted both the registry and the installer's actual behavior, and the header line itself was simply never added. Fix: add the header (matching every other managed .sh hook's format) and correct the comment so it no longer asserts the opposite of what the registry and installer actually do. Adds a regression test that iterates every MANAGED_HOOKS entry and asserts it carries a header matching the worker's own detection regex, so a future hook added to the registry without one fails CI instead of shipping silently. Fixes #4076 * chore(#4076): add changeset fragment for PR #4092 * fix(#4076): address review nits — drop unneeded exemption, fix blank line Per @trek-e's review on #4092: - tests/managed-hooks.test.cjs:96: the readFileSync call uses a loop variable (entry-derived hookPath), not a literal path, so local/no-source-grep's static literal-path detector never flags it — the allow-test-rule exemption comment was unnecessary. Replaced with a plain note explaining the source-read rationale. - tests/managed-hooks.test.cjs:121-122: dropped a stray extra blank line before the bug #2136 section divider. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b0572c0108 |
feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage golden script) as a regression safety net ahead of extracting partitionStages into a standalone module. Also adds the new module's unit and property tests (test matrix rows 1-11) against its expected public API, which does not exist yet and is added in the next commit. * feat(#3674): extract file-overlap partitioner into a shared, generic module Moves partitionStages' greedy first-fit file-overlap algorithm into a new, dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap), generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own Plan[] shape onto the generic input and back — behavior-preserving, no dependency ordering, no path normalization, no filesystem access moved or added. Enables a future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive without pulling in orchestration internals. * docs(#3674): register the file-overlap-partitioner module bookkeeping New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row (regenerated via gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry matching the convention set by similarly-scoped leaf modules (text-lines.cts, plan-dependency-graph.cts, spec-section.cts). * fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct - docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write, produced no diff (manifest is keyed by content, not row order) - tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is: tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier. * fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction The `no two plans in the same stage share a modified file` property reconstructs which physical item produced each output id via `remaining.findIndex(r => r.id === id)`. Under duplicate ids (an explicitly-supported input shape for `partitionByFileOverlap`) that reconstruction can pick the wrong physical occurrence, producing a false-positive overlap failure (observed counterexample: p0(f1), p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but misread by id-order as [[p0,p208#1],...], which do overlap). Properties (a) determinism and (b) totality already exercise duplicate ids correctly and are left unchanged; only this property's generated items are now constrained to unique ids via `fc.uniqueArray`, where the reconstruction is unambiguous. --------- Co-authored-by: sim <sim@local> |
||
|
|
fa107c0461 |
fix(#4172): add --merge-async to test:coverage:scripts-floor (#4173)
* fix(#4172): add --merge-async to test:coverage:scripts-floor The "Coverage gate (merged shards)" test.yml job has OOM-crashed (exit 134, SIGABRT) on every push to next since |
||
|
|
5c7243e54b |
fix(#3995): derive the review diff base from the phase directory (#4181)
* test(#3995): diff base keys on the phase directory, not commit subjects All three derivation sites (Tier 3, spawn_reviewer, fallow pre-pass) must anchor on the phase directory's first commit; the milestone-blind repro (an archived milestone's same-numbered phase commit capturing the base) is the failing-first row. #3191/#3503 rows reworked to the directory-anchor contract; T6 docs-parity forbids any remaining phase-scope message-grep site. * fix(#3995): derive the review diff base from the phase directory A phase number is unique within a milestone, not a repository; the message grep had no milestone bound and tail -1 deliberately selected the oldest same-numbered subject, dragging archived milestones phases into the scope (7 files to 3388 plus the >50 depth downgrade). All three lockstep sites now anchor on the first commit that added anything under the phase own directory — the same anchor class git-base-branch phaseStartCommit uses. ShellCheck baseline gains the escaped fragment shifted parse signature. Emitted-Drift-Ack-Growth: code-review.md — phase-directory anchor replaces the message-grep derivation at both sites (#3995) * chore(#3995): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |