23941a871cfdc1790d9740a39f4d0378dc10591a
5551 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
23941a871c |
Merge pull request #4075 from open-gsd/chore/backmerge-main-to-next-36cc1f1c
chore: back-merge main → next (
|
||
|
|
aebf948cc2 |
chore: back-merge main into next (36cc1f1c)
|
||
|
|
8df14f57de |
Merge pull request #4074 from open-gsd/chore/sync-next-version-1.12.0
chore: sync next package version to 1.12.0 |
||
|
|
4a705cf19f | chore: sync next package version to 1.12.0 | ||
|
|
36cc1f1cb2 |
Merge pull request #4073 from open-gsd/release/1.12.0
chore: merge release v1.12.0 to main |
||
|
|
ceed559fd4 | chore: promote CHANGELOG for v1.12.0 | ||
|
|
9fe51d4de9 | chore: finalize v1.12.0 | ||
|
|
4d70b4dc43 |
fix(#4068): add --merge-async to c8 coverage-merge invocations (#4069)
* test(#4068): failing-first regression guard for c8 --merge-async flag Adds a config-invariant test asserting test:coverage:unit and test:coverage:report pass --merge-async to c8, guarding against the coverage-merge OOM (release.yml finalize dry-run, exit 134/SIGABRT) regressing a third time (prior stopgap: #199). Also commits the sourced research memo backing the diagnosis. * test(#4068): rename regression test to avoid lint-test-file-count collision coverage-merge-async-flag.test.cjs's effective prefix (coverage-merge- async-flag) matched the unrelated gsd-core/bin/lib/coverage.cjs module's 2-file cap under lint-test-file-count.cjs's startsWith bucketing, failing FAIL_EXCEEDS_LIMIT (confirmed via the RED gsd-test run on 7b6e6ca9e). Renamed to c8-merge-async-flag.test.cjs -- no colliding prefix. * fix(#4068): add --merge-async to c8 coverage-merge invocations The `test:coverage:unit` script OOM-crashed (exit 134, SIGABRT) in the release.yml finalize dry-run of 1.12.0: all 1785 unit tests pass, then c8's report/merge phase crashes ~76s later against the 6144 MB heap ceiling. Root cause (verified against this repo's pinned c8@11.0.0 source, node_modules/c8/lib/report.js): the default sync merge path, Report._getMergedProcessCov(), loads every raw V8 coverage file for the whole run into memory as one array before merging. This is a recurrence of #199 (same OOM class at ~466 tests, "fixed" by raising the heap ceiling) -- the suite outgrew 6144 MB as it grew to 1785 tests, and test.yml's own coverage-gate job already independently hit and stopgap-fixed the identical class once (test.yml:614-620, 4096->8192 MB). c8 ships the upstream fix for exactly this: --merge-async (c8 v7.14.0, already inside the pinned c8@^11.0.0 range -- no dependency bump), which switches to Report._getMergedProcessCovAsync(), reading and merging one raw file at a time instead of loading them all at once. The merge arithmetic (mergeProcessCovs()) and --all's zero-coverage-file inclusion are unchanged between the two paths, so the coverage gate's accuracy is unaffected. Verified empirically (not just by source inspection): against a 223 MB synthetic raw-coverage corpus derived from this repo's actual 235 gsd-core/bin/lib/**/*.cjs files (same order of magnitude as the 358 MB figure documented in test.yml), c8's real Report class crashes with the same "Ineffective mark-compacts near heap limit" signature via the sync path at a 300 MB heap ceiling, while the async path completes at a flat ~40-50 MB peak down to a 100 MB ceiling. Full repro steps and numbers: .gsd/bug/fix-4068-coverage-merge-oom/10-diagnosis.md. Only `test:coverage:unit` and `test:coverage:report` are changed. `test:coverage` (line 150) is a non-CI dev-convenience script, never invoked as a command by any workflow. `test:coverage:unit:raw` (line 153) already runs --reporter none, which skips the merge/report phase this flag affects, entirely -- adding it there would be a no-op. Fixes #4068 * fix(#4068): add changeset fragment for the coverage-merge OOM fix The changeset fragment was generated locally (npm run changeset) but never committed -- isolated code review caught the gap (no .changeset/*.md on the branch, confirmed via git diff --name-status). * chore(#4068): backfill changeset PR number pr:0 -> pr:4069 --------- Co-authored-by: sim <sim@local> |
||
|
|
8793d307f6 |
fix(#4060): drive repo-baseline lint check in-process, not via a subprocess timeout race (#4065)
* test(#4060): failing-first regression for repo-baseline subtest timeout race The "repo baseline passes" subtest in lint-allow-test-rule-refs.test.cjs drives the script under test via a spawnSync subprocess with a fixed 30s timeout, which races the script's real wall-clock completion against unbounded CI-load contention -- it has already died at this race twice (#4060, and once before at a lower bound). Rewrites the subtest to call the script's `main` directly, in-process, removing the subprocess timeout race entirely. This commit only changes the test (main is not yet exported), so it fails first with `TypeError: scriptUnderTest.main is not a function`. * fix(#4060): export lint-allow-test-rule-refs main() for in-process drive The "repo baseline passes" subtest previously drove this script via a spawnSync subprocess with a fixed 30s timeout, racing the script's real completion time against unbounded CI-load contention -- it has now died at that race twice (#4060, and once before at a lower bound). A fixed timeout racing unbounded contention has no value that is both tight and safe, so raising it again would not fix the mechanism, only its odds. Parameterizes main() to accept an explicit argv (defaulting to process.argv.slice(2) only when omitted, so the CLI entrypoint is unaffected) and exports it, so the test can call it directly, in-process -- removing the subprocess and its spawnSync timeout kill race entirely for this one row. * fix(#4060): capture stderr too in the in-process repo-baseline subtest Code-review finding: the in-process rewrite captured only console.log, but main()'s real failure path throws a bare, messageless ExitError -- all diagnostic detail goes to process.stderr.write. The old subprocess-based assertion embedded both stdout and stderr in its failure message; this silently dropped that debuggability. Captures process.stderr.write the same way (restored in finally) and surfaces both streams in assertion failure messages and in a wrapped re-thrown error on an unexpected throw from main(). --------- Co-authored-by: sim <sim@local> |
||
|
|
6e681cb36c |
fix(#3901): commit-docs-guard suites isolate children from a global core.hooksPath (#4063)
* test(#3901): the guard suite must survive a developer's global core.hooksPath (failing first) * fix(#3901): the commit-docs-guard suite isolates children from the host git config A developer's GLOBAL core.hooksPath (~/.gitconfig) applies to every fresh repo, so the guard's correct refusal — it will not install a hook git would never run — failed this suite's beforeEach and 18 tests read as a commit-docs-guard regression on any machine that centralizes commit hooks. CI never saw it (runners have no global git config). before() pins GIT_CONFIG_GLOBAL to an empty file for the whole suite (runGsdTools and the git helpers propagate process.env to every child). A LOCAL repo value cannot isolate this: any non-empty local value trips the same refusal, and an empty local value makes rev-parse --git-path hooks resolve to ./ instead of .git/hooks. The regression test pins the seam: the isolation file is armed and empty, a child git resolves no hooksPath from the host, and the beforeEach enable installed the hook at the repo-local default path. A child given an explicitly hostile GIT_CONFIG_GLOBAL still refuses — that is the guard being correct. * fix(#3901): review fold-ins — isolation hoisted to all three guard suites; changeset Adversarial review: the pin covered suite A only — the B (#3588 B1-B15) and D (real git commit wiring) suites run enable against fresh repos too and failed identically on hostile-global machines; the issue's 18 failures could not have come from A's five tests alone. The pin is extracted to isolateGlobalGitConfig() and armed in all three suites (B9's LOCAL hooksPath test is unaffected — the pin neutralizes only the global scope; D's premise is the hook firing from .git/hooks, which the isolation preserves). * chore(#3901): backfill changeset PR number (4063) * fix(#3901): drop the now-unused before import — max-warnings 0 fails CI --------- Co-authored-by: sim <sim@local> |
||
|
|
73919526e9 |
fix(#3902): packaging guard resolves npm 12's pack --json shape — stays armed on Node 26 (#4064)
* test(#3902): packListToPathSet must resolve npm 12's object-keyed pack --json shape (failing first) * fix(#3902): resolve both npm pack --json shapes — the packaging guard stays armed on Node 26 npm 12 (bundled with Node 26) emits an object keyed by package name where npm <=11 emitted an array; parsed[0] was undefined, the before() hook threw, and all 6 tests in the packaging guard — including the two 'does NOT ship' guards that stop repo-only CI tooling leaking into the published package — went dark for contributors on Node 26. CONTRIBUTING pins Node 24 as the floor but requires Node 26 compatibility for code AND tests; this was the test's parse, not the code. packListToPathSet now resolves either shape and fails loud on anything else — a silently-empty set would be the exact dark-guard failure mode this guard exists to prevent. * fix(#3902): review fold-in — a multi-key object fails loud Single-package assumption documented and enforced: npm 12 emits exactly one key for this repo (no workspaces — verified); if workspaces were ever adopted, first-key selection would silently validate one package's file list and recreate the exact silent-disarm this fix kills. * chore(#3902): changeset fragment (pr number backfilled after PR creation) * chore(#3902): backfill changeset PR number (4064) --------- Co-authored-by: sim <sim@local> |
||
|
|
686c05a740 |
fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128 (#4059)
* fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128 64 had no principled derivation (no git trailer limit, no CI resource bound) and a wide-touching maintenance PR can legitimately accumulate more than 64 distinct emitted-drift acknowledgments, tripping the cap and failing Required tests even though nothing is actually wrong. The cap's underlying purpose is unchanged: parseAckTrailers() still throws rather than truncates on overflow, and still forces pruning of stale acknowledgments at some ceiling. Only the ceiling moves. * chore(#4058): backfill changeset PR number (4059) --------- Co-authored-by: sim <sim@local> |
||
|
|
90fff40f5a |
fix(#3900): overlay walker skips non-regular files — a repo-root socket no longer fails 32 tests (#4062)
* test(#3900): a unix socket under the repo root must not break the overlay (failing first) * fix(#3900): the overlay walker skips non-regular files (sockets, FIFOs, device nodes) The walker classified every not-a-directory entry as a file, so a unix socket anywhere under the repo root (an MCP server, a language-server daemon, a watcher) made copyFileSync throw ENXIO — the reporter measured 32 local test failures across 6 files, all in before() hooks, none about the code under test; a FIFO would block until a writer appeared. The hard-link path does not save it: the overlay sits under tmpfs, so cross-device EXDEV always falls through to the copy. One guard at the classification site (srcStat.isFile() before the copy/link branches) — sockets, FIFOs, and device nodes are not repository content; CI never saw this because a fresh clone has no working-tree daemons, which is exactly why it ate local debugging time. Also adds opts.root (test-only root override, default REPO_ROOT) so the regression tests build the overlay from a tiny synthetic root in milliseconds instead of paying a full repo walk. * chore(#3900): changeset fragment (pr number backfilled after PR creation) * fix(#3900): review fold-ins — sibling cold-fixture guard, changeset format, root doc buildColdInstallTree enumerated the live repo's hooks/ with a NAME-only filter — the same class #3900 documents: a daemon socket or FIFO inside hooks/ would reach cpSync and throw ENXIO (or block). Dirent type guard added. The changeset gains the mandated bold user-visible lead, and opts.root is documented at its definition (mirroring cold-runtime-lib- fixture's repoRoot convention). * chore(#3900): backfill changeset PR number (4062) --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
472f585f7c |
fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates `milestone complete <version>` is a one-way door — ROADMAP.md and REQUIREMENTS.md archived, every phase directory in the milestone MOVED, STATE.md rewritten — and ran unconditionally on first invocation through every invocation path, including `query milestone.complete <version>`, whose `query` meta-prefix reads as a read-only namespace but performs no filtering (#167's invocation-compatibility shim + #3243's dotted-form normalization). The gate lives on the destructive command itself, not on the `query` prefix (the prefix is an intentional invocation mechanism, not a permission boundary — restricting it would break dozens of shipped workflow callers). Without --confirm and without --dry-run the command now refuses via error() before reading anything beyond its arg checks, so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run still previews with no confirmation needed and is now documented in the usage block (it was only documented for the sibling archive-quick). --force keeps its narrow meaning — bypassing the TRUNCATED-scope and unstarted-phase guards — and does not double as the mutation opt-in. --confirm follows the existing `phases clear --confirm` idiom in the same module. complete-milestone.md's two invocations pass --confirm (the workflow has gathered explicit user intent by that step). Existing tests get --confirm appended — pre-change behavior is exactly confirmed behavior — and a #3726 regression block covers: refusal + full-tree byte-identity on both invocation forms, --force not satisfying the gate, --dry-run still passing without confirmation, and --confirm proceeding. The refusal tests fail against pre-fix code (negative control run). Fixes #3726 * docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS Cross-AI review of the fix diff (codex, pre-create) caught three shipped doc sites still instructing the now-refused bare invocation: the CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's two guard-override instructions (`--force` alone now refuses without --confirm). Localized CLI-TOOLS copies already lag the English synopsis (no --force/--dry-run either) and follow the translation pipeline, not this fix. * chore(#3726): set changeset fragment pr to 3774 * test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack Two CI reds from the --confirm gate, both this branch's own misses: - tests/qa/scenarios/milestone-rollover.json invoked `milestone complete 1.0 --force` as a JSON arg-array fixture — a caller shape the test sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never enumerated. Adds --confirm; the scenario's boundary-crossing contract is otherwise untouched. - complete-milestone.md's +420-byte --confirm note trips the emitted-attribution growth ratchet. Acknowledged as a #3726 append to the existing complete-milestone.md entry in 3409-unreachable-guard-arms.json (two ack sources may never name the same path, per that fragment's own precedent). Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed. * docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm Review Major 1: the truncated-window and unstarted-phase guard paragraphs still told the reader to "Pass `--force` to override", which now refuses (--force alone does not satisfy the confirmation gate), while the flag table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair so the file no longer contradicts itself. * docs(#3726): synopsis renders --confirm and --dry-run as alternatives Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as "a dry run still needs --confirm", the opposite of AC 3. Render the pair as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage docblock, and let the flag rows carry the rule. * test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack Rebase onto next (26 commits) surfaced three tests the gate now refuses: the #3685 write-flag contract pair in tests/milestone.test.cjs and the `milestone complete` boundary fixture in tests/state-contract.test.cjs all invoke the command bare. Each now passes --confirm (a mutating run is exactly what they assert on). The +420 byte complete-milestone.md growth ack rode on 3409-unreachable-guard-arms.json, which #3078 swept from next as fully spent — hence the modify/delete conflict. Re-filed under a fresh fragment named for this issue, never resurrecting the swept one. * test(#3726): pin the present-but-falsy arm of the confirmation gate Review Minor 1: the boundary triple covered absent and present but not present-but-falsy. The gate is an exact-token match, so --confirm=false and --confirm=0 refuse today — pinned (canonical + query forms, whole .planning/ tree byte-identical) so a future `=`-aware or prefix-matching parser cannot silently turn --confirm=false into a confirmed run of an irreversible command. * test(#3726): drop --confirm from dry-run-only invocations Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run invocations that never needed it, so each stopped standing as incidental proof that a preview needs no confirmation. Reverted to the pre-PR form; the dedicated AC-3 test carries the explicit assertion. * docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate REQ-I18N-02 (docs/features/internationalized-documentation.md) requires translations to stay synchronized with the English source. The four localized CLI-TOOLS.md guides still advertised a bare `milestone complete <version>`, which now exits 1. Render the English synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and `[--archive-quick]` flags the translations had also fallen behind on. * test(#3726): drop --confirm from the remaining preview-only invocations Round 2 reverted the --confirm appends on --dry-run-only invocations in tests/milestone.test.cjs, but four more sat in two files the sweep missed: tests/milestone-archive.test.cjs (three) and tests/milestone-window-single-owner.test.cjs (one). Each is a preview run whose whole purpose is to document that a preview mutates nothing, so `--dry-run ... --confirm` contradicted the semantics the test exists to pin. Dropping the token restores each as incidental proof that a preview needs no confirmation; the dedicated AC-3 test keeps the explicit assertion. No assertion added, relaxed, or removed — the change is four tokens. * chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer #3954 (ADR-3942) moved emitted-drift acknowledgments out of tests/emitted-drift-acks/ and into git commit trailers, and the fragment directory no longer exists on next. The reason this PR's fragment carried moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit; the fragment file is removed rather than resurrected. Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift. * fix(#3726): name --confirm in the version-required refusal The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke the command without args and the error lists what is required) stopped at `version required for milestone complete (e.g., v1.0)` — one required argument short. Discovering --confirm took a second round trip through the gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test that also asserts the version-less invocation leaves .planning/ untouched. * test(#3726): pin the milestone complete docs against a silent regression The changeset is `type: Fixed`, which the docs-required lint exempts, so nothing in CI would notice a later edit that reinstated the bare-`--force` override prose or dropped `--confirm` from the synopsis. Four tests in tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md and its four localized mirrors; the `--confirm` flag row; both guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md, by guard name (a substring match on each instruction's `--force --confirm` text); and — as an identity ratchet over the milestone-complete sections — every `--force` sentence or clause that lacks `--confirm`, so a new bare instruction in its own sentence or clause fails whatever its wording. Named residual: a bare instruction spliced into the same clause as a compliant one coalesces with it and passes the ratchet; the by-name pins are what keep the four known instructions from losing the pairing that way. The file is registered in scripts/docs-guard-registry.cjs so the pin runs on the PR that changes those docs, not only after merge. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
370cfc6680 |
enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local> |
||
|
|
b431ae9f0d |
fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry (#4057)
* test(#3898): a spaced-hyphen thematic break in ## Gaps is not an entry (failing first) * fix(#3898): a separator-shaped line in ## Gaps is skipped, not made an entry splitGapsEntriesCore's opener regex (/^(\s*)-\s/) matched a spaced hyphen thematic break, fabricating a gap named '- -' with result 'unknown' — surfaced by audit-uat as an outstanding finding that cannot be cleared by editing any entry, because there is no entry, only the separator the author wrote for readability. Five shapes were affected (- - -, -, and wider/indented variants); the unaffected ones (---, ----, * * *, ___) were safe only by accident — the path never matched them, not because it understood breaks. A line whose content after the opening marker is solely hyphens and whitespace (with at least one further hyphen) is now skipped entirely — neither an opener nor a continuation. Deliberately option 2 from the issue, not a full thematic-break concept: a break does not close the Gaps list, entries after it keep parsing, and the deliberately-frozen byte-for-byte Gaps behavior changes ONLY for documents carrying such a separator (which previously produced a phantom). A real entry whose truth begins with a hyphen (- truth: "-5 error budget...") is untouched — its remainder contains non-hyphen characters. * fix(#3898): review fold-ins — span-contiguous skip, property coverage The skip is narrowed to where the phantom came from: a separator-shaped line BETWEEN entries (nothing open, or it would open a top-level entry). One landing strictly inside a live entry (indent > baseIndent) folds back as a continuation line, so entry lines and the GapsEntrySpan agree byte-for-byte — the span invariant and the #3805 ack writer's identity re-verification both hold (the review traced the unconditional skip to a match_verification_failed refusal in that corner). Adds the parser- convention property test (arbitrary hyphen counts/indents/spacings) and a span-contiguity pin. * chore(#3898): changeset fragment (pr number backfilled after PR creation) * chore(#3898): backfill changeset PR number (4057) --------- Co-authored-by: sim <sim@local> |
||
|
|
44ddc6dc46 |
fix(#3865): init.* phase queries accept --phase <N> as the positional alias (#4054)
* test(#3865): init.* phase queries accept --phase as positional alias (failing first) * fix(#3865): init.* phase queries accept --phase <N> as the positional alias The phase-taking init.* queries read their phase token at args[2] blindly: '--phase 60' made the literal '--phase' the phase, which the locator can never resolve — the reported incident answered a well-formed phase_found:false, plan_count:0 with exit 0 for a phase holding seven committed plans (ADR-3473 §8.4's later strict validation turned the paired form into a usage error instead, still not the alias). normalizePhaseAlias (shared by all eight phase-taking handlers — the issue's four plus phase-op, review, discuss-phase-assumptions, todos, the same class) rewrites '--phase N'/'--phase=N' into the caller-owned positional slot before flag parsing, so handlers and strict validation see exactly the argv the positional form produces. A valueless --phase is a usage error naming the flag. Any other flag-shaped args[2] now resolves to undefined (the commands' designed use-the-current-phase input) instead of passing flag text down as a phase name. isFlagToken exported from command-arg-projection (single owner of the flag-shape predicate). * fix(#3865): review fold-ins — honest no-position-given comment, fail-closed throws The helper's comment claimed flag-shaped args[2] resolves to 'the commands' designed use-the-current-phase input' — no such cross-module behavior exists: execute-phase/plan-phase/verify-work usage-error 'phase required', the find-based queries answer phase_found:false, and todos drops its area filter. Reworded to what actually happens, so the contract-grade comment cannot mislead a future edit. Adds the parseNamedArgsOrExit-style throw after each error() call as a fail-closed backstop against a returning fail(). * chore(#3865): changeset fragment (pr number backfilled after PR creation) * chore(#3865): backfill changeset PR number (4054) --------- Co-authored-by: sim <sim@local> |
||
|
|
d9e906a744 |
fix(#3864): smart-entry classify() matches the verif* status stem (#4052)
* test(#3864): smart-entry classify must match the verif* status stem (failing first) * fix(#3864): classify() matches the verif* status stem, aligning with normalizeStateStatus "verified"/"verification" contain no "verify" substring, so the exact-word branch never matched: a STATE.md declaring status: verified fell through to situation "unknown" (or idle-stranded on a clean tree with unpushed commits — differently wrong, which made it look intermittent). state-document's normalizeStateStatus already matches any verif* stem; the classifier now uses the same stem, so the two owners agree on the invariant. verify_failed is tested earlier in classify(), so failed verification still wins. Negative controls (completed/executing/planning/paused) unchanged and pinned. * test(#3864): review fold-in — the real handler-written verification status pinned both ways 'Phase complete — ready for verification' (state-document.cts's own Status default) carries the verif stem but names completion: at 5/5 it must stay complete (isComplete beats the stem), mid-project it must route to verify-pending (pre-fix it fell through to unknown/idle- stranded — a real behavior change beyond the literal verified repro, sanctioned by the issue's own verification→verify-pending table). * chore(#3864): changeset fragment (pr number backfilled after PR creation) * chore(#3864): backfill changeset PR number (4052) --------- Co-authored-by: sim <sim@local> |
||
|
|
9932929e37 |
docs(#4027): record crush-runtime and codex-native-plugin as out-of-scope (#4045)
Records two triage dispositions from the 2026-08-29 /triage-review sweep: new runtimes and first-party add-ons are not accepted in-tree; both route to the EoS Registry or Capability system as community-maintained work. Closes #4027 Closes #4033 Co-authored-by: sim <sim@local> |
||
|
|
4048ba2e80 |
fix(#3860): Quick Tasks lookup accepts milestone-suffixed headings; schema-aware section selection (#4050)
* test(#3860): quick-tasks heading tolerance for milestone suffixes (failing first) * fix(#3860): prefix-match Quick Tasks heading + schema-aware section selection The exact-anchored predicate (^quick tasks completed$) never matched a milestone-suffixed heading ('Quick Tasks Completed (v1.1+)'), so every /gsd-fast append AND every reset failed with QUICK_TASKS_SECTION_ABSENT before the columns were ever looked at — the reported table was already canonical. A heading is a section label, not data: the predicate is now prefix-anchored with a word boundary (Completedness still excluded), hoisted to one shared isQuickTasksHeading beside QUICK_TASKS_SECTION_ABSENT so the two call sites cannot drift. Section selection is now schema-aware (the issue's deliberate-choice ask): among ALL matching sections, the first whose table parses with a recognized Quick Tasks schema wins — a legacy (v1.0) table first in document order no longer shadows a usable (v1.1+) one below it. When no section is usable, the FIRST is returned so the error names the real problem (unrecognized schema) instead of a false 'no section'. Fixes one over-assertion in the new tests (v1.1+'s own next ordinal IS 2; pins v1.0 rows byte-identical instead). * fix(#3860): review fold-ins — level-bounded bodies, splice guard, convention pin Adversarial review caught that collectSections ends a candidate's body only at the next MATCHING heading: an intervening '## Deferred Items' table (canonical STATE.md layout — templates/state.md puts it after the Quick Tasks section) would be swallowed into the body, and the append's last-table-line scan would splice the quick-task row into that WRONG table — silent corruption on the primary /gsd-fast path. Each candidate is now re-collected through collectSection with an offset-precise predicate, restoring the pre-#3860 level-bounded stop (next same-or- higher heading) for both probing and splicing. Also pins the both-schema-valid tie-break (document order, newest-on-top — the layout the issue itself demonstrates; this layer has no active-milestone signal) with a test, guards the Deferred Items layout with a dedicated splice-target test, and updates resetQuickTaskRows's doc comment to name the new pipeline. * chore(#3860): changeset fragment (pr number backfilled after PR creation) * chore(#3860): backfill changeset PR number (4050) --------- Co-authored-by: sim <sim@local> |
||
|
|
fd63889d1f |
fix(#3854): write normalization preserves tight multi-line lists (#4049)
* test(#3854): write normalization must preserve tight multi-line lists (failing first) * fix(#3854): no blank before a bullet whose previous line is an indented continuation _normalizeMd's 'separate a list from a preceding paragraph' rule inserted a blank before any bullet whose previous line wasn't a bullet — but an INDENTED CONTINUATION of the previous multi-line item also isn't a bullet. Every .md write (phase.complete in the report, but any write through platformWriteSync) therefore converted tight lists to loose ones: +61 blank lines on the reporter's 1015-line ROADMAP, one before each bullet following a wrapped item. Tight and loose lists render differently, so this was a rendering change plus misleading diff noise; one-shot (idempotent afterwards), which is why integrity checks on headings/content passed. The guard is the mirror image of the after-a-bullet rule two lines below, which already excludes indented next lines. Paragraph→list and heading→list separations — the rule's purpose — are pinned unchanged by the new suite. * fix(#3854): review fold-ins — ceiling tracks next's 281 + this branch's marker (282), header/require nits The ceiling is not ratcheted but must track the tree: origin/next raised it to 281 (sibling branch's marker file); this tree adds one more (shell-command-projection-md-normalize), so 282/282. Also fixes the test header's stale pre-rename filename and hoists the inline require to the file's single import. * chore(#3854): changeset fragment (pr number backfilled after PR creation) * chore(#3854): backfill changeset PR number (4049) --------- Co-authored-by: sim <sim@local> |
||
|
|
529480b4a5 |
fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model (#4048)
* test(#3895): no shipped agent may hardcode a model frontmatter pin (failing first) * fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model Exactly one of the 34 shipped agents carried 'model: sonnet' in its frontmatter; every other agent resolves through the model-profile system. The ship:post dispatch (#2684) resolves per-hook and — per #2517 — deliberately OMITS model= on inherit so the agent inherits the orchestrator's model; the frontmatter pin intercepted that inherit case, silently forcing sonnet where all 33 siblings would inherit, and operators could not durably remove it (install rewrites live copies wholesale). Deleting the line changes nothing for default profiles — the catalog entry (model-catalog.json agents.gsd-mempalace-curator: golden/balanced sonnet, budget haiku) preserves today's behavior — while restoring model_overrides and inherit authority. Pinned by a new agent-frontmatter guard: no shipped agent may hardcode a model pin, and the catalog entry must keep existing so the pin's deletion can never orphan the agent. * chore(#3895): changeset fragment (pr number backfilled after PR creation) * chore(#3895): backfill changeset PR number (4048) --------- Co-authored-by: sim <sim@local> |
||
|
|
400db94e02 |
fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults (#4047)
* test(#3894): research_before_questions must resolve globally and order quick.md (failing first) * fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults Two layers, one key. The quick workflow ran its discussion phase before its research phase unconditionally — neither quick.md nor its steps ever read workflow.research_before_questions, though the key is documented, schema-registered, /gsd-settings-writable, and honored by /gsd-discuss-phase and /gsd-new-project. A gray-area answer given without research is then written to <quick_id>-CONTEXT.md as a locked decision downstream agents are told not to revisit — an evidence-free choice made unfalsifiable (the reporter's #3714 misresolution). - quick.md Step 4 now carries the same research-before-questions check the two honoring paths make: when enabled, research-phase executes before discussion-phase; false/unset keeps the written order. Both sections stay section-manifest gated. - src/config-loader.cts forwarded workflow.post_planning_gaps from ~/.gsd/defaults.json but silently dropped this key — same file, same nesting, one resolved and one didn't. Now forwarded with the same flat + nested-alias fallback shape, added to the resolution-keys lockstep canary and the #3532 shadowed-warning set (nested alias reporting generalized over both keys). Emitted-Drift-Ack-Growth: quick.md — #3894: +Step 4 ordering rule (the research-before-questions check the discuss-phase and new-project paths already make); a real behavioral gate, not incidental bloat. * fix(#3894): review fold-ins — gate the CONTEXT.md reference, colon slash-forms - quick/steps/research-phase.md directed the researcher subagent to read <quick_id>-CONTEXT.md under DISCUSS_MODE with no existence hedge — but under the new ordering (research BEFORE discussion) the file cannot exist yet when the researcher is dispatched. The reference now says read-only-if-present with the #3894 reason; the alignment purpose still applies on the default ordering. - quick.md's new rule used the hyphen slash forms (/gsd-discuss-phase, /gsd-new-project); source artifacts under gsd-core/workflows must author the colon form the install-time converters key on — the same file already uses /gsd:new-project and /gsd:quick elsewhere. * docs(#3894): planning-config row names the flat CONFIG_DEFAULTS alias config-field-docs requires every CONFIG_DEFAULTS key to appear in the doc; the row documented the canonical namespaced form only. Adds the same alias sentence post_planning_gaps's row carries, plus the #3894 quick-path note. * chore(#3894): changeset fragment (pr number backfilled after PR creation) * chore(#3894): backfill changeset PR number (4047) --------- Co-authored-by: sim <sim@local> |
||
|
|
192eb1dfbd |
fix(#3886): git commit timeout reported as commit_timeout; stale lock surfaced; 30s band (#4046)
* test(#3886): a timed-out git commit reports commit_timeout, not commit_failed (failing first) * fix(#3886): git commit timeout reported as commit_timeout; 30s band; stale-lock surfaced cmdCommit's git commit invocation did not distinguish a spawnSync timeout from a real non-zero exit (#2608 fixed this for the staging loop only): a slow pre-commit hook crossing the 10s cap was SIGTERM'd mid-hook and reported as reason commit_failed with whatever partial stderr git had flushed (in the reporter's case an incidental CRLF warning), while the kill left a stale .git/index.lock blocking the next attempt. All three commit sites now check isSpawnTimeout before the nothing-to-commit/ordinary-failure branches: cmdCommit reports reason commit_timeout + timed_out:true and names the stale lock's path (surfaced, not auto-deleted — deleting a lock a live git holds is destructive; the caller recovers deliberately); the subrepo counterparts do the same within their per-repo result / rollback error. The commit calls also move to the 30s band the push call already uses — husky+lint-staged alone idles ~4s on Windows before any task runs. * fix(#3886): review fold-ins — git-path lock resolution, shared band constant, executor contract row, precedence pin - The stale-lock path is resolved via git rev-parse --git-path index.lock, never a literal .git/index.lock join (#3588 row 8's class: a linked worktree's .git is a FILE, so the literal path cannot exist there while the real lock — under <gitdir>/worktrees/<name>/ — blocks the next commit; this repo leans on linked worktrees). - COMMIT_TIMEOUT_MS hoisted; all three sites and their messages build from it (the subrepo variant also regains the stdout fallback the primary site had). - agents/gsd-executor.md's commit-result contract gains the commit_timeout row with the OPPOSITE retry advice from staging_timeout (remove the stale lock, then retry once) — an executor matching the doc previously had no handling for the new reason. - Precedence pin: a timeout whose partial output contains 'nothing to commit' must still read as a timeout (branch-reorder mutant). Emitted-Drift-Ack-Growth: gsd-executor.md — #3886: +commit_timeout row to the commit-result contract with the retry guidance OPPOSITE staging_timeout's (remove the stale lock, then retry once); the executor previously had no handling for the new reason. * chore(#3886): changeset fragment (pr number backfilled after PR creation) * chore(#3886): backfill changeset PR number (4046) --------- Co-authored-by: sim <sim@local> |
||
|
|
0bf778c352 |
fix(#3849): phase allocation counts numbers held by sibling git worktrees (#4042)
* test(#3849): phase allocation must skip numbers held by sibling worktrees (failing first) * fix(#3849): phase allocation counts numbers held by sibling git worktrees Both allocators (cmdPhaseAdd, cmdPhaseAddBatch) chose max+1 over numbers gathered from ONE checkout — headers, bullets (add only), on-disk dirs. Every sibling git worktree carries its own .planning/ on its own branch, so a phase minted there was invisible and the same number was allocated twice (the reported incident: two Phase 441s, one with six written plans, surfaced a day late by human memory). New shared horizon collectSiblingWorktreePhaseNums: one git worktree list --porcelain, then per sibling — phase-dir names (the cheap scan that would have caught the incident) and the WHOLE sibling ROADMAP.md headers (a row can predate its dir; milestone-scoping would be wrong — a number used under any milestone on another branch is taken). Widen, never refuse: unreadable sibling / no .planning / not a git repo / git unavailable each contribute nothing and allocation is unchanged. Reuses isSentinelPhaseId and the allocators' own patterns. Secondary (#1229 never reached batch): cmdPhaseAddBatch now also scans roadmap bullets — a bullet-only 'Phase N' row was invisible to batch allocation, exactly the condition #1229 was filed for. Also fixes the two new tests' result-key access (output.phases, not output.results). * fix(#3849): review fold-ins — subprocess band, bounded test git, linked-worktree fixture - execFileSync options now match the repo's git band (10s window, windowsHide, 4MiB maxBuffer) — a spurious 4s timeout silently reverted to the pre-fix collision. - test git helper bounded (15s) per local/no-unbounded-spawn. - new fixture: allocation FROM a linked worktree counts the main checkout — the incident's actual topology direction. - fixture-setup rmSync carries the sanctioned lint-disable (setup, not teardown; cleanup() still owns directory removal). * fix(#3849): exempt the sibling-worktree scan from the enumeration-drift guard collectSiblingWorktreePhaseNums reads a SIBLING checkout's phases dir — a different question from the cwd-scoped listMilestonePhaseDirs the guard routes everything to (which cannot see another worktree's .planning at all). Function-scoped, per ADR-3180 Decision 4(a): any other re-derivation in phase.cts is still caught. The GREEN bench caught the omission. * chore(#3849): changeset fragment (pr number backfilled after PR creation) * chore(#3849): backfill changeset PR number (4042) --------- Co-authored-by: sim <sim@local> |
||
|
|
519ac23ebb |
fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first) * docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts; its exit-2 block IS the contract — a post-tool hook cannot prevent a commit) and gsd-session-state.sh is registered SessionStart (session orientation, not post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the three INVENTORY locales; the issue asked for a neighbouring-row scan, which is how the session-state row was found. All other rows in the four tables verify against the surface. * fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins Adversarial review found the same two wrong rows shipped in five more files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now covers all ten shipped tables. The parity parser unioned only the Kimi mirror list, silently exempting agent-isolation-guard (registered via the dynamic preToolEvent push): probes are now parsed too, with bare hook names resolved against hooks/ ground truth and dynamic event variables resolved to their canonical (non-Gemini) events; an exact-set pin replaces the loose size guard. allow-test-rule marker carries the issue ref; unverified-ceiling 280→281 (audited: the new marker is legitimate — the suite reads product docs whose text is the contract). * fix(#3839): register the hook-table parity suite in the docs-guard lane The new suite reads ten docs/ paths, so lint-docs-guard-registration requires it in the docs-guard registry — the first GREEN bench run caught the omission (the RED run's docs-guard failures were the same signal, previously misread as marker fallout). * chore(#3839): changeset fragment (pr number backfilled after PR creation) * chore(#3839): backfill changeset PR number (4041) --------- Co-authored-by: sim <sim@local> |
||
|
|
3c08315a5e |
fix(#3827): new-mode routing gate — classify approval no longer authorizes scaffold writes (#4037)
* test(#3827): new-mode routing approval gate before roadmapper (failing first) * fix(#3827): new-mode routing gate — classify approval no longer authorizes scaffold writes The route_new_mode step delegated to gsd-roadmapper (PROJECT.md, REQUIREMENTS.md, ROADMAP.md, STATE.md creation + commit) with no gate: the discovery gate approved classification only, and the zero-conflict branch said 'proceed to routing silently'. Merge mode previews its diff and gates via approve-revise-abort; new mode now shows the exact destinations and requires Create planning setup | Keep synthesized intel only | Abort, per the skill contract's routing-gate requirement. The keep-intel-only choice is the analysis-only path the issue asks for: no roadmapper, no destination writes, intel preserved; finalize labels it 'new (intel only)' and points at /gsd:new-project instead of plan-phase. Ambiguous gate answers re-ask once, then treat as Abort — never infer Create. Zero-conflict wording is mode-aware: silence about conflicts is not authorization to write. Emitted-Drift-Ack-Growth: ingest-docs.md — #3827: +routing gate display block, AskUserQuestion contract, three disposition branches (incl. the intel-only no-write path), ambiguous-answer rule, intel-only finalize line, mode-aware zero-conflict pointer; a real behavioral gate, not incidental bloat. * chore(#3827): changeset fragment (pr number backfilled after PR creation) * chore(#3827): backfill changeset PR number (4037) --------- Co-authored-by: sim <sim@local> |
||
|
|
331747ea99 |
fix(#3817): count the truncation remainder — display truncates, counting must not (#4034)
* test(#3817): audit-open counts must include the truncation remainder * fix(#3817): count the truncation remainder — display truncates, counting must not * chore(#3817): changeset fragment (pr number backfilled after PR creation) * chore(#3817): backfill changeset PR number (4034) --------- Co-authored-by: sim <sim@local> |
||
|
|
ac0eed1267 | Merge pull request #4015 from open-gsd/fix/3889-instrument-chunk-timeout | ||
|
|
213a2fff63 |
chore(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path (#4029)
* test(#3813): pin the #1883 unreadable-milestones contract on the live planning-snapshot path * fix(#3813): delete the caller-less listMilestoneArchiveDirs seam; #1883 contract now pins the live path * chore(#3813): changeset fragment (pr number backfilled after PR creation) * chore(#3813): backfill changeset PR number (4029) * chore(#3813): docs-exempt marker — internal dead-code removal --------- Co-authored-by: sim <sim@local> |
||
|
|
b811ea16fc |
fix(#3807): advance-plan refuses an ambiguous multi-Phase Current Position (#4028)
* test(#3807): advance-plan must refuse an ambiguous multi-entry Current Position * fix(#3807): refuse an ambiguous multi-Phase Current Position before advancing * chore(#3807): changeset fragment (pr number backfilled after PR creation) * chore(#3807): backfill changeset PR number (4028) --------- Co-authored-by: sim <sim@local> |
||
|
|
3a4c3cb83e |
fix(#3805): audit-uat honours the audit_acknowledged marker via the shared predicate (#4025)
* test(#3805): audit-uat must honour the audit_acknowledged marker * fix(#3805): route audit-uat's UAT and VERIFICATION scans through the shared acknowledged predicate * chore(#3805): changeset fragment (pr number backfilled after PR creation) * chore(#3805): backfill changeset PR number (4025) --------- Co-authored-by: sim <sim@local> |
||
|
|
f4fefb0bef |
fix(#3804): audit-uat enumerates all three phase-archive layouts (#4022)
* test(#3804): audit-uat must see all three phase-archive layouts * fix(#3804): enumerate all three phase-archive layouts (flat, workstream-archived, workstream-active) * chore(#3804): changeset fragment (pr number backfilled after PR creation) * chore(#3804): backfill changeset PR number (4022) --------- Co-authored-by: sim <sim@local> |
||
|
|
c1a2c33886 |
perf(#4012): two subprocess spawns, not five
My instrumentation suite tipped ubuntu shard 1/3 over its 15-minute job cap.
Measured, not guessed: baseline on next (
|
||
|
|
80de48c319 |
enhance(#3914): every phase records a truthful guard ledger (#4018)
* fix(#3914): retire n/no-process-exit where its successor governs
Epic #3889 criterion 5 — no phase closes with a guard added and its
predecessor left standing — is violated in the tree by the epic that wrote it.
local/require-registered-exit was registered on gsd-core/bin/**/*.cjs and
scripts/**/*.cjs, while n/no-process-exit stayed 'error' over a nine-glob block
covering those same two. Only the hooks 'off' exemption ever came down; the
predecessor's registration never did. Both rules have been enforcing the same
property on the same surfaces since P6.
Narrowed, not deleted. Seven of those nine globs have NO successor —
eslint-rules/, bin/lib/, pi/, examples/, vscode/, .kilo/, .opencode/ — so
deleting the rule outright would silently drop enforcement on all seven. That
is the inversion this epic has already hit three times: removing a coarse guard
because a narrower one exists somewhere it does not reach. Flat config is
last-match-wins and both successor blocks come after the nine-glob block, so
'n/no-process-exit': 'off' in exactly those two retires the predecessor
precisely where the successor governs and nowhere else.
The successor is strictly more precise: it permits process.exit only inside
terminateNow in cli-exit.cts, the single sanctioned terminator (ADR-3889 §3),
where n/no-process-exit permits none and would flag terminateNow's own
generated copy.
Asserted at the consumer's altitude via ESLint.calculateConfigForFile on real
paths, with the positive control that matters: n/no-process-exit is still
'error' on six of the seven successor-less globs, so a future edit that turns
this into a blanket disable goes red. bin/lib/ has no file in this checkout and
is reported as untested rather than given an invented path. Severity is
normalized across the string/numeric/array forms the API can return, and the
normalized value asserted — not truthiness.
Verified by running calculateConfigForFile myself on both superseded globs and
four controls before trusting the test.
Found and fixed inline: the change made an eslint-disable directive at
gsd-tools.cjs:257 partially unused, which --max-warnings 0 rejects; narrowed to
the one rule still in force.
Verification runs on the remote runner.
Refs #3914
* docs(#3914): the epic added three guards, it did not remove one
The audit reconciled the epic ledger against what actually landed. The net is
+3, not -1: four lint:generated-sync --check arms (gen-scripts-cli-exit,
gen-hooks-cli-exit, gen-exit-code-registry, gen-exit-code-docs) plus one rule,
against two retirements.
An epic whose thesis was consolidation ended with a larger guard surface than
it started with. The additions are each defensible; the claim that the total
fell was never true.
Two of the three prior errors in this amendment are mine. It said "Net -1 by
count" above terms reading -1 -1 +1 +1 +1, which sums to +1 — an arithmetic
error in the paragraph directly below the sentence arguing that an ADR about
honest accounting must not pad its own ledger. And the term list omitted two of
the four --check arms, which is what turns that +1 into the real +3.
Recorded rather than quietly rewritten. This ledger has now been wrong three
times — the original -2, the -1 that replaced it, and #3914's own table, which
states -1 above terms summing to 0 — and a written claim nobody checked against
the thing it describes is the exact failure this epic exists to close.
Refs #3914
* fix(#3914): make the successor actually supersede before retiring the predecessor
An isolated security review found that the previous commit turned off a guard
that was still doing work. Reproduced by executing both rules against a
fixture, not inferred:
const exit = 'exit';
process[exit](1);
n/no-process-exit flags it; local/require-registered-exit did not, because it
early-returned on callee.computed. So retiring the predecessor on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs un-guarded that shape on precisely
the two globs this epic's exit contract cares most about.
This is the third time in this epic I have removed a coarse guard on the claim
that a narrower one covered it, without checking construct-level parity — after
the allowlist key-to-prefix-to-exact-membership sequence and the band
ranges-to-categories one. The rule is the same every time: a narrower guard
supersedes a coarser one only where it demonstrably reaches at least as far,
and "demonstrably" means executing both against the constructs, not reading
either.
The successor now resolves computed property access for the statically
determinable cases — a string Literal, and an Identifier bound once to a string
Literal, resolved through scope — and leaves genuinely dynamic properties
alone so the rule does not over-fire. Measured after the fix: plain
process.exit flagged, process['exit']() flagged, process[exit]() flagged,
process[globalThis.k]() not flagged. That makes it a strict superset of the
predecessor on these globs, since process['exit']() was caught by NEITHER rule
before.
The second finding is worse than the first, because it was reasoning rather
than oversight. My justification comment claimed n/no-process-exit "would flag
terminateNow's own generated copy here". It would not — that file is in the
global ignore list, so neither rule ever lints it. There was no conflict to
resolve; I wrote a rationale I had not checked, in a change whose entire
subject is written claims nobody verified. Both comment blocks now state the
real basis.
The tests that should have caught this asserted only rule SEVERITY per glob and
never construct REACH, which is exactly how a coverage hole passed. A parity
matrix now pins all five shapes, including a RED/GREEN regression pin against
an inlined reproduction of the pre-fix rule — inlined rather than loaded from
HEAD, because HEAD resolves to the fixed commit under the remote runner and
would silently stop testing anything.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the two exit rules are complementary — keep both
Reverts this branch's retirement of n/no-process-exit. The premise was wrong
twice, and the second review proved the change itself was wrong.
I claimed local/require-registered-exit was a strict superset on
gsd-core/bin/**/*.cjs and scripts/**/*.cjs. Measured, successor vs predecessor:
function f(exit) { process[exit](1); } 0 vs 1
let exit='exit'; exit='exit'; process[exit]() 0 vs 1
const { exit } = ...; process[exit](1) 0 vs 1
plus for-of bindings, let-then-assign, var redeclaration, catch params, and an
undeclared global named exit. The predecessor matches any identifier NAMED
exit however it is bound; the successor resolves only a string literal or a
single-write const. It never was a superset — I asserted the relationship after
fixing one construct and did not re-check the rest.
The justification was independently false: all three generated cli-exit copies
are in the global ignore list, so n/no-process-exit was never flagging
terminateNow. There was no conflict to resolve. I wrote a rationale I had not
verified, in the phase whose subject is written claims nobody checked.
So criterion 5 does not apply to this pair. They are not predecessor and
successor — they are complementary, each catching constructs the other misses.
The epic's criterion assumed a replacement relationship that does not exist
here, and retiring either rule loses real coverage. The ADR ledger now says so
with the measured shapes.
What survives is the genuine improvement: the computed-property strengthening.
local/require-registered-exit now catches process['exit'](1) and optional-chain
terminators like process?.[k]?.(1), which NEITHER rule caught before, while
correctly ignoring a genuinely dynamic property so it does not over-fire.
The parity tests are rewritten to assert what is true rather than what I wanted
to be true: a bidirectional matrix where each rule is shown catching shapes the
other misses. The previous matrix tested only the four shapes where the
successor wins, which is precisely why the regression shipped — a test set
selected to confirm the thesis.
Also corrected: a stale ADR sentence claiming a third wrong ledger version that
does not exist (the table it described now reads +3 over terms summing to +3),
and a changeset whose stated motivation was the false generated-copy conflict.
Verification runs on the remote runner.
Refs #3914
* fix(#3914): the exemption term was a no-op — the net is +4
Fourth correction to this ledger, and a fourth error of the same kind.
Every version counted removing the n/no-process-exit 'off' entry from the hooks
block as -1. Measured: calculateConfigForFile returns undefined for that rule on
hooks/**. It was never registered there, and no broader block sets it globally,
so the 'off' entry overrode nothing and removing it changed no enforcement at
all. A no-op removal, not a guard removal — the same category error as counting
baseline acknowledgement entries: a thing that is not a guard, in guard units.
It is misattributed too; that block came down in
|
||
|
|
51ca9f39ba |
fix(#3801): register inline_plan_threshold in the defaults manifest and correct the docs (#4019)
* fix(#3801): register inline_plan_threshold in the defaults manifest (default 2) and correct settings-advanced * chore(#3801): changeset fragment (pr number backfilled after PR creation) * chore(#3801): backfill changeset PR number (4019) * test(#3801): parse the defaults table with the shared markdown-table parser --------- Co-authored-by: sim <sim@local> |
||
|
|
ac7587287b |
fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true #3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue was reopened rather than noted. The section now states three things: every field is single-valued, the section is overwritten rather than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line appended in good faith is silently ignored rather than winning. Progress history belongs in ## Performance Metrics, two headings down, and the text now points there. The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before being written rather than inferred from the issue title: stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase") -> "1 of 5 (First)" The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags im with NO g, so String.match returns the first hit. Writing "first wins" without running it would have repeated the exact error I had to retract twice in this epic already. A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines and asserts it resolves independently - negative space, because a reader returning the first line of the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a last-match variant returns "9 of 9 (Appended later)" and T1 reds. No assertion checks that the document contains a sentence. That is what local/no-source-grep exists to stop, and it would pin wording that is allowed to improve. The point of the test is that if that regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly becoming a lie with nothing to notice. Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four locales. The constraint is structural, not stylistic - confirmed by running that comparison after the edit. The four locale copies are translated rather than left stale. They are not gate-enforced for prose, so "nothing fails" was available and is not the same as correct: leaving four documents asserting something the English one now contradicts is a correctness problem. Code spans and the anchor link stay untranslated - they name real tokens. The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been false the fix would have belonged in the generator, and a hand edit would have been silently reverted. One real gate failure fixed inline rather than reported: the new test's comments referenced docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and lint-docs-guard-registration failed lint:ci correctly. Registered. Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope and stays unclaimed. Closes #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): the rule I documented was false — replace it with the measured one An isolated review returned two blockers. Both mine, and the first is the worse kind: I wrote a falsifiable rule into a reference page and got it wrong. 1. "A duplicate resolves to the FIRST occurrence" is FALSE. stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence beats document order. Measured against the built reader, all intra-section: Phase: A (plain) / **Phase:** B (bold, later) -> B LATER WINS Phase: A (indented) / Phase: B (plain, later) -> B LATER WINS Phase: A (plain) / | Phase | T (table) | -> A first wins My original verification tested plain-versus-plain, saw first-wins, and generalized to all forms. Measuring one case and claiming the general rule is the same error I have had to retract twice already in this epic. It is also worse than silence. The sentence told authors an appended line is safely ignored; a bold line appended "for emphasis" silently overrides the original. Someone trusting the doc would have corrupted their own state file. And #3812 never asked for a resolution rule - it asked for single-valued, overwrite-not-append, and where history goes. The rule was my unrequested addition. Replaced with the measured truth: resolution is by FORM (bold anywhere, then plain at line-start, then table row), and only WITHIN the winning form does the first occurrence win. Both consequences stated plainly - a higher-ranked form wins regardless of position, and an indented `Phase:` is invisible to the plain form. All five claims in the new paragraph verified by execution before being written, including the two I had wrong. 2. The tests tested the wrong case and passed for the wrong reason. T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section case, which #2956 already fixed by scoping. #3812 says verbatim that #2956 "fixed the inter-section case and never addressed intra-section duplication", so the case the new prose describes was untested, and the fixtures passed because of section scoping rather than field resolution. They also called bare stateExtractField rather than the production chain, T2 could not discriminate first from last at all, and no fixture mixed forms - which is precisely why the false claim survived to review. Rewritten as four rows, all intra-section, all through the real stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first wins within a form), plain-then-bold (the bold LATER value wins - the row whose absence let the false claim ship), indented-then-plain (indented invisible), and sibling-field independence. Each proven to fail against a reader that disagrees. 3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the anchor their own heading generates. ja-JP/ko-KR kept the English heading, so theirs already resolved. 4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section designed to grow" with a bare demonstrative whose nearest referent read as Current Position - saying the opposite of the point. Rewritten so the clause attaches unambiguously to `## Performance Metrics`. 5. Cross-locale drift, flagged by the implementing agent rather than by me: after fixing EN, the four locales still stated the OLD false rule. Four documents asserting something measured to be wrong is worse than four saying nothing. All four now carry a faithful translation of the corrected paragraph, with code spans, each file's own anchor, and the ja/ko referent fix preserved. Verified: all five claims executed against the built reader; every rewritten test row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to date, so the edits stay outside the generated marker regions; locale heading parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): second false rule on the same page — scope the ranking to the section A second isolated review found a second false falsifiable claim, and the failure mode is the same one twice in a row: attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT Both times the claim covered a wider surface than what was actually executed. The fix each time was not a better sentence, it was executing the surface the sentence describes. BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false. ## Current Position / Phase: 1 of 5 + ## Archive / **Phase:** 88 bare stateExtractField(whole doc) -> "88 (other section)" PRODUCTION (slice then extract) -> "1 of 5 (in section)" #2956's section slice means production never hands another section to the matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not seen. The ranking is real but scoped: it applies WITHIN `## Current Position`. I verified against the bare function and wrote a claim about the system. Every existing test placed its bold line inside the section, which is exactly why nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive` and asserts production returns the in-section plain value, with the unscoped reader asserted to DISAGREE so the row proves the scoping rather than assuming it. BLOCKER — the changeset still shipped the ORIGINAL retracted claim. I corrected the page and left the release note saying "resolves to the first occurrence ... a second entry added in good faith is silently ignored". The note contradicted the page it announces, and the release note is what most people actually read. Rewritten to the corrected rule. MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the file", which is the vacuous direction; the surprising case, and the one the very next clause illustrates with an APPENDED bold line, is "even if it comes LAST". All four locales reproduced the inversion faithfully, so it was an EN-source defect rather than translation drift. Two sharp edges now named, both measured: a bold `**Phase:**` followed only by trailing spaces resolves to an EMPTY STRING and does not fall through to a valid plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form ranking has to say where the ranking bites. Also fixed: all five files labelled the link `## Performance Metrics` while the heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only the label's level was wrong. Every clause in the final paragraph re-verified through the PRODUCTION chain (stateCurrentPositionSlice -> stateExtractField), clause by clause, before being written: bold in another section does not win; bold in frontmatter does not win; bold appended last does win; first wins within one form; trailing-space bold yields empty. All four locales carry the same corrected rule. gen-state-md-docs --check reports all 6 targets up to date; heading counts unchanged at 20/20 across all five files, so locale heading-parity is untouched; build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3812): backfill changeset pr number Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1051c6d8d4 |
fix(#3799): scope legacy cleanup to the install's resolved config dir and add --no-legacy-cleanup (#4013)
* fix(#3795): read the interrupted agent id before clearing the stale marker (#4006) * test(#3795): the interrupted-agent read must precede the stale-id clear * fix(#3795): read the interrupted agent id before clearing the stale marker execute-plan's init_agent_tracking step ran `rm -f .planning/current-agent-id.txt` BEFORE the existence check that read it, so the interrupted-agent branch and the Task resume prompt it exists to offer were unreachable (#3795) — a kill -9 mid-executor left the file, and the next run deleted it before looking. The read now precedes the clear; fresh-run semantics (no stale id leaking into the new spawn) are preserved. A structural guard pins the order. Emitted-Drift-Ack-Growth: execute-plan.md — #3795: +bytes from reordering the interrupted-agent read before the rm plus the explaining comment * chore(#3795): changeset fragment (pr number backfilled after PR creation) * chore(#3795): backfill changeset PR number (4006) --------- Co-authored-by: sim <sim@local> * fix(#3799): scope legacy cleanup to the install's resolved config dir and add --no-legacy-cleanup * chore(#3799): changeset fragment (pr number backfilled after PR creation) * chore(#3799): backfill changeset PR number (4013) * test(#3799): mark the legacy-name fixtures with gsd-allow-legacy-name --------- Co-authored-by: sim <sim@local> |
||
|
|
d1d6c82a0f | chore(#4012): backfill changeset pr number to 4015 | ||
|
|
bf8905fcc9 |
fix(#4012): a hang never emits the events I was listening for
The init marker settled it. The diagnostic reported "THE REPORTER LOADED BUT RECORDED NO TEST EVENTS — the events file contains only the reporter's own reporter:init marker", which refutes the reporter-never-loaded hypothesis and leaves exactly one explanation. The runner spawns a child process per test file and surfaces a subtest's test:start / test:pass / test:fail to the parent's reporter only once the child REPORTS that test — which happens when it completes. The fixture hangs forever, so it never completes, so it never reports. I had recorded exactly those three event types: the precise set that a hang guarantees you will never see. The feature could not have worked for the case it was built for. test:enqueue and test:dequeue are emitted by the runner as it queues and begins each file, independent of anything inside finishing. test:dequeue is what actually means "in flight", and it is now the primary signal, with test:start kept as a secondary one. A file is in flight when it has been dequeued and has no terminal event. The four branches now describe states that are all real: the events file absent (reporter never loaded); the init marker alone (the runner dequeued nothing at all — genuinely surprising now rather than the expected outcome); everything dequeued and terminated (the files finished and the process hung afterwards, a handle leak); and one or more dequeued-but-unterminated files, named, which is the case this whole feature exists to report. Verified against the exact shape the real hang produces, by executing analyzeChunkEvents on a synthetic events file: init + enqueue + dequeue with no terminal event reports hangs.test.cjs as in flight, and appending a test:pass clears it. Four more unit tests cover the ordering and multi-file cases with no subprocess, so this logic is now checkable without a runner round-trip — which matters, because every defect in this feature so far was visible only remotely. T1 is untouched and should now pass for the right reason. Verification runs on the remote runner. Refs #4012 |
||
|
|
444137e63a |
fix(#4012): make the artifact say whether the reporter ever loaded
Down to 3 remote failures, all one chain. The explicit no-events reporting is working — the diagnostic now states the events file does not exist, instead of silently printing the generic message. But it then ASSERTED a cause: "the child was killed before the reporter wrote even one event (process/spawn startup stall, not a test hang)". That was a guess dressed as a finding, and the fixture contradicts it: it starts a real test, so test:start should fire in milliseconds against a 2000ms budget. Two hypotheses remained and I could not separate them locally, because the local test runner is hook-blocked here: either the custom reporter never LOADS in the child, or it loads and no event reaches it before the SIGKILL. Rather than guess a third time, the artifact now answers it. The reporter appends a reporter:init line as its first action, before consuming anything, so the file's contents discriminate: absent means the reporter never loaded; init-only means it loaded and saw no test events; init plus events means it works. The diagnostic has a branch for each and, where the cause is genuinely unresolved, names both possibilities instead of picking one. Also passes the reporter as a file:// URL via pathToFileURL. Node documents the --test-reporter value as an import()-style specifier, and a bare absolute path is not a portable one — notably on Windows. That is a correctness fix whichever hypothesis holds, and it is a live candidate for the first. FIXED_OVERHEAD is derived by reducing over the actual argv strings, so the longer URL is accounted automatically. Verified by execution, not assumption: composing the reporter against an EMPTY event stream writes exactly one line, the init marker. That is the whole point of the marker, so it is pinned by a test rather than left to inspection. T1 stays red and untouched. Verification runs on the remote runner. Refs #4012 |
||
|
|
5504724d7a |
fix(#4012): the reporter body must return nully, not an iterable
Fourth failure on this feature, and this one was caused by the previous fix's lint workaround. TypeError [ERR_INVALID_RETURN_VALUE]: Expected nully to be returned from the "body" function but got an instance of Array. When stream.compose is given an async FUNCTION as the body, that function must return nully. The reporter ended with 'return []' under a comment asserting Node "still requires the exported function to return an iterable" — exactly backwards, and the direct cause of 41 failures across every run-tests.cjs invocation. That return existed only to dodge ESLint's require-yield after the previous commit converted async function* to async function. A lint workaround became a runtime crash, and the comment written to justify it stated the opposite of the contract. Both are now corrected to what the runtime actually does. Verified by EXECUTION rather than by reading: composing the real reporter against a fake event stream completes with no error and leaves both handled events durably on disk, with the ignored event type skipped. The failing form was reproduced the same way first, so the diagnosis is not inferred from the stack trace alone. The new unit test requires the reporter directly and asserts the returned value is nully — the assertion that would have caught this before it reached the runner — plus the exact NDJSON written. No subprocess, so this half of the feature is verifiable without a full runner pass, which matters because every defect in this feature so far has only been observable remotely. T1 remains untouched and red. Verification runs on the remote runner. Refs #4012 |
||
|
|
7f2af28639 |
fix(#4012): the reporter sink must be a regular file, not devNull
Third failure on this feature, and this one broke everything rather than just the diagnostic: 43 failures across every run-tests.cjs invocation. Error: EINVAL: invalid argument, fsync Emitted 'error' event on WriteStream instance Node opens a WriteStream for a --test-reporter-destination and FSYNCS it on close. fsync on /dev/null is EINVAL — it is a character device, not a regular file. So os.devNull is not a usable reporter destination at all, and every chunk crashed on exit. The sink only ever needed to be a regular file that stays empty, since the reporter writes its real output through appendFileSync to the path in GSD_RUN_TESTS_EVENTS_FILE. It is now one fixed file inside the existing events dir, pre-created rather than relying on the stream's create-on-open, and left empty by design. One sink for the whole run, not per chunk, so its path length stays constant — reporterOverhead feeds FIXED_OVERHEAD, which is computed once before chunking, and a variable-length path would silently mis-account the Windows argv ceiling. The comment that named devNull now says why the destination must be a regular file, so this does not get re-optimized back into the same crash. Reaching for devNull was the mistake: it looks like the obviously correct way to discard output, and it is, for a pipe or an fd — but not for something Node is going to fsync. Each of the three failures on this feature was a different edge of the same assumption, that a reporter destination behaves like ordinary output. The regression test asserts the closest externally observable consequence — a normal run must not surface the EINVAL/fsync text. The argv construction lives inside main() with no exported seam, and the sink is swept before a test could stat it; adding a seam purely to assert that is left out rather than reshaping production code for the test. Stated plainly rather than implied. The failing T1 is untouched and still red. Verification runs on the remote runner. Refs #4012 |
||
|
|
0abd137ec7 |
fix(#4012): the events reporter has to survive SIGKILL
The remote run proved the instrumentation did not work. Timing landed — "chunk 1/1 was killed after 2006ms" — but the in-flight-file naming produced nothing and fell through to the pre-existing generic message. The feature I wrote to diagnose a kill was itself destroyed by the kill. Root cause, confirmed rather than assumed. The reporter yielded strings, which node pipes into the --test-reporter-destination WriteStream. That stream BUFFERS. execFileSync's timeout sends SIGKILL, which is uncatchable and gives nothing a chance to flush, so the events sat in a buffer that died with the child. The parent's own timer reported correctly because it lives in the parent — which is exactly why half the feature looked fine. The reporter now writes each event with fs.appendFileSync, unbuffered and durable at the moment it happens, to a path passed through GSD_RUN_TESTS_EVENTS_FILE. Env vars do not count toward the Windows 32,767-char argv ceiling, so moving the path out of argv also REDUCES FIXED_OVERHEAD; the accounting moved with it rather than being left stale. The destination is now a fixed devNull sink that stays empty by design. Silence was the reason this was invisible for a whole run. Failing to read the events file now says so explicitly, and distinguishes a file that could not be read at all from one that exists but is empty — the generic fallback firing quietly is what let a broken feature look like a working one. A write is unbuffered but not atomic, so a kill can still interleave a partial line; the reader tolerates exactly one unparsable trailing line and reports the complete ones before it. The failing T1 was left red and untouched rather than weakened to pass. Three new unit tests cover the reader directly, with no subprocess, so the parsing half is verifiable without a full runner pass: missing file, existing-but-empty file, and a truncated final line. Also adds ndjson-reporter.cjs to GSD_SCRIPTS_LIB_FILES in bin/install.js — scripts/lib/ ships, and omitting it meant the file would install everywhere and orphan on uninstall. That single omission caused 4 of the 7 remote failures. Verification runs on the remote runner. Refs #4012 |
||
|
|
11c24ce973 |
chore(#4012): regenerate the 19 install-tree goldens
scripts/ ships, so a new file under scripts/lib/ moves every install-tree golden. The remote runner caught this as 26 failures across all 19 runtimes. My fault in the dispatch: I limited the implementation's verification to build:lib and lint:ci and left regen:derived out. Both of those passed, which is precisely why a green local gate is not a substitute for the runner — the goldens are only exercised there. One line per golden: scripts/lib/ndjson-reporter.cjs joining the shipped tree. Refs #4012 |
||
|
|
8497833a15 |
fix(#4012): a killed chunk now names the file that was hanging
The per-chunk timeout fired correctly but reported almost nothing, so every diagnosis cost a CI round-trip. run-tests.cjs logged chunk START only — no timestamp, no duration, no end line — then on a kill printed all ~55 basenames and asked the operator to work out whether output kept flowing (slow) or stopped early (hang). It could not name the in-flight file because the child is spawned with stdio inherit, deliberately, per #3597/#1051. Three additions. Per-chunk elapsed timing on every path, not just failures, so drift toward the cap is visible before it becomes a kill. Every timing number in the investigation behind this had to be reconstructed by hand from GitHub log timestamps. A second, machine-readable reporter running ALONGSIDE the human one, writing NDJSON to its own file. On a kill that file is read back and the files with a test:start and no matching completion are named, with the staleness of the last event, so "stopped 480s ago at X" reads differently from "still emitting at kill". stdio stays inherit and nothing is piped or tee'd — the maxBuffer and live-output risks that shaped the original design are untouched. Ranking of the killed chunk's files by known weight, flagging any absent from tests/test-timings.json, since an unweighted file is an unknown quantity. Two details that are correct rather than lucky. Passing --test-reporter at all replaces node's implicit default, so the human reporter is now named explicitly and reproduces node's own selection (spec on a TTY, tap otherwise) — visible output is unchanged. And the destination path's chunk index is zero-padded to a fixed width because FIXED_OVERHEAD is computed ONCE before chunking; a variable-length path would have silently mis-accounted the Windows 32,767-char argv ceiling and reintroduced #3597. The reporter flags are added to FIXED_OVERHEAD exactly as --test-force-exit is. The multi-reporter pairing and the stream.compose reporter contract were confirmed against Node's v24 documentation, not recalled — the first draft carried them as an unverified assumption and said so. Also corrects a stale comment claiming the 600s cap sits "below the 20m job cap". The lane is sharded 3x at timeout-minutes: 45; the windows shards were at 19m when chunk 1/5 was killed on |
||
|
|
83273f9642 |
fix(#3798): the profile closure follows command references into workflow spawn surfaces (#4009)
* test(#3798): tiered profiles must install the agents their workflows spawn * fix(#3798): the profile closure follows command references into workflow spawn surfaces * chore(#3798): changeset fragment (pr number backfilled after PR creation) * chore(#3798): backfill changeset PR number (4009) --------- Co-authored-by: sim <sim@local> |
||
|
|
ac3668e4b7 |
fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow (#4008)
* test(#3797): the roadmapper must follow one write-first contract * fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow The roadmapper contradicted itself: role blurb, output format, and completion checklist described an approve-first flow while its execution flow said "Write Files Immediately" with reactive-only revision (#3797). The approval gate belongs to the ORCHESTRATOR — both callers read the written ROADMAP.md, present it, and gate on approval (with an auto-mode bypass a subagent cannot host) — so write-first is the contract. All approve-first text now describes the write-then-return reality, the old "Draft Presentation Format" (whose ## ROADMAP DRAFT header matched no orchestrator branch) is folded into the ## ROADMAP CREATED structured return as a preview block, and the duplicate checklist lines are merged. A structural guard pins the single contract. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — #3797: +bytes — approve-first wording replaced with write-first descriptions; the DRAFT presentation template folded into the ROADMAP CREATED return as a preview block * chore(#3797): changeset fragment (pr number backfilled after PR creation) * chore(#3797): backfill changeset PR number (4008) --------- Co-authored-by: sim <sim@local> |
||
|
|
004c7532b6 |
fix(#3796): write the audit report to the single-version filename every reader expects (#4007)
* test(#3796): the audit report writer and readers must agree on the filename * fix(#3796): write the audit report to the single-version filename every reader expects * chore(#3796): changeset fragment (pr number backfilled after PR creation) * chore(#3796): backfill changeset PR number (4007) --------- Co-authored-by: sim <sim@local> |