bc5619dd272efe48d6933b48af4eea8410f8e326
2561 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a5b5860b93 |
fix(#3238): bump js-yaml to the patched 4.3.1 (high-severity devDep advisory) (#3246)
* chore(#3238): bump js-yaml 4.3.0 -> 4.3.1 (GHSA-5p4m-2wfm-xmqj)
Dependabot alert 14. js-yaml 4.3.0 sits inside the vulnerable range
>=4.0.0 <4.3.1 of GHSA-5p4m-2wfm-xmqj -- high, CVSS 7.5
(AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H), CWE-407 Inefficient Algorithmic
Complexity.
resolveYamlOmap() enforces `!!omap` key uniqueness with a linear
objectKeys.indexOf() scan inside the per-element loop, so resolution is O(n^2)
in entry count. `!!omap` is registered in the DEFAULT schema, so a plain
yaml.load(untrustedInput) with no options is affected. The loop is synchronous,
so it blocks the event loop -- amplification is per-process, not per-request.
Same weakness as CVE-2026-59870, fixed in 5.x at 5.2.1 and only now backported
to the 3.x/4.x lines.
Reproduced locally against the installed 4.3.0, advisory PoC, default schema:
n= 5000 load= 22ms -
n=10000 load= 57ms 2.59x
n=20000 load= 199x 3.49x
n=40000 load= 768ms 3.86x <- ~4x per doubling = quadratic
After the bump, same machine, same PoC:
n= 5000 load= 33ms -
n=10000 load= 33ms 1.00x
n=20000 load= 49ms 1.48x
n=40000 load= 96ms 1.96x <- ~2x per doubling = linear
Duplicate-key rejection is preserved (YAMLException still raised), so the
upstream indexOf -> Set swap kept the semantics it was guarding.
Scope is development-only and stays that way: js-yaml is a devDependency and is
absent from package.json's `files` allowlist, so it never ships to consumers.
`npm audit --omit=dev` reported 0 vulnerabilities before this change and still
does; `npm audit` went 1 high -> 0.
Targeted install rather than `npm audit fix`, so the blast radius is auditable:
npm reports "changed 1 package", and the lockfile diff is 4 insertions /
4 deletions touching only js-yaml. The declared floor moves ^4.2.1 -> ^4.3.1 so
a future resolution cannot land back on a vulnerable 4.3.x -- the operative fix
is the lockfile, since npm ci is lockfile-driven.
Stayed on the v4-legacy line (4.3.1) rather than jumping to `latest` 5.2.3: 5.x
is a rewritten module layout and a separate change with its own blast radius.
The v4-legacy dist-tag exists precisely so 4.x consumers can take this patch.
Regression guard mirrors tests/issue-2765-brace-expansion-lockfile.test.cjs --
the repo's existing precedent for a dev-scope lockfile bump against a
high-severity DoS advisory. It walks `npm ls --json --all` so a transitive copy
left behind still fails, and carries a vacuity guard so an empty version list
cannot pass silently. No timing assertion: wall-clock assertions are barred by
the clock-seam rule and would be load-sensitive on shared benches, so the
measurements live in the diagnosis artifact instead.
Refs #3238
* fix(#3238): require 5.2.1 on the 5.x line; add the release-notes fragment
Three review findings, all fixed inline.
SPEC AXIS (the serious one): the guard's `(maj > 4)` clause accepted ANY 5.x.
GHSA-5p4m-2wfm-xmqj names only the 3.x and 4.x ranges, so the isolated
adversarial pass -- checking strictly against that advisory -- rated 5.0.0 as
correctly accepted. But the advisory's own body records that the SAME weakness
in the 5.x line is CVE-2026-59870 / GHSA-724g-mxrg-4qvm, fixed in 5.2.1. A
guard whose purpose is "this tree has no quadratic !!omap resolver" must
require 5.2.1 there too, or an accidental major bump to 5.0.0 silently
reintroduces the exact bug the test exists to prevent. The two reviewers
disagreed and the disagreement was load-bearing: taking only the adversarial
verdict would have shipped the hole.
ISOLATED ADVERSARIAL (item 4): Number('4.3.1-beta.1') produced NaN, and NaN
comparisons made the predicate return false. That failed safe, but by accident
rather than design, and majors outside {3,4} were reported vulnerable despite
being outside every advertised range. The predicate is now explicit -- build
metadata stripped, prerelease fails CLOSED (4.3.1-beta.1 sorts below 4.3.1 and
may predate the fix), unparseable fails closed, maj<3 accepted as predating the
affected lines.
Validated by a standalone harness over 27 version strings (12 accepted, 14
rejected, 1 build-metadata): all 27 agree with the advisory ranges. Boundary
rows on all three affected lines -- 3.15.0/3.15.1, 4.3.0/4.3.1, 5.2.0/5.2.1.
STANDARDS AXIS: the precedent this change mirrors,
|
||
|
|
c75ce93be9 |
feat(#1954): flag undeclared coupling between same-wave plans (#3237)
* test(#1954): failing-first contract for plan-checker undeclared-coupling check * feat(#1954): flag undeclared coupling between same-wave plans * docs(#1954): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
70e50eb1b2 |
chore(#3235): hoist the preamble-strip conditional out of the replace operand (#3236)
* chore(#3235): hoist the preamble-strip conditional out of the replace operand
CodeQL js/identity-replacement (alert 53, medium, CWE-116) fired on
src/roadmap-parser.cts:637. The #2947 fix (
|
||
|
|
2a73f53cb3 |
fix(#3204): milestone sectioning is vocabulary, not heading position (#3230)
* test(#3204): failing-first suite for the clobbered phase count A project declaring six phases with four phase directories on disk had state.record-session write progress.total_phases: 4 — #2828 regressing at 1.9.1, reported in #3204 with a deterministic reproduction. Before the fix in the following commit, these rows FAILED (wrote 4, expected 6): a flat roadmap carrying `## Progress`; one carrying `## Overview` and `## Phase Details`; the CRLF variant of the first. Two more, found by adversarial review and added after the first fix attempt, failed against that attempt: structural headings interleaved among flat phase headings, and this repo's own bundled-template shape (a `## Phases` wrapper around a single nested milestone). The #1761 control — sibling milestone sections must keep falling back to the disk count — passes both before and after, so the fix has something it must not break. Assertions read progress.total_phases through the product's own frontmatter parser via `state json --raw`, never a regex over STATE.md. Rows 12 and 13 are hostile: a phase heading carrying a version token, and a version heading inside a fenced code block; neither may count as milestone sectioning. Refs #3185, #3204 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3204): milestone sectioning is vocabulary, not heading position buildStateFrontmatter chooses total_phases between the ROADMAP's declared phase count and the on-disk directory count, and refuses the roadmap count when hasMilestoneSectioning says the document is milestone-sectioned — because a whole-document count would then conflate sibling milestones (#1761). That predicate returned true for ANY non-Phase level-2/3 heading, so a flat roadmap carrying an ordinary `## Progress` was called sectioned and the disk count clobbered the declared one: six declared phases, four directories, total_phases written as 4, converging on the truth only once the last directory happened to exist. That is #2828 regressing at 1.9.1, and it came from this epic — #3184 replaced state.cts's hand-rolled #2828 guard with this predicate, and the replacement is strictly more permissive than the guard it retired. Three position-based models were tried and all failed, because position does not carry milestone-ness: - any non-Phase heading (shipped) — over-detects, giving #3204; - strict nesting/ownership — misses same-level siblings, regressing #1761, and false-positives on the bundled template, where `## Phases` wraps a single `### v1.1`; - adjacency — reproduced live: `## Overview` and `## Notes` interleaved among six phase headings are two owning candidates, so a 6-phase roadmap with 2 directories wrote 2. A heading is now a milestone heading iff it is a non-Phase heading carrying a milestone signal: a version token, a status marker, or the word Milestone. Sectioning means two or more, since one cannot conflate siblings. Known limit, recorded in the doc comment rather than hidden: two milestone sections carrying none of those three signals are not detected. Also drops buildStateFrontmatter's local dedup-key regex, flagged in-source as diverging from the canonical token rule, for phaseKeyFromDir — the remainder of #3185, since #3222 had already routed the enumeration itself through listMilestonePhaseDirs. #1514, #2445 and #3017 are preserved untouched. Closes #3185 Fixes #3204 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3185): changeset, glossary entry and ADR status for the phase-count fix CONTEXT.md's Roadmap Parser Module entry never named hasMilestoneSectioning, so the predicate whose semantics this change reverses had no glossary presence at all — a PR gate for a module/seam change. Added, covering the vocabulary model, the three position-based models that failed, and the residual limit. ADR-3180 recorded the fifth enumeration copy as unowned in four places. It is owned now. Amendment 4's scope table row 1 also carried an error worth keeping visible rather than rewriting: it claimed Phase 3 merged without routing the state writers, when #3222 had in fact routed the enumeration — the audit read Amendment 3's silence about the symbol names as absence of the work. The real gap was the trust discriminator one layer above, which is what #3204 was. Changeset is Fixed and leads with the symptom a user sees — a phase count that shrinks to match how many phase directories happen to exist yet — and carries the known limit forward rather than leaving it in a source comment. Refs #3185, #3204 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3185): stop quoting the retired phase-token regex in a comment The remote runner failed tests/phase-id-drift-guard.test.cjs: the comment explaining that the local dedup regex had been replaced by phaseKeyFromDir quoted that regex verbatim, and scripts/lint-phase-id-drift.cjs scans for the literal token without caring whether it sits in code or in prose. That is the guard being right, not over-eager — a quoted pattern is one paste away from being live again, which is exactly how the copy it replaced spread. Described in prose instead. Worth recording: this guard is check:phase-id-drift, which lint:ci does not run — it is enforced by tests/phase-id-drift-guard.test.cjs. A green lint:ci is therefore not evidence the drift guards pass. Refs #3185 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3185): backfill changeset PR number (#3230) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
86bebcefa2 |
refactor(#3216): bind milestone identity to the canonical locator (#3226)
* refactor(#3216): widen milestone-window guard to literal-## matchers The guard keyed only on the `#{N,M}` quantifier plus a literal version or phase-lookahead token. getMilestoneInfo hand-rolls its milestone-heading match with a literal `^##`/`## ` and an interpolated ${escapedVer}, so it satisfied neither token and the guard reported a clean zero on a file carrying live re-derivations (#3171, #3197) — a zero it did not earn. Widen token (a) to a literal 2-6 `#` run, admitted ONLY inside a heading-MATCHER literal (a regex literal, or a string/template handed to new RegExp) so a heading-BUILDING template is not mistaken for a re-derivation. Widen token (b) with the grouped `v(\d+(?:\.\d+)+)` shape and an interpolated version placeholder. Ships BEFORE the consolidation per ADR-3180 s7.2: a guard widened afterwards measures an already-cleaned surface. It is expected to be RED until the consolidation lands. * test(#3216): failing-first milestone-identity single-owner suite 63 tests across two files, from the matrix in .gsd/phase/. Section H of milestone-window-single-owner.test.cjs covers the 21 input classes of the design's behavior table plus its negative space; milestone-window-drift-guard covers the widened tokens and proves the exemption is function-scoped, not file-scoped. Copy count is 3 found by the guard, not 1 per the epic (ADR-3180 Amendment 3's standing rule, holding for the fourth consecutive phase): both getMilestoneInfo sites plus cmdRoadmapAnalyze's milestone enumeration at roadmap.cts:454, which carries the same #3171 truncation and #3197 phase-heading confusion. Expected RED until the consolidation lands. * refactor(#3216): bind milestone identity to the canonical locator getMilestoneInfo hand-rolled two milestone-heading regexes inside the owner's own file. Both were wrong, differently: the STATE-version site's ^## anchor is level-blind so [^\n]* absorbs a third #, and the fallback site had no anchor at all, so '## ' matched from the second # of '###'. Against '### Phase 7: Close v3.3 gaps' the fallback returned {v3.3, gaps} (#3197). Both captured names with [^\n(], truncating at a parenthetical (#3171). Bind both to the canonical grammar. locateMilestoneHeadings becomes a version-filtered view over one shared source, and a new version-agnostic listMilestoneHeadings enumerates milestone headings for callers that need all of them. getMilestoneInfo returns ScopedResult<MilestoneInfo|null>; the {v1.0,'milestone'} default, which was output-identical to a real v1.0 project, is deleted. The #2245 never-throws invariant is preserved. Copy count: 3 found by the guard, not 1 per the epic. The third was cmdRoadmapAnalyze's own milestone enumeration (roadmap.cts:454), carrying both defects in the implementation the epic blessed. buildStateFrontmatter and archivePhaseDirectories branch on scope: the first writes null rather than a fabricated identity, the second falls through to its dated-label fallback. A fabricated v3.3 passes ARCHIVE_VERSION_LABEL_RE, so it would otherwise misfile phase history. Also fixes an unsafe cast in init.cts that masked these type errors across five call sites, which would have shipped undefined milestone fields under green tsc. * fix(#3216): restore the #1761 unbounded guard and bullet precedence Review and the first full-matrix run surfaced five real defects in the consolidation, all fixed here rather than by relaxing the tests that caught them: - buildStateFrontmatter gated its isMilestoneBoundedInRoadmap check on the scope-gated milestone value, which is null on any non-COMPLETE scope, so the #1761 unbounded guard was silently skipped and state json reported a percent it must omit. It now gates on the STATE-asserted version, independent of identity scope. - The rewrite lost #2135's precedence: the name-bearing progress-marker bullet is consulted before the heading again. - A single-segment version (v3, no dot) did not resolve; the name-extraction fallback now accepts it. - A version carrying regex metacharacters, or a $& / $1 replacement pattern, is matched literally. - listMilestoneHeadings' heading field trimmed, so a CRLF roadmap no longer leaks a trailing carriage return into roadmap analyze's output. Also emits milestone_version / milestone_name / current_milestone as explicit null rather than omitting the key, so the prompt layer cannot render a bare placeholder, and corrects an init.cts comment plus a cast left inconsistent. * test(#3216): update milestone-identity expectations to the scoped contract getMilestoneInfo returns ScopedResult<MilestoneInfo|null> and the {v1.0,'milestone'} default is deleted, so the suites asserting the old shape assert removed behavior. Updated rather than weakened: every touched call site now asserts the scope explicitly against the frozen SCOPE enum. roadmap-parser.test.cjs: 20 expectations moved to {value,scope}. The #1881 unreadable-vs-absent diagnostic assertions are untouched and still prove their original point — only the return shape moved. One pre-existing assert.ok(info) is now a specific UNSCOPED assertion, so that case is stronger than before. new-milestone-clear-phases.test.cjs: the test asserting phases clear archives under the v1.0 default now asserts the dated archived-<YYYYMMDD> fallback, which is the deliberate consequence of deleting that default. Two of this branch's own tests were also corrected after they drove the implementation the wrong way: the parity test compared raw heading text and so pushed a stray ## prefix into roadmap analyze's public output, and the hostile metacharacter row demanded a pathological version resolve, which pushed a widening of the ADR-locked \b boundary. Both now assert what the contract actually requires. * docs(#3216): document milestone identity and correct the CONTEXT.md entry ADR-3180 s7.2 moves to Enforced and gains two rules that were unstated: the name derives from the heading's own version token and drops a trailing status marker, and a free-form legacy ROADMAP with no version anywhere is UNSCOPED with no identity rather than a defaulted v1.0 (decided by the maintainer before implementation, per s7's own rule that an unstated behavior is not decided). Amendment 4 records Phase 6's validation, including that the copy count was a lower bound for the fourth consecutive phase. CONTEXT.md's Roadmap Parser entry described locateMilestoneHeadings as boundary-matched with (?![\w.-]) — the alternative Amendment 2 tried and REVERTED. The code uses \b and says so, and the ADR agrees; the revert updated code and ADR and missed CONTEXT.md, which is the epic's own fixed-on-one-copy failure class in the docs layer, on a file that is itself a PR gate. * fix(#3216): persist the real version on a truncated identity buildStateFrontmatter wrote null for BOTH milestone and milestone_name on any non-COMPLETE scope, discarding a real version. ADR-3180 s7.2 rule 6: a version known with no resolvable name is TRUNCATED carrying {version, name: null} — 'the version is a real answer, the name is a non-answer, and collapsing the two is the failure this contract exists to prevent.' The two fields are now gated by what is actually known: the version whenever one exists (COMPLETE or TRUNCATED), the name only on COMPLETE. Never fabricated. Caught by this phase's own Decision 4(c) consumer-output test, which is the argument for asserting at the consumer rather than the owner — the owner was correct throughout; only the consumer collapsed its answer. * refactor(#3216): extract helpers and make cmdCommit's scope gate explicit From the two-axis code review: - init.cts repeated the identical getMilestoneInfo cast at five sites with copy-pasted comments — duplication inside a PR whose thesis is that duplicates get deleted. Extracted milestoneRecord(cwd); the one site-specific comment is kept, the four generic copies removed. - getMilestoneInfo hand-built its { value, scope } literal at ten return points; a local scoped() constructor now does it once. Every per-branch rationale comment is preserved and no returned value or scope changed. - cmdCommit gated the milestone branch name on plain truthiness, which is also true for TRUNCATED, so an unresolved identity drove branch creation incidentally rather than deliberately. It now gates on the SCOPE enum, accepting COMPLETE or TRUNCATED because both carry a real version, and the comment records why that differs from archivePhaseDirectories — which demands COMPLETE because it uses the value as a filesystem path component. * test(#3216): cover the bare-version-in-prose truncated path The spec review found the bareVersionMatch path — no STATE version, no milestone heading, a version token only in prose — returning TRUNCATED with no test exercising that exact shape, violating Decision 4's boundary-coverage requirement. * docs(#3216): record the missed Tier-2 surfaces and rule 5's corollary Decision 3 requires an explicit call-out for EVERY Tier-2 change, and Amendment 4's first draft named eight surfaces while the change touched thirteen. Adds cmdCommit's branch-name construction and the four init JSON bundles, an incomplete list being the same defect in miniature that this epic removes. s7.2 rule 5 gains a corollary separating two cases the original wording ran together: no version token ANYWHERE is UNSCOPED, while a bare version token in prose or a non-milestone heading is weak but real evidence and yields TRUNCATED under rule 6. * chore(#3216): set changeset fragment pr to 3226 --------- Co-authored-by: sim <sim@local> |
||
|
|
b9f51836e6 |
refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a lower bound for the third consecutive time, and that two derivation families had never been named at all. ADR-3180 gains Decision 7 — a normative behavior contract that says what the right answer IS for each derivation, not merely who owns it. A reviewer with no written rule can only ask "does this look like the others", which is how a fifth copy passes review. Decision 4 gains (d) scan surface is every authored surface and an owner FILE is never exempt, only its named functions; and (e) a surface that cannot be consolidated today ships ratcheted, never unguarded. Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined copies of its own body across five modules. All six now route through it; `clampPercentFromFraction` is added for the one caller that already held a fraction. Every migration is behaviour-identical — clampPercent's first line IS the `total > 0 ? … : 0` ternary each copy carried. Guarded by lint-completion-ratio-drift.cjs, which reports zero re-derivations with no file-level exemption. Prompt layer: workflow markdown re-derives live-plan counting in raw shell (#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs scans it with a shrink-only baseline of the 7 sites that exist today — new sites fail, and a baseline entry that stops firing fails too, so an acknowledgment can never outlive the thing it describes. lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only the four named canonical functions are exempt now. The blanket exemption was pointed at the one file most likely to grow the next copy, and it had. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage Five findings from the two orthogonal review passes, all fixed. Decision 4(c) breach: the completion-ratio identity test asserted at the OWNER, which is exactly the bypass that decision exists to close — a consumer can call clampPercent and then post-process locally, leaving both the lint and an owner-level test green. It now drives `roadmap analyze`, `query progress` and `stats` and asserts on their own output, over a fixture containing a `status: superseded` plan so a consumer that re-counted raw files would report 60 where the owner reports 75. Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the issue that removes them. They name Phase 8 (#3218) now. The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical sites were one indistinguishable key and migrating either would have left the guard green with the other alive. Entries carry an occurrence count; fewer than acknowledged fails as a partial migration, more fails as a new copy. Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's test already had, and the fast-check property tests CONTRIBUTING requires for clamp/budget-limit functions. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes) `tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs` under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`. A fixed wall-clock budget around a double spawn, running inside a container that is concurrently executing the full ~31k-test suite, fails by construction under load. Confirmed against three full matrix runs. Every failure was shaped `null !== 0` — the child was KILLED, never an assertion about the thing under test. One captured probe had already printed the correct resolution (`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The victim subset varies by run and by lane. What these tests are actually about is suite-token RESOLUTION — `unit` as a bare token in --files/--files-from. Executing the seeded trivial files is incidental and is the entire timeout surface, so the assertions move in-process against the same functions `main()` calls, in the same order. `parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are exported for that; no behavior, signature or logic changed. No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the harness for real and asserts exit codes end to end, on a 120s budget. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: delete the three elapsed-time assertions CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all three are load-sensitive: on a saturated bench each can fail while the code under test is correct. In every case the load-bearing assertion sits on the line above and the timing line adds no discrimination. run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s harness backstop?" — is already answered by the assertion above it. A backstop kills by signal, which surfaces as status null, never 124. Observed directly this session: three matrix runs produced exactly that null shape from killed children. normalize-test-command and context-predicates: both bounded a ReDoS check. A threshold only ever separates "fast" from "slightly slow", which is bench load, not correctness — catastrophic backtracking on 800 KB of input does not take 251ms, it does not finish at all. A real regression therefore shows up as the suite being killed on that test, which is louder and more reliable than a number. The structural assertions (returned unchanged; cleanly rejected) are what actually carry those tests, and they stay. The sweep now reports zero elapsed-time assertions in tests/. The remaining Date.now() uses are unique-path suffixes, barrier deadlines, fixture timestamps and fake mtimes — none of them assertions. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3180): backfill changeset PR number (#3223) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows The baseline keys on (file, trimmed text). `file` came from scanTree's `path.relative()`, which uses NATIVE separators, while the committed baseline stores POSIX. On Windows every violation was therefore unmatched — reported as FRESH — and every baseline entry matched nothing — reported as STALE. The guard failed 100% of the time there, on both CI shards: ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale + { file: 'gsd-core\\workflows\\execute-plan.md', ... } The remote runner this repo gates on is Linux-only and cannot see this class at all; the GitHub Actions Windows lane is what caught it. Normalization is unconditional — never gated on process.platform. A platform-conditional normalizer makes the POSIX path the special case and leaves the Windows branch unexercised on every other OS, which is the same blind spot in a different place. It is applied at one seam inside findPromptDrift, which builds `file` on every returned violation, so the baseline key, the --update writer, the stderr report and the tests all consume one normalized value. The regression tests drive a Windows-shaped relPath directly and run on every OS rather than skipping off-Windows — a test that only runs on the platform where the bug lives is why this escaped. They include a sanity check that un-normalized input does NOT match, so the assertion cannot pass vacuously. Audited the three sibling guards: none keys against a committed cross-platform baseline, and their exemption keys are path.join-built, so producer and consumer share the native convention. Left correct code alone rather than making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing there would break those three on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
636ec92107 |
refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite Covers the enumeration rows with direct code evidence: 999.* backlog dirs listed by progress/stats, the phase-0 sentinel divergence, the #1324 letter-prefixed-decimal negative space, and the destructive-path find — cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes 999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there. Also covers the pass-all degrade, which is where the defect actually lives: when the milestone window declares no phases the filter becomes a literal () => true and its heading-side sentinel exclusion is unreachable. A fixture carrying phase headings keeps the filter active and never reaches that path. Named for the derivation, not a module: the suite drives commands, phase, milestone, workstream-inventory and state, and both the phase and phase-locator buckets are already at the per-module test-file cap. Committed alone so the remote runner records the failure before the fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): phase enumeration has one owner and a decidable scope Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner of "which phase directories belong to the current milestone". It applies the milestone window AND the sentinel filter and returns a ScopedResult, so a caller can tell a genuinely-empty milestone from an enumeration that could not be scoped. The sentinel test now runs against DIRECTORY NAMES and is unconditional. getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but degrades to a literal () => true pass-all predicate when that set is empty -- at which point the heading set is never consulted and its sentinel exclusion is unreachable exactly when it is needed. That degrade is the #3167 path, and it is why stats already used the filter and still listed backlog directories. The narrowing is sentinel-only: pass-all stays over-inclusive otherwise. Sentinel copies deleted, canonical isSentinelPhaseId adopted: - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded 999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a 999 heading produced a row with no directory; that seed is filtered now. cmdPhasesList routes only its ENUMERATION. --phase lookup searches the physical set (scoping it would report an out-of-window phase as not found) and --include-archived still merges archived dirs (they are by definition from other milestones). Both exempt by documented reason, never a file allowlist. Fixed inline, found while building: isDirInMilestone could not match a #1324 letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0 heading, so stats reported the phase with plans: 0 while its directory held plan files. Defers to phase-id's extractPhaseToken rather than widening a fourth bespoke regex; additive, so it can only admit directories. Refs #3180. Closes #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): route the last two enumeration re-derivations workstream-inventory countRoadmapPhases counted every `Phase` heading across the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.* backlog and Phase 0 and spanned every milestone the document ever had. Its own caller already resolved a currentVersion and passed it to getMilestonePhaseFilter elsewhere in the same file; this was the sibling copy that never got the fix. state.cts phaseInventoryProvider enumerated phase dirs with its own /^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md inventory carried backlog and sentinel directories as current-milestone phases. A non-COMPLETE enumeration scope now throws to the outer catch as a real scan failure rather than reporting a confident undercount, mirroring the per-phase scanPhasePlans contract beside it. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the sentinel rule re-implemented 23 times across 8 modules, in three regex variants plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped through them while roadmap.analyze and the engine-wide convention (#1580) both treat 0 and 999 alike. That disagreement is the defect class this epic removes. All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites: init recommended-actions and backlog counts, milestone phase scan, the phase-lifecycle progress table, phase.cts used-number collection and the four renumber-on-remove guards, roadmap-parser's heading and bullet milestone counts, roadmap get-phase fallbacks, and state's heading denominator. Excluding Phase 0 at these sites is a deliberate behavior change and the point of the consolidation — several carried comments already saying 0 should be excluded while the literal beside them caught only 999. Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the whole src/ tree with no file allowlist and reports both shapes: an independent phases-dir enumeration, and an independent sentinel literal. Exemptions are function-scoped with a written reason. The guard is comment-aware — its first pass flagged JSDoc and a comment documenting that the code below uses the canonical owner, which would have trained readers to exempt prose. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero Per-site triage of the 31 remaining whole-repo guard hits, applying the rule generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or MUTATION pass wants the physical set; only "which phases belong to this milestone" wants the scoped set. Routed (10): init new-milestone phase_dir_count, init milestone-op fallback count, init manager, init progress, milestone complete stats/dry-run/archive move, phase complete's next-phase scan, state update-progress, state frontmatter stats, and uat audit's active set. Exempt with a written function-scoped reason (never a file allowlist): the audit/UAT/verification sweeps that deliberately scan every directory to report gaps, phase create/insert/rename/renumber mutations, single-phase lookups, roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree destructive pass, and the reads that list a phase dir's FILES rather than enumerating the phases dir at all. Latent defects fixed by the routing: sentinel directories leaked into cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone filter with no sentinel exclusion, so `milestone complete` was archiving backlog directories. scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and npm run lint:ci is green. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3 Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the scoped output of progress, stats, phases list, phases clear and milestone complete, the CONTEXT.md Phase Locator glossary entry naming listMilestonePhaseDirs, and ADR-3180 Amendment 3. Amendment 3 records: the SCOPE contract held unchanged; the declared deviation from Decision 1's provisional signature (the window needs cwd/ws, which the locked roadmapContent parameter cannot supply); the copy count being a lower bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing finding that the sentinel exclusion sat on the heading set and was unreachable under the pass-all degrade; the two destructive-path defects; and the generalized exemption rule. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): wire scope to consumers; revert two wrong routings the suite caught Review + remote runner findings, all fixed: The three consumers computed the enumeration scope and threw it away, so TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE -- reproducing this epic's own output-identical-failure defect one layer up. progress, stats and phases list now emit phase_scope (null on the phases list --phase lookup path, which performs no enumeration). Two routings were wrong and the suite proved it: roadmap-parser's two milestone phase-count scans are reverted to the 999-only literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's decimal phase ids as sentinel milestone 0 and stopped counting them. state.cts phaseInventoryProvider is reverted to the physical disk scan. `state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy trees whose fixture resolves no window, swallowed the raw readdirSync fault message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md rows, which is the job. Both are now function-scoped guard exemptions with written reasons, not silent reverts. This is the consolidation trap named in the epic: a canonical rule can cover MORE than the copy it replaces, and only real inputs show it. Adds phases list coverage, a scope-branch test, and a drift-guard unit suite; backports comment-awareness to the milestone-window and plan-count guards so all three siblings share one false-positive profile; names #3161 alongside #3167 in Amendment 3's subsumption record. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification An isolated security review caught this branch committing the epic's own sin: the over-broad predicate was worked around at ONE call site and left live at the destructive ones. isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id whose leading digit run is all zeros before a non-digit captures 0 -- "0.1", "00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts disagree with that: #2554 requires "00.1" to be counted as a real phase, and the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel. The rule is asymmetric and now says so explicitly: 999 is sentinel with or without a decimal part; 0 is sentinel only when bare. A decimal phase under either is a real phase for 0 and reserved for 999, because 999 reserves a MILESTONE while 0 reserves a PHASE. Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two scans route through isSentinelPhaseId again and the guard exemption that existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild exemption stays -- that one is a genuine reconciliation-wants-the-physical-set case. Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted isSentinelPhaseId('0.1') === true and so had encoded the defect as expected behavior. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong Reverts the previous commit. The remote suite failed six tests proving it wrong, and the reason is the sharpest finding of this phase. An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1 as sentinel milestone 0 and judged that a defect against #2554. Correcting the canonical predicate broke #2949. Both contracts are pinned and both are right, because they ask different questions: #2554 is this dir part of the current milestone's phase SET? -> count 00.1 #2949 must this phase COMPLETE before the milestone closes? -> 0.x sentinel No single global predicate answers both. isSentinelPhaseId keeps its semantics (0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower 999-only rule (#2554) as a function-scoped guard exemption with a written reason — not a second silent copy. That corrects how Decision 1 reads: "one owner per derivation" governs who computes an answer, not how many questions share it. An over-broad canonical rule is as much a defect as a divergent copy and fails worse, because it looks like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5. Where a review's inference about intent conflicts with a pinned contract, the pinned contract wins; the finding is adjudicated, not fixed. The boundary tables in the enumeration suite are corrected to assert 0.x IS a sentinel, with the layering explained. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * chore(#3185): set changeset fragment pr to 3222 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e705652ba1 |
Merge pull request #2677 from 0xdhx/fix/2665-test-env-base-config-location-vars
fix(#3156): derive the config-location scrub set and close the leaks it cannot reach |
||
|
|
66a4940d6f | Merge branch 'next' into fix/2665-test-env-base-config-location-vars | ||
|
|
b421e95434 | Merge branch 'next' into fix/3174-quick-verification-status-query | ||
|
|
342590c70e |
refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite Covers the 50 input classes in the phase test matrix: scope classification (genuinely-empty vs truncated vs unscoped vs unreadable), the section-end owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c), the milestone.complete refusal with negative proof that no directory moved, the version-token boundary defect, drift-guard behavior, and three fast-check properties over document-shaped generators. Committed alone so the remote runner records the failure before the fix lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * refactor(#3184): milestone windowing routes through one owner Three copies of the milestone section-end walk lived in roadmap-parser.cts — two distinct computeSectionEnd function nodes plus an inline third in getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is now the sole owner and the other two are deleted, not kept in sync by comment. The whole-repo drift guard found what the epic did not: state.cts held three more re-derivations of the same vocabulary — two byte-identical milestone bounding checks carrying a defect neither reported copy has (no boundary after the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning predicate. All three route through the owner now. A composition-level duplicate appeared inside this change's own first pass: getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out of the owner's primitives, and had already diverged on whether to skip a closed milestone heading. sliceMilestoneWindow is the one composition. Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is distinguishable from a genuinely empty milestone — those were output-identical, which is the whole failure class. roadmap analyze emits it (#3165), and milestone complete refuses to archive on anything but COMPLETE rather than pass-all moving every phase directory on disk (#3166). The pass-all degrade is preserved where its premise holds: making the filter deny-all would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. extractCurrentMilestone keeps its signature — 200+ affected symbols across 41 files and 25 process flows — and is a one-line wrapper over the scoped owner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): fence-aware phase detection and one heading-selection owner Review fixes from the two orthogonal passes. The blocker: hasPhaseEntries matched ATX phase headings fence-aware via tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown, so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely empty milestone then classified TRUNCATED and milestone complete refused a legitimate archive — a false positive in the destructive direction, worse than the defect this phase set out to fix. Both that path and getMilestonePhaseFilter own pre-existing bullet scan now run on stripFencedCode, since leaving one meant the owner file gave two different answers to the same question. The selection rule — locate, prefer the non-closed heading, else the first — had been written three more times inside the file whose thesis is single ownership. selectMilestoneHeading owns it; all three sites route through it. The copies were behaviorally identical, so this is de-duplication with no observable change, verified by probing that all three paths select the same heading. roadmap analyze emitting a scope no consumer read left #3165's actual symptom alive, so Route 0 in next.md now treats a non-complete scope as scan-failed rather than as a clean empty scan, and the ADR amendment no longer overstates what shipped. Also: the scope refusal moved above the archive-directory create, so a refusal leaves nothing on disk; the versionOverride comment names all four consumers; COMMANDS.md documents the new guard beside its sibling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#2658): exclude the changelog from the malformed-path scan The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships into that tree, and its #2658 entry quotes both malformed paths while describing the fix that removed them — so the release note documenting the fix trips the fix's own regression test. Red on next before this branch. The installer is correct: a probe over a real --trae --local install found 621 emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope was the defect, not the product. Excluded by exact relative path rather than by loosening the patterns or skipping all markdown — the emitted agent and command markdown is precisely what #2658 was about, so the gate stays strong everywhere it matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): regenerate install-tree fixtures for the shared drift scanner scripts/lib/ ships in the npm package and installer, so extracting the shared tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's install tree. Regenerated via npm run gen:install-tree; the delta is exactly that one path per fixture. The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded), so only the extracted library moves. This matches the existing scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only helper carried in the shipped tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal The remote runner caught two regressions this branch introduced. Both were mine, and neither review pass found them — only running the existing suite did. The version-token boundary. I replaced locateMilestoneHeadings' \b with (?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562 fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a closed v8.0-A sibling (#730), and \b is what allows it while the stricter boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to \b; the state.cts consolidation is now a straight merge with no behavior change, and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and the design doc no longer claim otherwise. The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is about the TRUNCATED window specifically — the heading is found and the section closes before the phase region, so pass-all archives everything. UNREADABLE and UNSCOPED are pre-existing, legitimately handled states, and refusing on them broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed to TRUNCATED; docs corrected to match. One of the new tests was also wrong: its fixture gave the shipped and current milestones' phases the same numeric id, and the filter matches on that id, so it could not have distinguished the two windows. Fixture corrected to exercise what it claims to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): enumerate drift-scan.cjs for uninstall The installer copies scripts/lib/ wholesale, but uninstall removes an explicit set — deliberately, so a user's own helpers in that directory survive. The extracted drift-scan.cjs was copied in and never enumerated, so it outlived uninstall, left the directory non-empty, and the rmdir that follows failed. Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is likewise a lint-only helper that ships there and is enumerated. Verified with a real install-then-uninstall into a temp target: scripts/lib/ held exactly the three GSD files and was gone afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset Found while shipping this phase, and fixed here rather than noted. install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE — the comment at the copy site literally says "and any future lib helpers". uninstall() removes them by hardcoded enumeration, deliberately, so a user's own helpers in those directories survive. A wholesale writer paired with an enumerated remover cannot stay in sync by construction: any file added to either directory ships to every user and is then orphaned in their repo forever, since it survives uninstall, leaves the directory non-empty, and the rmdir that follows fails. Nothing reported this. 31,225 tests were green over it. That is the same divergence class this epic exists to delete, sitting in the installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity assertion that fails the moment the two surfaces disagree. The test compares each directory's real contents against its enumeration and names the offending file plus the constant to add it to. Both enumerations are hoisted to module scope and exported, so the test asserts on the actual arrays rather than pattern-matching the installer's source — no allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on the current tree, correct report when an unenumerated file is injected. scripts/changeset/ turned out to carry the identical defect and is covered too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * chore(#3184): backfill changeset PR number Also narrows the wording to match the shipped behavior: the refusal fires on a truncated window specifically, not on any non-complete scope. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f3ce2dbab5 |
docs(#3156): name the isolation this helper does NOT provide
Found by the pre-push adversarial review of this round, and worth recording in the code rather than only in the PR thread. installSpawnHome() creates one sandbox home per test-FILE process, not one per spawn, so two installer spawns in the same file share .gsd state. The containment claim is unaffected -- nothing reaches the developer's real home -- and it is strictly better than the status quo it replaces, which shared the real home and every byte of its state. But "contained" and "isolated from each other" are different properties, and only the first is claimed. |
||
|
|
78330e505c |
fix(#3174): read quick's verification status via the verification.status query
`gsd-core/workflows/quick/steps/quick-verification.md` read the verifier's result with `grep "^status:" F | cut -d: -f2 | tr -d ' '` and routed it through a table whose only arms were passed / human_needed / gaps_found. That read fails two ways. Driven against the old pipeline: never written (verifier died) -> empty -> no arm off-schema value -> weird_value -> no arm `status:` in frontmatter AND prose -> two lines -> no arm valid `passed` on a CRLF checkout -> passed\r -> no arm stale report still reading passed -> passed -> SUCCESS `status: passed` in the prose only -> passed -> SUCCESS off-schema `passed:bogus` -> passed -> SUCCESS The first four leave the orchestrating agent improvising at the moment the pipeline failed. The last three are silent false passes: staleness was never evaluated, the match was not anchored to frontmatter, and `cut -d: -f2` splits an off-schema value at its own colon. The CRLF row is a pre-existing Windows bug this change closes as a side effect. The unanchored match is DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, which the code side already fixed by name — `readVerificationStatus` parses frontmatter only, anchored at byte 0, and is total over its input space, returning `missing`, `unknown` and `stale` sentinels. execute-phase.md, verify-work.md and progress.md all read this same artifact through that query already; quick was the remaining second mechanism. Route quick's read through it and add an explicit terminal arm. Three details a naive swap misses: - The step file carries the runtime shim bootstrap itself. Step files are read and executed as their own units, so quick.md's bootstrap does not reach here. Copied byte-identically from gsd-core/workflows/_runtime-launcher.snippet.sh, the source sync-runtime-launcher.cjs generates every workflow's copy from. Without it the call resolves to nothing, 2>/dev/null swallows the error, and the fix degrades to a permanently-taken recovery arm. - No jq. `--pick status` returns the bare value. Per #2589 a `| jq -r` pipe yields an EMPTY variable with no diagnostic wherever jq is absent — the Windows/Git-Bash default — which here would route a passing verification into the recovery arm, strictly worse than the grep being replaced. - $VERIFICATION_STATUS is a DISPLAY string ("Verified" / "Needs Review" / "Gaps") consumed at quick.md:619 and quick.md:684, not the raw status. The raw value lands in $STATUS and the new arm sets both, so the failure path does not emit an empty index-table cell. next_action / next_command are deliberately not surfaced. readVerificationStatus discovers and parses shape-agnostically, which is what makes the status half correct for ${QUICK_DIR}; but it also reads the directory basename as a phase token to build those commands, and a quick dir is `${quick_id}-${slug}` with a date-derived quick_id — so the projection carries the date as a phase argument. Quick supplies its own recovery actions instead. Adds tests/fix-3174-quick-verification-status-read.test.cjs under `allow-test-rule: source-text-is-the-product` (CONTRIBUTING.md's exception matrix; the pattern tests/verify-work-auto-transition.test.cjs already uses for verify-work's status-query ordering). It pins five properties: the query replaces the grep, the bootstrap precedes the call, the bootstrap matches the canonical launcher snippet, the status-read fence is jq-free, and the terminal arm names all three sentinels and sets the display string. Verified as a negative control against pre-fix next: 0/5 pass there, 5/5 here. |
||
|
|
6ac4d2e6ab |
fix(#3156): bound the cold-require probe the rebase turned into a violation
Post-rebase validity finding, not a new defect. The base range added local/no-unbounded-spawn (#3143, |
||
|
|
af39f13be2 |
test(#3156): pin the ambient-HOME leak the scrub set cannot reach
Two halves, and the second is the one that matters.
The contract half asserts installSpawnEnv() and installerEnv() both replace the
ambient HOME, keep USERPROFILE tracking it (os.homedir() reads that one on
Windows), and still let an explicit override win.
The behavioural half drives the REAL installer against a REAL ambient HOME: it
points process.env.HOME at a canary, runs `install.js --cursor --local`, and
asserts the canary gains no .gsd. No assertion about the scrub set can stand in
for this, because writeNonClaudeDefaults() resolves through os.homedir(), which
reads no GSD variable -- a set-membership test would pass with the bug fully
live.
Negative-controlled against the pre-fix tree rather than assumed: reverting
installerEnv() to `{ ...process.env, ...overrides }` fails BOTH halves, the
behavioural one reporting the actual artifact
("the installer wrote GSD's user store into the ambient HOME: defaults.json").
|
||
|
|
35df0891af |
fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates $HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard is right — this is #2665's class arriving through the one door the scrub set is structurally unable to close. bin/install.js writeNonClaudeDefaults() (#2834) writes path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime. os.homedir() consults NO GSD variable, so: - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is derived; and - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back to exactly that homedir(). Only a sandboxed HOME contains it, and HOME is deliberately excluded from TEST_ENV_BASE because blanking it would break far more than it fixed. So the containment belongs per-spawn, which is the discipline the suite already applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment naming this exact hazard for the --codex spawn, while the parameterized --${runtime} spawn 130 lines below it does not. Instance fixed, class open. Rather than add a fifth hand-synced env shape to a PR whose subject is that hand-synced copies drift, this adds ONE export -- installSpawnEnv() in tests/helpers.cjs -- and routes every raw installer spawn through it, including the shared tests/helpers/install-shared.cjs installerEnv(), which every install suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are unaffected: overrides spread last. Census (measured, not reasoned): of the 119 test files that spawn bin/install.js, exactly four wrote into a sandboxed $HOME before this commit -- install, copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu went red on one file while the macOS full lanes went red on four. Attribution: the leak reproduces unchanged at upstream/next itself, so the defect is base-owned and pre-existing; only the detector is new. The guard found a real leak on next within one run. No new failures: the surviving names under a sandboxed HOME (folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too, and base additionally fails folded:bug-3288-model-catalog-install-path, which this tree does not. |
||
|
|
fae0c6ae1a |
fix(#2665): stop watching shared ground, and derive the artifact prefix too
The previous commit widened the guard's watch set and claimed the enumeration was complete. Re-running the pre-push adversarial gate on that commit -- which I should have done before pushing it, and did not -- refuted the claim on four counts. All four were real. 1. FALSE POSITIVES, which is the worse polarity. `hooks/lib`, `hooks/package.json`, `scripts/lib` and `scripts/changeset` were watched WHOLESALE. The installer preserves foreign files in every one of them -- it removes the CommonJS marker only on an exact content match, because "a user-authored package.json is never deleted" -- so a user editing their own helper mid-suite tripped the guard. A driven probe produced four violations from touching only user-owned files. Watching shared ground is exactly what the module's SCOPE note refuses: a guard that cries wolf gets switched off, and then catches nothing at all. Now only exact GSD filenames inside those dirs are watched, and a test asserts foreign edits stay silent. 2. THE PREFIX WAS HARDCODED, which is this PR's own defect one level down. Each artifactLayout declares its OWN prefix, and kimi's `kimi-agents` layout declares `gsd` with no hyphen, writing `agents/gsd.yaml` and `agents/gsd.md`. A fixed `gsd-` scan is structurally blind to both, as it is to pi's `extensions/gsd.js`. The prefix is now derived per parent, as a SET -- the same destSubpath carries different prefixes across runtimes (`agents` appears with both `gsd` and `gsd-`). `extensions` joins the non-registry parents; pi declares no artifactLayout at all, so no registry walk could find it. 3. THE ENTRY BOUND FAILED OPEN on a non-finite limit: `Math.max(0, NaN)` is NaN, and every budget comparison against NaN is false, so the walk was unbounded -- the single thing the constant exists to prevent. Clamped with Number.isFinite. The walk also kept invoking itself for every remaining sibling after the budget was gone; it now returns. 4. THE RESIDUAL LIST WAS WRONG AGAIN. `agents/subagents/**` (kimi stages under an unprefixed intermediate dir), the loose capability generators, and the `extensions`/`plugins` CommonJS markers are all unwatched and were unnamed. They are named now, and the four shared dirs are recorded as DELIBERATELY not watched -- a different thing from missed. Each fix is negative-controlled and each control fires. The NaN control did not fire on its first form: the test asserted `truncated: false`, which the broken code also produces on a small tree, so it discriminated nothing. Repaired with a NaN perTarget against a small finite ceiling, where the two behaviours differ. |
||
|
|
766480967e |
fix(#2665): derive the guard's artifact targets, and close the fallback hole in the extras
A pre-push adversarial review refuted this round's own completeness claim, and it was right on all three counts. Fixes, in the order they matter: 1. The watch enumeration was still a hand-list, and it was measurably incomplete. It missed kilo's SINGULAR `command/`, hermes' `skills/gsd` (a whole directory whose name carries no `gsd-` prefix, so no prefix rule could ever reach it), `plugins/gsd-core.js`, and the unprefixed subtrees the installer fills -- `hooks/lib`, `hooks/package.json`, `scripts/lib`, `scripts/changeset`. The parents are now DERIVED from the capability registry's own artifactLayout.global destSubpath values, exactly as TEST_ENV_BASE derives its keys, plus a named list for the non-registry paths the installer writes directly. A capability declaring a new destination now extends the watch set in the commit that declares it. Scope note: only `global` is walked -- `workflows` is declared LOCAL-only (windsurf) and is not a config-root parent. 2. resolveExtraWatchTargets carried the identical ambient-only defect that Blocker 3 closed one function over: it resolved $GSD_HOME/.gsd and each kimi descriptor from the ambient env alone, so a child that BLANKED those vars wrote to the HOME-derived fallback while the guard watched the override. Both legs are now unioned, matching resolveLiveConfigRoots. 3. The order-independence claim for the scan budget was too strong. It holds BELOW the global ceiling; once MAX_TOTAL_ENTRIES is exhausted, which targets get curtailed still depends on iteration order -- inherent to any shared aggregate bound. The residual is now named in the docblock and the test title says which regime it pins, instead of asserting the general claim. Negative limits are clamped at 0 so an injected value cannot masquerade as a scan bound. The module's KNOWN GAP now names its remaining residuals (the loose generator scripts, the kimi native-root hook bundle) rather than implying completeness -- an unqualified claim here just invites the same refutation next round. Both under-watch, which fails quiet. Reverting the derivation fails two tests; reverting the fallback leg fails a third. |
||
|
|
104fc76f70 |
fix(#2665): watch the hook bundle and the install markers the census found
Self-found by re-deriving the guard-shape census against bin/install.js's own
write sites, not by a review finding. Three artifacts a global install writes
into a live config ROOT were watched by nothing:
hooks/gsd-check-update.js, hooks/gsd-context-monitor.js,
hooks/gsd-update-banner.js -- `hooks` was absent from GSD_PREFIXED_PARENTS
.gsd-source, .gsd-profile -- absent from GSD_OWNED_ENTRIES, and an
exact-name list does not match a dot-prefixed
name via the `gsd-` prefix rule
This is the SAME shape as the leak that motivated the prefixed-parent scan in
round 1 -- a gsd-prefixed child under a parent nobody had listed -- one parent
over. That it recurred is the argument for re-deriving this list from the
installer each round instead of trusting it: the enumeration is the weak point
of an enumerate-and-block mechanism, and it does not announce when it falls
behind.
Ownership is unchanged, only coverage: `hooks/` is shared with the host agent,
so only `gsd-`-prefixed children are watched. A test asserts a host-owned
hook is still ignored, because widening the parent list must not widen
ownership -- a guard that flags the host's own files gets switched off, and
then catches nothing at all.
Reverting the widening fails the new test.
|
||
|
|
e31f706ceb |
docs(#2665): document the two live-config-guard env vars
The changeset for this PR is typed `Added`, and CONTRIBUTING requires a docs/ change for that type. The only docs/ file in the diff was CONTEXT-INDEX.json -- a GENERATED index -- so the Docs Required gate passed while no human-readable documentation existed for either new variable. A gate satisfied by a generated artifact is satisfied vacuously. docs/TESTING-SUITES.md now carries a section on the guard: what it watches and why it is ownership-scoped rather than whole-root, the two env vars in a table, why the default is report-only and what the promotion condition is, and what each violation label means (including that UNVERIFIED is not clean). GSD_SKIP_LIVE_CONFIG_GUARD is named explicitly because it is a bypass on a safety check. An undocumented bypass is one people eventually set without knowing what they turned off. A test asserts both variables appear in that doc -- checked as permitted by local/no-source-grep before writing it, rather than assumed forbidden. It fails when the section is removed, so the doc cannot rot back to the state the review found. Addresses review finding: Major 4. |
||
|
|
f0ef9063d5 |
test(#2665): restore the three agent-skills tests this PR deleted
Commit 2bed9fd8 ("replace the hand-synced TEST_ENV_BASE copies with the
canonical import") also removed markLocalGsdInstall and three behavioural tests
from tests/agent-skills.test.cjs -- 79 lines, no replacement, and no mention in
the commit message or the PR body:
- unconfigured Codex reads its local companion agent from a descendant cwd
- workstream runtime selects the local Codex companion when root config differs
- unconfigured Claude remains empty when a local Codex companion exists
RULESET.TESTS.delete-bad-tests permits deleting a bad test only when it is
replaced with compliant tests in the same PR. Nothing was replaced, and these
were not bad tests -- they were collateral in a mechanical edit. A silent net
loss of behavioural coverage inside a PR whose subject is test hygiene is the
one thing that should not pass here, and the reviewer was right to block on it.
Restored verbatim. They need no adaptation to the canonical TEST_ENV_BASE
import: they pass their env explicitly, and an explicit env still spreads last
over the base. The file goes 80 -> 83 tests, all green.
Non-vacuity checked rather than assumed: dropping the local-install marker the
first two depend on fails both. The third is a negative assertion and correctly
stays green, which is why it is named here rather than counted as covered.
Addresses review finding: Blocker 1.
|
||
|
|
e4f79c32b0 |
fix(#2665): wire the fourth suite lane, and derive the lane list instead of naming it
qa-loop-walk runs `npm run test:qa`, which is `run-tests.cjs --suite qa` -- so it runs the live-config guard like every other suite lane, and it set no GSD_STRICT_LIVE_CONFIG_GUARD. A leak of exactly the class this PR closes would have printed a warning there and left the lane green. The test that is supposed to prove the guard is wired everywhere could not detect that, because its job list was three literals (`test`, `test-full`, `test-inert`). A hand-list certifying its own completeness is the defect this whole PR is about, reproduced inside the test guarding the fix -- so the list is now DERIVED from the workflow: every job with a step reaching run-tests.cjs, directly or through an npm script resolved transitively through package.json. The indirection is the load-bearing half; a grep for the filename alone is what made qa-loop-walk invisible. The derivation asserts a floor (>= 4 jobs) before ruling on any of them, so a selector that silently matched nothing fails loudly instead of passing vacuously. Windows lanes keep their carve-out, keyed on whether the job's matrix mentions windows rather than on the job's name. Negative-controlled: un-wiring qa-loop-walk fails the new test. The literal version passed with that lane unwired, which is how it shipped. Addresses review finding: Major 5. |
||
|
|
4bc6b0a2a3 |
fix(#2665): defer the built-lib require so an unbuilt tree fails one test, not all
tests/helpers.cjs required gsd-core/bin/lib at module scope to derive the
config-location scrub set. That lib is BUILT, so on an unbuilt tree the require
threw inside `require('./helpers.cjs')` -- before a single test() had registered
-- turning one missing `npm run build:lib` into a whole-suite crash with no
message naming the remedy. This is the file ~370 test files import, so the blast
radius is the suite. `npm test` builds via its pretest hook; the shape that
reaches this is a direct `node --test` invocation, which is exactly what a
contributor reaches for when running one file.
The require is now memoized behind builtLib(), and the two derived exports
(TEST_ENV_BASE, CONFIG_LOCATION_ENV_KEYS) are enumerable lazy getters, so
destructuring and Object.keys() behave as before. Reading either is what forces
the build; a test file that needs neither now imports cleanly. When the build IS
missing, the error names `npm run build:lib` instead of surfacing a bare
MODULE_NOT_FOUND.
Verified by a cold-child probe rather than by inspection -- this process has
already loaded everything, so an in-process assertion would pass vacuously. The
probe checks require.cache before and after touching TEST_ENV_BASE, and fails
when the require is moved back to module scope.
Addresses review finding: Major 7.
|
||
|
|
4eb29b9741 |
test(#2665): pin the MAX_DEPTH boundary on both sides, not just above it
RULESET.TESTS.boundary-coverage asks for {limit-1, limit, limit+1}. The depth
bound was exercised only at limit+2, which pins neither side of the edge: an
off-by-one that truncated a tree sitting exactly AT MAX_DEPTH would have passed,
and a truncation is not a cosmetic miss here -- it reports `unverified`, which
in strict mode fails the run.
Negative-controlled by weakening the guard to `depth >= MAX_DEPTH`: the new
limit case fails, where the previous single limit+2 assertion did not.
Addresses review finding: Minor 8.
|
||
|
|
f054c85fb0 |
fix(#2665): budget the scan per target, so order stops deciding the verdict
MAX_ENTRIES was a single running budget threaded across every watch target. One large early target exhausted it, and every target scanned afterwards reported truncated -> `unverified` -- which under GSD_STRICT_LIVE_CONFIG_GUARD=1 is a failed run. The guard's verdict therefore depended on directory iteration order and on unrelated local state, neither of which says anything about whether the suite leaked. Each target now draws a fresh allotment, so a pathological tree truncates itself and nothing else. MAX_TOTAL_ENTRIES keeps the aggregate bounded -- which is what the single budget was actually for -- and when that ceiling engages, the targets it curtails are still reported `unverified` rather than attested clean. The limits are injectable so the boundary is testable without materialising 20000 entries, matching the `deps` seam the resolvers already use. Two of the three new tests fail when the shared budget is restored; the third asserts the retained global ceiling, which is deliberately unchanged behaviour. Addresses review finding: Major 6. |
||
|
|
e82a15a852 |
fix(#2665): watch the fallback root a scrubbing child actually resolves to
resolveLiveConfigRoots resolves what THIS process sees, and getGlobalConfigDir is env-first -- so with an ambient CLAUDE_CONFIG_DIR the guard watched that path. A spawned child does not see it: TEST_ENV_BASE blanks the config-location vars precisely so the child cannot follow them, and a blanked var is falsy, so the child resolves its HOME-derived root instead. A child that blanks the var and does NOT also sandbox HOME therefore writes into the developer's real ~/.claude, which the guard was not watching. That is this PR's own escape route, taken one process deeper -- and the guard is the artifact that is supposed to make it loud. Both resolutions are now unioned: the ambient one, and the fallback one obtained by handing the REAL descriptor resolver an EMPTY env. Deriving it that way is deliberate -- a hand-listed copy of the scrub set inside the guard is a second list to drift, which is the defect this PR spent three rounds closing one layer up. grok resolves through a hardcoded branch rather than a descriptor, so its fallback is stated explicitly for the same reason it is named in the ambient loop. Addresses review finding: Blocker 3. |
||
|
|
6a1fbf96fd |
fix(#2665): let the guard see deletions, in both shapes it can take
diffLiveConfig walked `after` alone, so it had no branch for a path that
existed before the run and does not after. A test run that DELETES a file from
the developer's real config dir passed the guard silently -- the least
recoverable case in the threat model this guard exists to cover.
The review named the missing `pre.exists && !post.exists` branch. That branch is
necessary and not sufficient: deletion arrives in two shapes and it reaches only
one of them.
- A FIXED owned entry (GSD_OWNED_ENTRIES x roots, plus every extra target) is
recorded at both ends whether it exists or not, so a deletion reads
{exists:true} -> {exists:false}. This is the shape the named branch fixes.
- A gsd-prefixed child is DISCOVERED by readdirSync, so a deleted one is
absent from `after` entirely and never enters an after-keyed loop at all.
The named branch is unreachable for it.
So the walk is now over the UNION of both key sets, with the explicit branch for
the first shape and an `!post` branch for the second. Both are covered by a
test, and reverting the fix fails both -- the prefixed-child test is the one
that would still fail with only the prescribed branch in place.
Addresses review finding: Blocker 2.
|
||
|
|
ecea537194 |
docs(#2665): the guard watches config.toml but GSD also writes <root>/hooks/ there
Found pre-push by this round's third adversarial review pass. Not a rebase regression — round 3 shipped it and #2755 doubled it. resolveExtraWatchTargets watches one config.toml per non-registry descriptor, and its comment asserted "GSD writes ONE named file into these third-party roots". That is false: bin/install.js also calls installSharedHooksBundle on the same root, populating <root>/hooks/ with GSD's hook scripts and a CommonJS marker. So a suite-produced leak of a hook bundle into a developer's real ~/.kimi or ~/.kimi-code passes this guard silently — #2665's own hazard, in #2665's own safety net. Behaviour is deliberately unchanged and the gap is disclosed instead. Closing it is a layout decision rather than one more path, for the same reason getGlobalSkillsBase is already a deliberate non-target: the snapshot applies the config-root layout beneath every root it is given, and these roots are not ours. Happy to fix it here or take it as a separate issue — the maintainer's call. The enumeration-relative test could not have caught this: it asserts one target PER DESCRIPTOR and nothing about whether one per descriptor is enough, because its expectation is derived from the same array it checks. That is exactly the scope boundary round-2 Nit 7 asked to be marked, biting one layer up from where it was marked; the test now says so. 479979c4's message says "there are three" — that is three WATCHED targets, not a count of write surfaces. The hooks bundle is a fourth, and unwatched. lint:ci rc=0; tests/live-config-guard.test.cjs 24/24. Comments and catalog only. |
||
|
|
12cfd27f53 |
docs(#2665): the guard's own comments still described one kimi home, not two
Same drift as the CONTEXT.md seams, one layer over: #2755 took NON_REGISTRY_CONFIG_HOME_DESCRIPTORS from one entry to two, and five comments across three files were left describing the one-entry world — "two live write surfaces", "today's only entry", "today's single entry", and a <kimi>/config.toml bullet naming only Kimi CLI's KIMI_SHARE_DIR. The sharpest one was a wrong pointer rather than a stale count: run-tests.cjs cited "scripts/lib/live-config-guard.cjs" for why the scope is narrow. That path does not exist, and it names the one directory this module is deliberately NOT in — the installer copies scripts/lib/ to users wholesale while uninstall removes only an allowlist, which is the whole reason the guard lives one level up. A reader following that pointer would have concluded the opposite of the decision. Comments only; no behaviour change. lint:ci rc=0, tests/live-config-guard.test.cjs 24/24, tests/run-tests-harness.test.cjs 138/138. |
||
|
|
c95b817cb6 |
fix(#2665): scrub GSD_ALLOW_SYMLINKED_DEST — a write-escape permission
Found by the pre-publication claim audit of this round's response comment, which refuted the sentence "none of the unscrubbed env reads names a write destination" on the grounds that naming a path is not the same property as influencing where writes land. GSD_ALLOW_SYMLINKED_DEST is boolean and names no path, so every rung of the derivation is structurally incapable of reaching it: not a registry configHome, not descriptor-shaped, not one of GSD's own location vars. It is still a #2665 leak vector. install-engine.cts reads it env-first (:214) and threads it as allowOptInFollow into the symlink-escape guard at four call sites, each gating a write (:361/:367, :416/:424, :785/:790, :927/:932). That guard is what stops a write leaving the install root, so an ambient =1 disarms it for the whole suite — the #2665 hazard arriving through a permission rather than a path. Added as its own named family (WRITE_ESCAPE_PERMISSION_ENV_KEYS) rather than folded into a location rung, for the same reason GSD_HOME got its own family in round 3: the list should not misdescribe what its members are. Blanking is fail-safe in the only direction that matters — '' is neither '1' nor 'true', so a blanked value makes the guard stricter, never looser. That asymmetry is what licenses scrubbing it wholesale rather than reasoning about each call site. The guard test names the variable literally rather than iterating the family constant: a test that asserts over the constant shrinks its own expectation when the family is emptied, which is the enumeration-relative failure that let the kimi-code descriptor go unwatched earlier in this same round. #2393's opt-in suite is unaffected (99/100, 0 fail): it sets process.env directly in-process and never routes through scrubConfigLocationEnv, and an explicit env argument still wins over TEST_ENV_BASE in childEnv. |
||
|
|
d1c8b32689 |
fix(#2665): watch kimi-code's config.toml, and pin it by name
The rebase onto next brought in #2755, which added a SECOND Kimi config home — kimi-code's `~/.kimi-code`, overridden by KIMI_CODE_HOME — declared as an inline object literal inside resolveKimiHooksTomlDir's body. That is the resolvable-but-not-enumerable shape round 3 hoisted KIMI_SHARE_DIR out of, so the hoist is extended to cover both descriptors rather than reverting #2755's parameterization. The scrub set was already complete: KIMI_CODE_HOME is declared in capabilities/kimi-code/capability.json, so the registry rung covered it and CONFIG_LOCATION_ENV_KEYS is 28 keys both before and after the rebase. What was NOT covered is the guard — resolveExtraWatchTargets iterates NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, so with only one entry it watched Kimi CLI's config.toml and never Kimi Code's. Targets go 2 -> 3. The existing 'extra targets are DERIVED from the descriptor array' test cannot catch this: it builds its expectation FROM the array, so removing an entry shrinks the expectation with it. Verified — with the kimi-code descriptor removed that test still passes while the new named test fails. This is the enumeration-relative scope boundary the suite already documents one layer down, biting one layer up. Also rewrites NON_REGISTRY_OWNED_FILE's docblock, which asserted "today's only such descriptor is kimi's ~/.kimi". There are now two, and its named residual is load-bearing rather than vacuous. |
||
|
|
652b99b71b |
test(#2665): reversion guards for all three round-4 fixes
None of the round-4 fixes had a test that fails on reversion: re-shipping the test-instrumentation chain passes #2858 (everything ships, so every require resolves), dropping the strict env from test.yml demotes the guard to report-only with nothing red, and the skillsHome derivation tests are enumeration-relative over declarations that are all empty today. One guard each, every one negative-controlled against its reverted fix (fails pre-fix, passes post-fix): - packaging-shipped-scripts-require-only-shipped.test.cjs asserts the four chain files are absent from the npm pack file list (reuses the tarball set the #2858 gate already resolves — no second npm pack). - live-config-guard.test.cjs asserts all three test jobs wire GSD_STRICT_LIVE_CONFIG_GUARD, matching the WHOLE expression anchored — a prefix match accepted both a Windows-silently-strict tail and a malformed one. - helpers-process-isolation.test.cjs cold-requires helpers.cjs in a child with sentinel skillsHome env vars injected into both enumerations, so the walk itself is under test rather than today's empty declarations. |
||
|
|
7b8c36f904 |
fix(#2665): walk skillsHome.env on both descriptor rungs of the scrub derivation
Review round 4, Minor 3. A configHome descriptor can nest a second, independently-resolved descriptor (skillsHome -> resolveSkillsBaseFromDescriptor) carrying its own env array, and the derivation walked configHome.env alone — the identical walk-one-field gap-shape rounds 2-3 closed for the registry and the non-registry set. Inert today (only kilo declares skillsHome, with env: []), closed before it is live rather than after. The guard's root enumeration deliberately does NOT gain the skills base: getGlobalSkillsBase returns a skills directory (codex: ~/.agents/skills), not a config root, and the snapshot applies the config-root layout beneath every root — adding it false-positives on <skillsBase>/gsd-core while missing a real <skillsBase>/gsd-help write (found by this round's pre-push adversarial review). Watching skills bases needs its own layout, like resolveExtraWatchTargets; a comment in resolveLiveConfigRoots records the non-action. New derivation test asserts both skillsHome rungs land in TEST_ENV_BASE, with an anti-vacuity check that at least one runtime actually declares the field. |
||
|
|
706bd2ab4e |
refactor(#2665): derive the guard's non-root targets from the descriptor array too
Follow-up to 38c9395d, found while fact-checking the round-3 response rather than by a test. That commit made TEST_ENV_BASE derive its keys from NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, but had the guard call resolveKimiHooksTomlDir directly. Both halves covered kimi, so nothing was broken — but only one of them would pick up a SECOND descriptor. That is the same partial-enumeration defect that put KIMI_SHARE_DIR outside the scrub set, reintroduced one layer over, in the very commit that closed it. resolveExtraWatchTargets now iterates the array and resolves each descriptor through resolveConfigHomeFromDescriptor, so the scrub set and the guard derive from one source and cannot drift apart. Verified: a synthetic second descriptor is picked up automatically (it was not before); kimi's target is unchanged on both the default (~/.kimi/config.toml) and KIMI_SHARE_DIR override paths. The new test asserts one target per descriptor plus the store root. The COUNT is the load-bearing half — every per-descriptor assertion passes vacuously today with a single entry, so only the count fails when the array grows and the guard does not follow. NAMED RESIDUAL, documented at NON_REGISTRY_OWNED_FILE: this assumes every non-registry descriptor is written the same way (config.toml). A descriptor whose owned file differs needs a per-descriptor mapping. It fails toward under-watching rather than false positives, so it is called out rather than left to be discovered. |
||
|
|
fec42e9a7d |
test(#2665): cover the widened derivation, and mark what these tests cannot prove
Two tests for round 3's change (round 2 Blockers 1 and 2): KIMI_SHARE_DIR must arrive via NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than a literal, and every GSD_LOCATION_ENV_KEYS entry must be blanked. Both assert a floor on their source first, so a renamed export fails loudly instead of passing vacuously. Negative-controlled against the registry-only derivation: both fail there, and only those two. And the Nit, which is the more useful half. This block asserts that TEST_ENV_BASE is not narrower than the enumerations it derives from. It cannot prove those enumerations are complete — a var no enumeration carries is invisible to every test here, and they stay green. That is exactly how round 2 found GSD_HOME and KIMI_SHARE_DIR while this block was fully green: one belonged to no enumeration at all, the other sat inside a function body where nothing could enumerate it. So the scope boundary is now written down at the top of the block, naming where the completeness question is actually answered — a source census re-derived each round, and live-config-guard.cjs observing real writes at runtime — so that a green run here is not misread as "the set is exhaustive." Deliberately NOT added: an assertion per reviewer-named variable. That is the hand-maintained list wearing a test's clothes, and it fails the same way. |
||
|
|
1f6d827e48 |
fix(#2665): scrub config-location env in the #2624 in-process install block
Self-found during the rebase onto next, not from the review.
The base range added `describe('#2624 .gsd-source marker is rewritten before
staging reads it')`, which calls the real `install(true, 'claude')` IN-PROCESS
and sandboxes HOME/USERPROFILE/GSD_EXPLICIT_CONFIG_DIR — but not
CLAUDE_CONFIG_DIR. That is exactly the Blocker-1 shape this PR exists to close,
reintroduced in new code written after the round-1 review.
Measured on the rebased tree, same file, same commit:
CLAUDE_CONFIG_DIR unset -> 50/50 pass
CLAUDE_CONFIG_DIR set -> 47/50, and a COMPLETE global install lands in it
(gsd-core/, agents/, skills/, hooks/, scripts/,
gsd-file-manifest.json, gsd-install-state.json,
.gsd-source, .gsd-profile)
The three failures are the honest symptom rather than the problem: the install
goes to the ambient config dir, so the assertions look for a marker under
tmpRoot that was never written there.
With scrubConfigLocationEnv() wired into the block's beforeEach/afterEach —
the same pattern the sibling block at :753 already uses — the file is 50/50
under BOTH conditions and leaks zero entries.
Worth stating plainly: CI cannot catch this class, since CI never has these
vars set. It surfaced here only because the rebase brought the base's new tests
under an ambient CLAUDE_CONFIG_DIR, which is the condition #2665's own
acceptance criterion runs under.
|
||
|
|
a294ec2a2b |
test(#2665): widen the hermeticity guard to its two blind surfaces, and cover its budget
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.
Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:
$GSD_HOME/.gsd — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
this root is exclusively ours, so the shared-root
false-positive trap the module documents does not apply.
<kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
file GSD writes is watched, never the root.
That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.
Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.
Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.
Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
|
||
|
|
3c580b77dc |
fix(#2665): derive the second config-location family instead of hand-adding it
Review round 2 named GSD_HOME and KIMI_SHARE_DIR as missing from the derived scrub set. Both premises confirmed; the prescribed remedy is not adopted verbatim, because adding two more literals to a four-item hand list is the pattern that reopened this bug three times. The census the derivation generalizes over was partial, so the census is what widens. Two structural gaps, both closed at the source: 1. KIMI_SHARE_DIR lived inside resolveKimiHooksTomlDir's body as an inline descriptor, resolvable but not ENUMERABLE. Hoisted to an exported KIMI_HOOKS_TOML_DESCRIPTOR and collected in NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, which TEST_ENV_BASE now derives from. kimi is the sharp case: it owns TWO config homes (KIMI_CONFIG_DIR, already registry-visible, and this one), so a registry-only derivation looks complete and is not. 2. GSD_HOME is a different FAMILY, not a missing registry entry. The registry describes where third-party runtimes keep config; GSD_HOME decides where GSD keeps its own user-owned state ($GSD_HOME/.gsd/ — consent.json, defaults.json, capability overlays), read env-first ahead of os.homedir() by capability-loader, capability-consent, capability-state, capability-writer, config-loader, install-profiles and bin/install.js. Named as GSD_LOCATION_ENV_KEYS rather than folded into the descriptor array, since it does not resolve through resolveConfigHomeFromDescriptor. GSD_AGENTS_DIR joins the same family (round 2, Minor): env-first and unconditional in getAgentsDir, misdirecting a read rather than a write. A census of every env-first first-party location var — the guard-shape question this PR owes each round — now yields exactly one remaining unguarded name, GSD_MODEL_CATALOG, and it is dead by precedence: the co-located candidate is index 0 and the loop breaks on first success, so the env var can only win on a tree that is already broken, and it redirects a read even then. Derived set: 39 -> 42 keys. resolveKimiHooksTomlDir behaviour unchanged on both the default and the KIMI_SHARE_DIR override path. |
||
|
|
e2eed1c58a |
test(#2665): ship the hermeticity guard at report level, not fatal
Its first CI run found PRE-EXISTING leaks on the Windows lane — C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all 1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows, and ~190 test sites across 31 files sandbox HOME alone, so the suite has been installing GSD into the runner's real home directory invisibly. That is exactly the class the guard exists to surface, and exactly the class this PR's review said CI could never catch. It is also a different defect from the one #2665 closes, and too large to fold in here. A brand-new gate that immediately reds an unrelated lane gets bypassed or reverted rather than obeyed, so the guard reports by default and fails only under GSD_STRICT_LIVE_CONFIG_GUARD=1. This is the repo's own established ratchet, not a hedge: the local/no-source-grep ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep (ADR 452). Promote this the same way once the USERPROFILE sweep lands. |
||
|
|
5863351281 |
test(#2665): separator-safe containment in the regression assertion
startsWith(ambientConfigDir) also matches a sibling like <tmp>/ambient-live-config-2, so it can report a leak that did not happen. Use path.relative and check for '..' or an absolute result, the repo's usual shape. The readdirSync assertion already carried the test, so this is cosmetic. Addresses review finding: Nit 9. |
||
|
|
771980b661 |
test(#2665): restore fallback-branch coverage in the #2003 regression
This PR fixed the test's real defect -- it compared the child's answer against the PARENT process's getGlobalConfigDir(), two different environments, agreeing only because the child inherited the developer's ambient CLAUDE_CONFIG_DIR -- but fixed it by INJECTING CLAUDE_CONFIG_DIR, which moved the test onto the env-first branch and silently dropped the only coverage #2003 had of the home-derived fallback. Sandbox HOME/USERPROFILE and leave the config vars blank instead: the expectation stays test-controlled AND the branch under test is unchanged. Drops the notStrictEqual against the codex dir. It could not fail whenever the strictEqual on the line above passed. Addresses review finding: Minor 7. |
||
|
|
a4efa4deda |
test(#2665): cover the derivation and the in-process scrub
The single regression test exercised CLAUDE_CONFIG_DIR only, so deleting GSD_RUNTIME or CODEX_HOME from any literal broke nothing -- a mutation of 24 of the 27 added key-value pairs survived. Four tests here: every configHome env var the registry declares is scrubbed; every scrubbed key is blanked rather than merely present; the four non-registry vars are named explicitly so deleting one is a failure rather than a silent narrowing; and scrubConfigLocationEnv round-trips both a set and an unset var (restoring an originally-unset var as '' would itself be a leak). Both derivation tests assert a floor on the registry first, so a renamed registry shape fails loudly instead of making the assertions vacuously true. Negative-controlled against the hand-written 3-key list this PR shipped: the parity test fails there and names all 19 missing vars. Addresses review findings: Major 6, Minor 8. |
||
|
|
a02462e050 |
test(#2665): fail the suite when it writes into a live config dir
The recurrence guard, and #2665's own "Optional hardening". This class is silent by construction: TEST_ENV_BASE cannot see an in-process caller, and CI cannot see the class at all because CI never has these env vars set. It damages the developer's machine and reports nothing -- which is how two prior authors each diagnosed it and fixed only the instance in front of them. run-tests.cjs snapshots GSD's install footprint in every live runtime config dir before the suite and re-checks it after, failing the run on a create or a modify. Roots come from the product's own getGlobalConfigDir, so the guard watches wherever the product actually points, including through an ambient var. Scope is ownership-based, not whole-root: the top-level install footprint plus gsd-prefixed children of dirs GSD shares with the host agent. A config root like ~/.claude is shared, and watching it wholesale would false-positive on the host's own history.jsonl or settings.json -- a guard that cries wolf gets disabled, and then catches nothing. The prefix test is load-bearing: the first version watched only the three top-level entries and MISSED a real leak into skills/gsd-*. It earned its place immediately -- it is what found the fifth in-process leak in runtime-artifact-layout.test.cjs, which no amount of reading the review would have surfaced. Known gap documented in the module: a write to a file GSD does not own is out of scope by construction. Lives in scripts/, deliberately NOT scripts/lib/ -- the installer copies that dir into every user's config dir wholesale while uninstall removes only an allowlist, so a test-only module there would ship to users and survive uninstall. Addresses review finding: Minor 8. |
||
|
|
36f4f9479a |
fix(#2665): scrub config-location env on raw spawns that sandbox only HOME
Same class as the in-process leak, on the child-spawn side. These call spawnSync
directly rather than through runGsdTools, so TEST_ENV_BASE never applies to them
and an ambient config-location var survives into the child.
- install-runtime-artifacts: the dev-preferences writer resolved env-first and
wrote SKILL.md into the live config dir instead of its tmp HOME. This also
fixes a test that FAILS today on any machine with CLAUDE_CONFIG_DIR set.
- issue-766: the `claude` CLI is third-party and bootstraps its own config into
whatever CLAUDE_CONFIG_DIR names, so even a bare --version probe wrote there.
It gets an explicit throwaway dir rather than a blank value -- GSD's resolvers
treat '' as falsy and fall back to the home dir, but a third-party binary
offers no such guarantee (blanking it produced a stray backups/ in the repo
root).
These two were the last writers standing between the suite and #2665's stated
acceptance criterion.
|
||
|
|
466db2f1a3 |
fix(#2665): scrub config-location env on the parent for in-process install()
The blocker the review said decides this PR. These tests call the real installer IN-PROCESS with only HOME/USERPROFILE sandboxed. getGlobalConfigDir is env-first, so an ambient CLAUDE_CONFIG_DIR beats the sandbox and a complete global install -- agents/, commands/, skills/, gsd-core/, manifest, settings -- lands in the developer's live config dir. No child-env scrub can reach it; only clearing the parent's env can. The review named two describe blocks in install.test.cjs. A sweep for the shape found four there (bug #3571 and bug #3288 each appear twice in the file), and the post-suite guard added later in this series found a fifth in runtime-artifact-layout.test.cjs, which the review did not name. All five now save/clear/restore via scrubConfigLocationEnv(). Verified: codebuddy-install, cline-install and codex-config already guard their own config-location var around in-process install(), so the class is closed. Addresses review finding: Blocker 1. |
||
|
|
bc5c362610 |
fix(#2665): replace the hand-synced TEST_ENV_BASE copies with the canonical import
Eight declarations were kept in sync by hand with no parity assertion. The drift was already in the tree: api-coverage-gate-e2e and representative-corpus declared TERM_SESSION, but the real variable is TERM_SESSION_ID, so both scrubbed nothing for that slot and leaked TERM_SESSION_ID into every child. Deleting the copies removes the dead key with them -- there is no longer a second place to get wrong. run-tests-harness keeps its local session-identity literal: that helper mirrors the production runGsdTools to prove its CONTRACT, so it must not re-import what it is testing. The config-LOCATION keys are a safety scrub rather than part of that contract, so it spreads the canonical derived set and keeps the rest local. Addresses review findings: Blocker 2, Blocker 3. |
||
|
|
0f97a26f76 |
fix(#2665): derive the config-location scrub set from the capability registry
TEST_ENV_BASE listed three config-location vars by hand. The resolver reaches 25: every runtime descriptor's configHome.env, plus GROK_AGENTS_HOME (a hardcoded branch of getGlobalConfigDir), GSD_RUNTIME, and GSD_PROJECT/GSD_WORKSTREAM (planningDir). A hand-written list can only ever be as complete as the author's recall, and every one of those resolvers is env-FIRST, so a missing key is a live escape hatch rather than a cosmetic gap -- which is why this bug has now been diagnosed three times. Derive the set from the same registry the resolver reads. The scrub list becomes structurally incapable of being narrower than the surface it guards: adding a capability that declares a new configHome env var extends it in the same commit. Also export TEST_ENV_BASE (a parity test could not previously import the canonical copy) and add scrubConfigLocationEnv(), the in-process counterpart -- TEST_ENV_BASE only ever reaches child processes. Addresses review findings: Blocker 2, Major 4 (GSD_WORKSTREAM/GSD_PROJECT), Major 5 (XDG_CONFIG_HOME). |
||
|
|
08021b02c0 |
fix(#2665): scrub config-location env vars in every TEST_ENV_BASE declaration
TEST_ENV_BASE blanks session-identity variables but none of the three that
decide WHERE a child process writes: CLAUDE_CONFIG_DIR, GSD_RUNTIME and
CODEX_HOME. The config-home resolver is env-first (runtime-homes.cts, the
dot-home case consults the env var before the home-derived fallback), so an
ambient CLAUDE_CONFIG_DIR in the developer's shell beats a call site that
sandboxes only HOME. The suite then writes into the developer's real config
directory -- including a registered skill under <configDir>/skills/ whose
body carries behavioural directives that load into later sessions.
Blank all three alongside the session-identity vars. `...env` still spreads
last, so the five call sites that already constrain these locally keep
winning with their explicit values.
TEST_ENV_BASE is re-declared in nine files, so the three lines are added
nine times rather than once. Consolidating the nine into a single exported
constant -- and fixing the TERM_SESSION / TERM_SESSION_ID drift between the
copies -- is deliberately left out of this change; see the PR body.
One call site needed adjusting. capability-state.test.cjs's
`capability state --runtime claude` CLI test passed no env at all and
compared the CHILD's resolved config dir against the PARENT process's
getGlobalConfigDir('claude'). That agreed only because the child inherited
the developer's ambient CLAUDE_CONFIG_DIR -- i.e. it passed *because of*
the leak. It now redirects both runtime homes into the sandbox and asserts
against values the test controls, so it is hermetic with the variable set
or unset.
Regression case folded into the owning module's test file rather than a new
bug-NNNN file, per scripts/lint-regression-test-names.cjs. It sets the
variable on the PARENT process, which is the actual vector; setting it in
the per-call env argument would exercise a path that was never broken.
|
||
|
|
343835facc |
refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one independent re-derivations across seven modules now route through it, and scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather than an allowlist (ADR-3180 Decision 4a). The epic scoped this at three copies. A whole-repo guard found twenty-six sites across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3 narrows to window plus sentinel enumeration. Two sites are exempt with a documented reason rather than a bare allowlist: audit.cts scans one quick task's own directory for a single completion record, and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time import. Neither is a phase directory. scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one owner answers both questions: verify.cts's numbering-gap check wants every plan on disk, its pairing check wants the live set. Both fields are additive. Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave scheduling, was scheduling status:superseded plans into waves and reporting zero plans for the post-#3139 nested layout. filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned them and only their own tests still called them. New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator, with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and the CONTEXT.md glossary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice The contract held; the phase boundary did not. The whole-repo drift guard found 26 re-derivations across 9 files against the epic's estimate of 3, and cmdProgressRender re-derives both enumeration and plan counting on adjacent lines, so DW4 was unsatisfiable within Phase 1's original file scope. Records the amended scope, scanPhasePlans's new allPlanFiles field, findOrphanSummaries, the two documented exemptions, the re-derived Tier-2 table, and the describeNonCanonicalPlans trap for later phases. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * fix(#3183): complete the canonical pairing rule and gate the naming diagnostic The remote runner went red with 13 deterministic failures on both lanes, and they were right: replacing verify.cts's canonicalPlanStem pairing with summaryCandidates dropped a case the bespoke rule covered. A plan carrying a descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no such candidate, so the plan read unsummarized. The fix is to complete the one rule rather than restore a second: summaryCandidates gains a canonical-id candidate, narrowed to fire only when an id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans and findOrphanSummaries all inherit it. The two-plans-one-summary collision behaviour of the original rule is preserved deliberately and documented in place. Second defect, independently root-caused while verifying: routing the #2893 naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i fallback, which is correct for counting and wrong for a naming check — a non-canonically-named file was accepted as a valid plan and the diagnostic went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect with a strict isCanonicalPlanFile predicate before reporting names. Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180: a question about file naming wants the physical, strictly-matched set; only a question about outstanding work wants the live set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * chore(#3183): register planning-scope.cjs in the eslint migration list tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor ignored per its ADR-457 migration state. The new planning-scope module closed five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md and the CONTEXT.md glossary - but not eslint, because that one is enforced by a test rather than by lint:ci, so the local pipeline stayed green while it was missing. Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is linted, matching every other migrated module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * fix(#3183): replace the plan-count drift detector with a literal tokenizer CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an escaped \.md". Five review rounds found it had two defects, not one: - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.` pair be consumed either as one escape or as two class characters, which is exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`. Excluding `\` from the class killed that but left a cubic path — 23ms at N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI runs on fork pull requests, so a crafted src/*.cts could stall the job. - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g. `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the literal at that `/`, so the scan never reached `\.md` and the guard missed it entirely. (Classes holding an ESCAPED `\/` were already matched; the tests cover those separately as parity, not as regressions.) Both defects have one root cause: regex-literal grammar — `\x` escapes, and `/` inside `[...]` not terminating — is not expressible in a backtracking regex. So the detector is now a tokenizer, not a regex. readRegexLiteralAt reads the literal at a given `/` in a single left-to-right pass with no backtracking, treating escapes as two-character units and suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch restarts it at every `/` on the line, preserving the old "find anywhere" behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the trailing-flag scan — which keeps the whole-line cost linear. Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now caught. Differential against the old regex over 28,474 lines (those matching FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/ gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the old regex hangs): 6 differences, all the tokenizer returning the fuller or newly-correct literal, 0 old-only misses. The `\.md` token stays case-insensitive, matching the `/i` the old regex carried. Also closes three holes in the same new file: - walk() tested entry.isFile(), false for a symlink, so a symlinked src/*.cts was silently unscanned — an evasion of a guard whose stated principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist. It now resolves symlinks, but confined: file links must resolve inside the repo root, directory links inside the scanned dir itself. Every sibling drift guard in scripts/ uses the Dirent classification and never follows links, so following them unconfined would have made this the only linter able to read outside the tree — on fork PRs an arbitrary out-of-repo read whose matched fragments reach a public CI log. The narrower directory rule additionally stops `src/up -> ..` from sweeping the whole repo, and the skip list is now checked against resolved paths so `src/g -> ../.git` cannot reach .git/** or node_modules/**. Real paths are de-duplicated and files reported canonically, so a symlink alias cannot shift which FUNCTION_SCOPED_EXEMPTIONS key applies. - Both the reported fragment and the reported FILE PATH are attacker- controlled source text written straight to a CI log, and git permits control bytes in a filename. Both are now escaped — C0/C1/DEL plus the bidi and zero-width controls — so a crafted literal or filename cannot recolour the log, overwrite a line with CR, or fabricate a line that looks like this guard's own success output. Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process probe over both pathological shapes (catastrophic backtracking is synchronous and would freeze the suite rather than fail one test), the bare-`/` class shapes verified to fail against the parent-commit blob, root-confinement tests covering the outside-file, outside-directory, cycle, broken-link and duplicate cases, direct isInsideRoot coverage including the sibling-prefix case that a bare startsWith would let through, sanitizeForReport coverage, and limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the exported constant. The earlier structural assertion was dropped — it checked for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats while staying exponential. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * chore(#3183): backfill changeset PR number Restores b77931869, which a force-push during the ReDoS remediation dropped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |