e2bfc06558513b96c1dbde1d820ea1ef8a2ae3f9
38 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e2bfc06558 |
fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase (#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until this lands. THE DEFECT, MEASURED Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking value, indistinguishable from the same call with a canonical id. Measured on |
||
|
|
5d4c98cde7 |
chore(#4729): guard the retired-runtime name, and finish the locale residue (#4753)
* chore(#4729): guard the retired-runtime name, and finish the locale residue Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one concern: make the tree clean, and keep it clean. The guard is inert until the tree is clean, and shipping the cleanup without the guard is the one-bug-at-a-time pattern this epic exists to end. WHY A GUARD, AND WHY LAST Nothing in CI answered "does any shipped surface still present a retired runtime as live?", and the two gates that look like they should cannot. checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it reported 0 violations while all four locale mirrors still documented --gemini as a live reviewer flag, with usage examples. And tests/gemini-runtime-removed.test.cjs is scoped by construction - its own docblock limits it to the installer CLI contract and the runtime-name-policy exports; it never reads docs/**, gsd-core/workflows/**, commands/** or agents/**. Every extension to it during this epic was a hand-added assertion for a surface somebody had already noticed. A guard written earlier would have red-flagged the very references phases 1b-4b were removing, which is why it lands last. PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE Each site was judged against its ENGLISH counterpart, not on its own: README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46 English README.md has ZERO occurrences -> substituted "Antigravity CLI, Kimi CLI" how-to/execute-a-phase.md:88 x4 locales fixed in #4728 -> substitute how-to/verify-and-ship.md:89 x4 locales fixed in #4728 -> substitute FEATURES.md cross-AI CLI list :1419 no Gemini -> DELETE FEATURES.md REQ-MULTI-RT-01 :1709 -> substitute FEATURES.md REQ-SKILLS-03 :1952 -> rewrite FEATURES.md REQ-QUOTA-02 :3256 deleted upstream -> delete VERSIONING.md:133 stale manifest -> see below The twelve README occurrences were an adversarial reviewer's BLOCKER, and the reason they survived my own sweep is structural: root-level *.md was outside the guard's scan set, so the repo's most-read runtime-advertising surface was invisible to the guard meant to police it. :46 is a live installer-runtime claim - it tells the reader the installer will offer a runtime that no longer exists. Checked for the duplicate-name trap before substituting: neither Antigravity nor Kimi appears anywhere in those four files. Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work, not a deferral. Two more show why "substitute Gemini -> Antigravity" is the wrong default: in the cross-AI list and REQ-QUOTA-02 English DELETES the name, because Antigravity was already in the list or the classifier had dropped it. Substituting would have duplicated a name - the identical trap ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot. VERSIONING.md:133 is a different and worse defect than translation lag. Under "Manifest Version Sync" it listed gemini-extension.json as a version-synced manifest. That file is ABSENT from the repo, and scripts/sync-manifest-versions.cjs says so in its own comment - "#1928: gemini-extension.json was removed with the gemini runtime ... it is no longer a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json, marketplace.json and vscode/package.json. So the doc named a manifest that does not exist AND omitted the one that replaced it. Both fixed, verified against the owning code rather than inferred from the name. The replacement bullet cites #1942, the issue that actually registered vscode/package.json, matching the convention of its neighbours. pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has no REQ-QUOTA-02 line. Skipped and recorded, never invented. PART B - THE GUARD scripts/lint-retired-runtime-name.cjs, modelled on scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem shape (forbid a retired token, allowlist frozen content, self-exempt via a split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1). Case sensitivity IS the mechanism, not an accident. The naive guard - "the string gemini must not appear" - is WRONG, not merely noisy: that string is load-bearing across Antigravity's real on-disk contract. A case-sensitive, standalone, capitalised name works because every legitimate reference is spelled differently and therefore cannot match: lowercase config homes (~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and GEMINI.md. Table-driven, so the next retired runtime costs one row. THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two rounds of isolated adversarial review reshaped it; both are recorded in .gsd/bug/chore-4729-gemini-drift-guard/60-review.json. ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can never be complete, so "no runtime word found" silently exempted every phrasing nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`, `Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0, as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and `Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime` or `运行时`. The fix was to stop enumerating exceptions and invert the evidence direction: Tier 1(a) - the hook DIALECT Antigravity inherits. Position is language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The marker must now form an ADJACENT COMPOUND with the name, not merely sit in a +/-24-character window - that window let `| Antigravity | Gemini-style hooks | Gemini support is live |` exit 0, one legitimate reference licensing a fresh live claim 21 characters later. The runtime-word veto is now LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files" in the same sentence); each is an explicit pin rather than a reason to loosen the veto for everyone. Measured: widening it surfaced exactly those ten and no others. Tier 1(b) - the provider/model axis. A version optionally followed by a qualifier, including full-width digits and CJK punctuation, AND positive model-axis evidence on the line, AND no runtime word. The positive requirement is the part that matters: all eight real model-axis lines in the repo name a model explicitly, so requiring it costs nothing on the real tree while flagging every laundering attempt. It is also the honest resolution of the agent/target tension below - rather than guess at an exhaustive veto list, stop treating an empty veto as evidence. Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match falling INSIDE an occurrence of its own snippet. Line-level containment let `Known provider menu update: Gemini CLI is once again a selectable GSD runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your GSD runtime.` both exit 0, because a short snippet elsewhere on the line pre-approved a brand-new claim. Span scoping makes short snippets safe: `Google (Gemini)` can only ever excuse the match inside those 15 characters. A LOAD-TIME validator now requires every pin to contain a retired name, and it immediately caught five of MY OWN pins whose snippets sat BESIDE the name rather than covering it - each would have shipped permanently inert and permanently reported stale. All pins were then reconciled in one pass. A pin is also marked used by PRESENCE on the line now, rather than only on the Tier-2 branch. Previously a pinned line that a general rule also matched never marked its pin used, producing a provably FALSE "no line matches pinned snippet" whose printed remedy told the maintainer to delete a pin that was still needed. Tier 3 - blanket trust, and a new occurrence inside it IS invisible. CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one surface - plus six append-only directories. All 21 `.changeset/` hits were measured to be fragments DESCRIBING the retirement or a fix to it, 464 of them under archived/; a fragment can only describe what already shipped and is deleted at release, so pinning them would be friction with no signal. The cost is stated in the guard's own header rather than hidden. THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and `references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a live claim introduced there was invisible at BOTH ends. The escape hatch must now carry a justification (`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is checked first, excuses the whole line, and the failure message advertises it, so an unexplained one is indistinguishable from a silenced defect. Plus an anti-vacuity floor counting files actually READ, not files listed - a candidate count stays healthy-looking even if every read failed. A FALSE NEGATIVE I INTRODUCED, AND CLOSED The model-display escape began as a blanket /^ \d/ - "space then a digit" - which also matched "Install for Gemini 2.5 CLI as a supported runtime.", laundering a genuine stale-runtime claim through an attached version number. That was the THIRD appearance of one failure shape in this epic: an exclusion added to suppress false positives creating a false negative. #4716's sweep excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby hid a stale review.models.gemini row whose example value was "gemini-2.5-pro" ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which is why the fix this time was to invert the rule's evidence direction rather than to enumerate more exceptions. The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI" matched inside "client" and would have vetoed legitimate model lists - and raw for CJK terms, where \b is ASCII-word-based and would never fire beside an ideograph, so anchoring them would silently disable the veto in ja/ko/zh. "agent" and "target" were deliberately left OUT: both occur throughout ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so vetoing on them would red correct content instead of catching runtime claims. The reasoning is in the guard's comment, not just the omission - and Tier 1(b)'s positive-evidence requirement is what makes that omission safe, since the rule no longer depends on the veto list being complete. COVERAGE tests/lint-retired-runtime-name.test.cjs drives the guard through its GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a guard, and this epic already shipped one that was vacuous for 2 of its 5 files, so properties are paired against BOTH failure modes - too broad silently absorbs a future defect, too narrow reds on legitimate content. Floor boundaries are covered at 149/150/151. The round-2 reviewer's sharpest point was about that claim, and it was right: the first matrix's pairing was "true of the properties chosen, not of the predicate's actual surface" - not one of its twenty properties could see the dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an over-broad pin colliding with a new line. Every one of those is now a committed regression using the reviewer's own attack line verbatim, and the local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0). That harness earned a finding of its own. Its first run reported PASS=2 FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so nothing staged, every fixture hit the empty-walk error path, and the two checks that assert an ABSENCE passed off that error path rather than off real guard logic. A staging failure is now fatal and every absence-asserting check first proves the walk ran and the expected violation was flagged. Later, one case failed because its fixture supplied only one of a pinned file's two approved lines, so the stale-pin check fired correctly - the expectation was wrong, not the guard. Telling those two apart is the whole value of running a matrix rather than reasoning about one. On the two orthogonal reviews: the isolated adversarial pass executed a great deal of code, across two rounds, against its own fixture repos. The security pass did NOT - it self-discloses that it verified by reading only, because node --test is hard-blocked here. Saying so plainly, because "two orthogonal reviews" without that caveat overstates what the second one established. It also raised, and I cleared by measurement, a concern that importing escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt clone: six other tracked scripts already require that exact path, three of them already in lint:ci, and .github/workflows/test.yml:192-193 runs `npm run build:lib` immediately before it for exactly this reason. Part A has no new test deliberately - those edits are covered by the guard itself inside lint:ci, and a separate per-locale assertion would duplicate it and then drift from it. The one exception is the root README case, which IS pinned: that residue was invisible to the guard rather than merely unasserted, so the fix is a scan-set change and needs its own regression test. No mode-bit read-failure fixture was added on purpose: the benches run as root, where chmod-based IO injection is vacuous, so such a test would assert nothing. The test's fixture helpers write throwaway docs/ paths, which trips lint-docs-guard-registration's reader-name heuristic. Resolved the way that lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry - because the fixtures only WRITE scratch data and never read shipped docs; the baseline was re-confirmed, not merely extended, each time locale and adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own --write path, since a new test file changes the count lint:generated-sync reads. Fixes #4729 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4729): backfill changeset PR number (#4753) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
48271de43f |
fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis (#4744)
* test(#4660): pin the letter-axis parity defect across all 6 shell/markdown phase-id sites Extends tests/nsegment-phase-grammar.test.cjs (#4568) one axis over: for each of the six sites, reads the live regex off disk and asserts it agrees with src/phase-id.cts's PHASE_NUMBER_TOKEN_SOURCE on the letter axis in BOTH directions — accepts `12A` / `3A` / `03A` / `23A.1.2`, still rejects `3a`, `3AB`, `A3` and the other canonical-invalid shapes — and that the two extracting sites return the full letter-suffixed token rather than its digit prefix (or nothing). Negative control against the unfixed tree: 22 failures, exactly the "(fails before the fix)" cases; every reject-parity case already green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis Adds `[A-Z]?` after the leading digit run at all six sites #4568 widened — the ERE translation of src/phase-id.cts's `\d+[A-Z]?(?:\.\d+)*` — so a documented, canonical-valid id like `12A` or `23A.1.2` is no longer refused by the four validating sites (code-review.md, code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md) or truncated to its digit prefix by the two extracting sites (execute-plan.md's plan-filename grep, plan-phase.md's --research-phase capture). Behaviour is byte-identical for every id that matched before; the adjacent comment and error-message text now names the grammar it mirrors. Driven: `init code-review 3A` on a fixture with a `03A-slug/` directory and a `### Phase 3A:` heading emits `padded_phase: "03A"`, which the old regex rejects and the widened one accepts — nothing upstream of the validator mangles the id. At execute-plan.md the trailing `-[0-9]+` is the PLAN number and stays digit-only; plan and milestone dimensions are out of scope per the brief. `CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE` derives from the canonical source by a literal `.replaceAll('A-Z', 'A-Za-z')`, so src/phase-id.cts is deliberately untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4634): extend lint-phase-id-drift to ban a letter-less phase-id mirror in workflows/ and agents/ Adds findLetterlessPhaseMirrorDrift — the letter-axis twin of the #4568 single-segment rule — flagging the unbounded-segment shape `[0-9]+(\.[0-9]+)*` (and its \d / doubled-backslash near-variants) whose digit run is NOT followed by the `[A-Z]?` class, on any phase-carrying line across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md and agents/**/*.md. Sanctioned the same way (`<!-- phase-id-owner: ... -->`), tolerates the case-flexible `[A-Za-z]?` directory-scanning variant so it cannot force that separate axis to narrow, and is wired into scanAll. Confirmed zero violations against the real tree post-#4660 fix, and one violation when a single site is reverted. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * docs(#4660): add Fixed changeset Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore: regenerate conformance-tier manifests for the extended grammar test tests/nsegment-phase-grammar.test.cjs now requires the compiled gsd-core/bin/lib/phase-id.cjs (to assert the canonical grammar agrees with each site's live regex), which moves it to a different platform-conformance tier; `gen-platform-conformance-tier.cjs --check` in lint:ci flagged the macOS manifest as stale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * test(#4660): reword a comment that tripped lint-docs-guard-registration The comment mentioned `docs/CONFIGURATION.md` between two backticked tokens, which the lint's template-literal detector read as a docs/ path expression. The test reads no docs/ file. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): refresh the compact-content benchmark baseline and acknowledge emitted growth plan-phase.md grew by 4 bytes (`[A-Z]?`), which moves the committed compact-content benchmark; refreshed with `benchmark-compact-content.cjs --write`. The six shipped files below grew by the widened regex literal plus the comment and error-message text that now names the canonical grammar. Emitted-Drift-Ack-Growth: code-review.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: code-review-fix.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example Emitted-Drift-Ack-Growth: gsd-code-fixer.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the defense-in-depth comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the comment and error message updated to the canonical grammar Emitted-Drift-Ack-Growth: execute-plan.md — #4660: `[A-Z]?` in the plan-filename phase extraction (6 bytes) Emitted-Drift-Ack-Growth: plan-phase.md — #4660: `[A-Z]?` in the --research-phase capture (6 bytes) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 * chore(#4660): set changeset fragment pr to 4744 * chore: re-trigger Validate Branch Name The required check-branch context was cancelled on this head by the workflow's cancel-in-progress group when the changeset pr-field backfill push landed three seconds after the PR opened; no completed run exists for the current head, and a fork contributor cannot re-run it. Empty commit to re-run it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
eb49ff98df |
fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime
#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.
The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:
- install-on-your-runtime.md English has NO `### Gemini CLI` section -> deleted
- USER-GUIDE.md :843 "…, Antigravity CLI, Kilo)" -> substituted
- ARCHITECTURE.md English has NO Gemini CLI table row -> row deleted
- ARCHITECTURE.md :24 English holds `Kimi CLI` in that slot -> Kimi CLI
- context-monitor.md :3 "`AfterTool` for Antigravity CLI" -> substituted
- spike-and-sketch.md :93 "(Codex, Antigravity CLI, etc.)" -> substituted
- configure-model-profiles "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
- COMMANDS.md English keeps only hyphen + Codex bullets -> colon bullet deleted
- FEATURES.md source docs/features/multi-runtime-support.md:10
lists no Gemini CLI -> name removed
ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.
The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.
The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.
Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.
Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.
Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.
Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.
Fixes #4728
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate
A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.
1. I committed the exact error I claimed to have avoided. The commit message
boasted that ARCHITECTURE.md:24 proved the value of reading English rather
than substituting blind, because Antigravity already appeared later in that
list. Five hundred lines further down the SAME four files, my
`Gemini:` -> `Antigravity:` substitution produced TWO consecutive
`- Antigravity:` bullets, because an Antigravity bullet was already there.
English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
locales, reusing each locale's existing words.
2. `--gemini` survived in the runtime-detection CLI flag list in all four
locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
already lists `--antigravity` later, so this is another place where
substituting Antigravity would have duplicated it. Now `--kimi`.
3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
line I had already corrected -- the "Adaptive (Recommended)" option in
settings.md:192 and new-project/steps/auto-mode-config.md:95.
4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
`non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
without the word "runtimes", and execute-phase.md:1028 has no parenthetical
at all. The reviewer proved it by re-introducing Gemini at both lines and
watching the assertion stay GREEN. That same blind spot is what hid finding 3.
Replaced with a case-sensitive `/\bGemini\b/` walk over every
`gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
reference in that tree is spelled differently and cannot match: Antigravity's
paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
`Gemini` there means the retired RUNTIME is being named. The walk asserts it
found at least 50 files so an empty walk cannot pass vacuously, and it now
covers the nested `new-project/steps/` directory where finding 3 lived.
Two allowlist entries, both by line CONTENT and both justified:
reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
`Known provider` menu. The second was escalated by the agent rather than
decided: Section 8 of that file says model policy is defined "independently"
of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
the same axis as the lowercase model ids -- and must keep working.
Proven to fail, not just asserted: the predicate reports 0 offenders on the
real tree and exactly 2 on a /tmp copy with Gemini re-injected at
health.md:52 and execute-phase.md:1028.
Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.
The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.
Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @
|
||
|
|
9c2927bff4 |
fix(#4733): derive the win32 chunk cap, isolation bar, and unknown-file weight (#4737)
* test(#4733): pin the cap, unknown-file weight, and isolation rules Failing-first coverage for the three defects that let a Windows conformance chunk be killed at the 600s per-chunk backstop with zero failing tests. The previous boundary rows were VACUOUS: they asserted literal arithmetic (21 * 18122 <= 400000) that cannot fail, and in doing so masked a shipped win32 cap of 23 -- a value that violates the very inequality they claimed to pin. These rows constrain defaultMaxFilesPerChunk itself, from both sides, so the shipped value is a derived maximum rather than a magic number. A second vacuous row was caught by review and removed: it recomputed the isolated set from the function under test using the identical predicate, so it was empty by construction. It is replaced by an exact deepEqual against the expected basenames, a cross-platform identity row, dynamism rows in both directions, an inclusive boundary triplet, and invalid-threshold throw rows. The cross-platform identity row is the regression guard for a threshold that was briefly anchored to the per-platform file-COUNT cap; it fails if isolation ever becomes platform-dependent again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4733): derive the win32 cap, isolation bar, and unknown weight A Windows conformance chunk was killed at the 600000ms per-chunk backstop with no test having failed, taking next red. Three compounding defects. The win32 cap of 40 permitted 40 * 18122 = 724880ms against a 600000ms backstop -- 121% of it -- so two rounds of budget-tuning could not hold. The cap is now derived: 22 is the largest value satisfying cap * 18122 <= 400000. The budget is 400000, not the raw backstop, because the chunk that died summed to only ~348328ms of per-file time -- a per-chunk overhead gap of at least 1.72x that no per-file table models. A file absent from the timings table was priced at medianWeight. The table is skewed 18.8x, so an unknown weighed 0.0533 -- 19x cheaper than average, and measured 17.5x under its real cost. Unknowns are now priced at the mean. ISOLATED_HEAVY_FILES was a static Set, stale by construction. Isolation is now derived from an absolute ms bar (0.3 * 400000 = 120000ms) converted to weight units via the live table's mean, so a file that gets heavy is isolated automatically instead of waiting for someone to edit a list. Review caught that an earlier cut anchored that bar to the per-platform file-COUNT cap -- a category error, count vs weight, which silently returned seven of the historical eight files to the shared pool on linux/darwin. Since macOS runs the full matrix only after merge, that would have planted a red next no PR could catch. The bar is absolute and platform-independent. Also from review: isolation no longer requires unit-suite membership, so fragment-single-edit-propagation.install.test.cjs -- 575000ms, 96% of the backstop in one file -- is eligible; partitionIsolatedFiles throws on a non-finite or non-positive threshold instead of silently isolating nothing; and stale per-shard figures no test pinned are removed rather than recomputed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4733): backfill changeset pr number --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ca8d9d4459 |
fix(#4429): stop a large commit_types config blocking or bypassing the gate (#4723)
* test(#4429): regression coverage for three defects in the commit hook Failing-first coverage. Every conforming-subject row is red against the unfixed hook, and each defect gets an explicit CONTROL row that reconstructs the pre-fix form and asserts the defect reproduces -- without those, the passing rows would pass with or without the fix. 1. SIGPIPE (the reported defect). The pre-fix first-line extraction used a `head -1` pipeline; once CONFIG_OUT exceeds the 64 KiB pipe buffer printf is killed and `set -euo pipefail` aborts the hook. That fix is already on next -- it landed incidentally in #4537, whose message never mentions #4429 -- and nothing in the tree would notice its removal. 2. regcomp. The commit-type alternation grew with the CONFIGURED list and exceeded bash's 64 KiB compiled-pattern cap. Boundary rows pin the cliff at 6051/6052, with controls on BOTH sides so limit-1 is not vacuous. 3. Ambient subprocess statuses (found by this change's security review). Defects 1 and 2 cannot be separated: each configured type adds len+1 bytes to CONFIG_OUT and len+1 to the alternation, so the smallest payload that overflows the pipe (N=6059) already puts the alternation past the ceiling. The SIGPIPE control accepts either SIGPIPE (141, Linux) or a reported write error (macOS bash 3.2's builtin printf, exit 1). Asserting only the message would go red on every CI lane, since the remote matrix is Linux-only. Named to bucket with gsd-validate-commit-crash-policy.test.cjs, which covers this same hook: lint-test-file-count derives a test's owning module from its filename prefix, and `validate-commit-*` collided with the `validate` module, already at its 2-file cap. Harness note, learned from three vacuous control runs: hooks/lib/git-cmd.js requires ../gsd-core/bin/lib/token-scanner.cjs relative to the hooks dir's parent, so a copy in a bare tmpdir fails open and returns 0 for any input. The layout symlinks gsd-core beside the copy, and every row that can prove it asserts the run was substantive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4429): bound the commit-type regex and isolate subprocess statuses Two fixes in the same file, both of the same shape: a value computed for one purpose was being read as authority about something else. 1. The commit-type alternation could not be compiled. COMMIT_TYPE_ALT joined every CONFIGURED type into one regex, so the pattern grew without bound. bash caps a compiled pattern at 64 KiB. Bisected on bash 3.2.57 (this repo's macOS target): a 65504-byte alternation compiles, 65515 fails. `[[ =~ ]]` returns 2 on a compile failure, and `if !` cannot tell that from "the subject does not conform" -- so the hook blocked a valid `feat(auth): ...` with CONVENTIONAL_COMMITS_VIOLATION while printing `feat` in its own valid_types. Match the shape with a fixed-size pattern, capture the type, then test membership against the COMMIT_TYPES array. The character class is exactly the `^[a-z][a-z0-9-]*$` safe-token filter the config loader already applies, so it captures every type that can legally reach COMMIT_TYPES and no token that cannot. Review verified equivalence over 46 handcrafted plus 6000 randomized adversarial subjects against a type list containing prefix-overlapping, digit-bearing and trailing-hyphen types: zero divergences. The loop adds no subprocess and no pipe, which is the hazard class #4429 is about. COMMIT_TYPE_ALT is now unused and removed. types pre-fix `feat(auth): ...` fixed 10 accept accept 6051 accept accept 6052 BLOCK accept 20000 BLOCK accept 2. Subprocess statuses were inherited from the environment. Each status is captured as `... || VAR=$?`, which assigns ONLY on the failure branch; on success the variable kept whatever it already held, and `${VAR:-0}` defaults only when unset or empty. So an EXPORTED CONFIG_STATUS, CMD_STATUS or CLASSIFY_STATUS -- from a CI wrapper, a .envrc, or another hook -- survived into the success path and was read as "the subprocess failed". Since the hook fails OPEN on a genuine subprocess failure by design (#3838), the result was a silent bypass. Measured: `CLASSIFY_STATUS=3 git commit -m "nope: bad"` printed "validator disabled for this call" and exited 0. The three are now initialised before use. The fail-open path is unchanged and verified byte-identical to origin/next with a failing node. hooks/dist/ is gitignored and rebuilt from hooks/ by scripts/build-hooks.js, so there is no second copy to sync. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): register the new suite with the conformance manifests Both conformance-tier manifests embed the test-file list, so adding a test file makes them stale. Regenerated with their own generators: node scripts/gen-platform-conformance-tier.cjs --write node scripts/gen-platform-conformance-tier.cjs --target macos --write Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): pin both fail-open-prone controls to their named cause Two rows in the ambient-status block asserted `status === 0`, which the hook also returns when the harness layout is broken -- so either row could have passed for entirely the wrong reason. This is the same vacuity trap the rest of the suite already guards, applied inconsistently to the rows added last. Measured, rather than reasoned about: genuine ambient bypass (pre-fix hook, CLASSIFY_STATUS=3) rc=0, no CLASSIFIER_THREW orphaned layout (no gsd-core symlink) rc=0, CLASSIFIER_THREW genuine fail-open (node shim exits 3) rc=0, no CLASSIFIER_THREW So assertSubstantive separates the intended cause from the harness failure in both rows, and each now pins its pass to the cause it names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4429): stop asserting a macOS-only regex cap on every platform First verification run was RED: 45636/45638 passed, both failures in this new suite on linux-node24. Cause is mine -- I measured the compiled-pattern ceiling on macOS and encoded it as a cross-platform expectation. Measured in the tester image itself: engine 6051 6052 20000 bash 3.2.57 / BSD libc (macOS) compiles rc 2 rc 2 bash 5.2.15 / glibc (Linux) compiles compiles compiles (228943 B) glibc has no reachable cap, so the regcomp defect cannot occur there and the control asserting a block at 6052 was red for a behaviour the platform cannot produce. The control now calibrates at runtime: it runs the pre-fix form and, when this engine compiled the alternation, it SKIPS with a message naming the reason rather than asserting. Skipped out loud, never silently passed -- a green row there would read as "the defect is covered" on a platform where it cannot occur. Both branches verified: the capped branch asserts (macOS 17/17, zero skipped), and the uncapped branch was exercised by forcing the payload to a size that always compiles, producing a skip and not a failure. Consequence stated rather than hidden: the remote matrix is Linux-only, so this one control is skipped in CI and really runs only on a macOS workstation. The rows that run everywhere are the ones carrying the regression weight -- the shipped hook accepting a conforming commit at every payload size, the gate still blocking unknown types, the SIGPIPE control, and all seven ambient-status rows. Note this also narrows the coupling claim: SIGPIPE and regcomp are coupled only on a capped engine. On glibc the SIGPIPE defect is directly testable without the regcomp fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4429): scope the regex-cap claim to the platform it applies to The changeset told users the validator "built a regular expression bigger than bash can compile" past ~6,000 configured types. That is false on Linux: glibc compiled a 228,943-byte alternation without complaint, so a Linux reader would have been misled about their own exposure. These are user-facing release notes, so the claim is now scoped to macOS (bash 3.2 / BSD libc) and says explicitly that glibc was never affected by this half. The hook's own comment led with the same overstatement -- "bash caps a compiled pattern at 64 KiB" -- before qualifying it. Reworded so the first clause states what is actually true: the limit is a property of the platform's regex engine. Text only; no behaviour change. Suite 17/17, eslint and lint:ci clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4429): backfill changeset PR number (#4723) * chore(#4429): backfill changeset PR number (#4723) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c0b2a05d2f |
fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser The isolation guards decided whether a run-scoped sentinel applied to a dispatch by regex-scraping model-authored prose. The scrape returned values in a different namespace from the ones the sentinel records, so the comparison could never succeed: sentinel { phase: "03", plan: "03-02-hardening" } <- $PHASE_NUMBER / $plan_id prose "Execute plan 02 of phase 03-auth." scraped { phase: "03-auth.", plan: "02" } <- greedy (\S+), both wrong #4594 reports only the phase half. Measured against a real phase-plan-index run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose carries a bare in-phase plan number — so the plan field mismatches too, and the Claude path is dead rather than latent. A fresh sentinel was therefore discarded on every executor dispatch and every legitimate ISOLATION=none degrade was denied, leaving the work unrun. hooks/lib/dispatch-identity.js is now the single owner of both halves. The two prompt-body producers emit a canonical marker carrying the same shell values the sentinel records, so producer and consumer agree by construction. The prose frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and deliberately reporting no plan — an absent identifier means "cannot compare" and is safe; a wrong one is a false mismatch and is not. The prose sentence itself is byte-identical: the executor agent reads it too, so the marker is purely additive (Hyrum's Law). An inapplicable sentinel is now named in the guards' deny reason instead of being dropped silently — the silence is why this survived three producers and two consumers unnoticed. Interpolated values come from a sentinel file and from prompt text, so both are length-bounded and stripped of control characters. ADR-4630 locks the seam and maps the epic's three phases. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4594): resolve eight review findings across the dispatch-identity seam Three orthogonal review engines ran on 43418af144 — the code-review skill's Standards and Spec axes, and an isolated adversarial security pass — plus a self-review of the committed diff. Every finding is fixed here; none deferred. F1 (major, reproduced). A keyless or unknown-key-only marker — the literal "[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and returned source:'marker' with both fields null, suppressing the prose fallback entirely. Any prompt text containing that literal silently disabled identity narrowing, so a fresh sentinel applied to a dispatch it was never scoped to, defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker that yields neither recognized key is no longer a marker: the scan continues to later markers, then later texts, then prose. Forward-compatible tolerance of unknown keys is unchanged. F2/F3 (major). The first cut duplicated sanitizeForReason, describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across both guards — the exact defect class this epic exists to delete, and with no cold-load justification, since both hooks already require hooks/lib/. They now live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in isolation-sentinel.js beside the comparison it mirrors, returning the nested {sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke four-field bag that renamed the pairs already flowing through the seam. F4 (hard violation). The visibility test asserted on the deny reason's prose. CONTRIBUTING.md prohibits raw text matching on hook output, which is why every deny carries a stable reason_code. The discard is now a structured sentinel_discarded field on each hook's stdout JSON, and the test asserts that; the sentence stays for the operator but is no longer the contract. F5 (hard violation). The 64-character truncation limit had no boundary coverage. 63/64/65 are now exercised against the single consolidated helper. F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or the bidi overrides, so a crafted value could still reflow or reverse the message. Both classes are stripped, with a test each. F7 (major). The producer/template parity test was vacuous — it rendered a marker and re-parsed its own output, and would have passed with both templates deleted. It now reads the two workflow templates, extracts each marker line, substitutes the measured values and asserts the owner's parser returns them. Proven red by deleting one template's marker line before being proven green. F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the orchestrator-worktree path because that prompt is built in shell. It is not: executor-isolation-dispatch.md:131 says plainly that those are template placeholders, not shell variables, so {plan_id} is model-substituted there too. A false guarantee in a design lock is worse than a stated limit. Both documents now say the marker is model-substituted on both paths and that the prose fallback is the real floor everywhere. The "3 workflow templates" count was also wrong — 3 prose sites across 2 files, 2 of which carry the marker. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth Refs #4630. The dispatch-identity marker and its substitution note grew gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts two real-tree guards that lint:ci does not run: - tests/benchmark-compact-content.test.cjs asserts the committed baseline is "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on 23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline regenerated with scripts/benchmark-compact-content.cjs --write. - tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer for any emitted file that grows, keyed on the bare filename. Added below. The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line inside the Agent() prompt's <objective>, and the note telling the orchestrator to substitute {plan_id} with the plan's id verbatim. Both are load-bearing -- the marker is what lets a guard hook match a dispatch to the sentinel the per-plan gate wrote, and without the note the orchestrator has no instruction telling it the value must not be paraphrased. Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): set changeset fragment pr to 4693 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bbdf7e8e84 |
chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold. THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser, not grep) found what the epic never enumerated: ADR-4650 named seven containment implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13 files, several guarding a write or an `fs.rmSync`. Two verified by reading rather than pattern-matching — `research-store.cts` comments its own as "ensure the resolved file path stays inside the store dir" immediately before a write, and `capability-lifecycle.cts` gates `fs.rmSync` with one. So the epic's Done-when "one containment predicate, used at every site" was FALSE when Phase 3 reported it satisfied. It is true now: the rule is clean across src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist. WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose first argument is a managed root and whose later arguments derive from argv. That is a taint analysis over 2046 call sites, in ESLint, without type information; "derives from argv" is not locally decidable. Any approximation either floods or is trivially evaded, and a rule that fires on hundreds of correct sites earns an allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is the COMPARISON, not the join, and that has one recognizable shape. Arm 1 X.startsWith(Y + sep) the hand-rolled containment idiom Arm 2 a containment predicate called as a bare statement, answer discarded Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper was called". Its example `validatePath(x, root).resolved` is already structurally impossible — Phase 3 un-exported `validatePath` — so the remaining expressible failure is ignoring the answer, which is the defect that recurred five times in this epic. The census found exactly one live instance (`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers stop re-deriving the path the comment above it was extracted to stop them re-deriving. The rule deliberately does NOT try to catch validate-one-path-use-another where the answer is used but a different variable flows onward. That needs flow analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the two are complementary. PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a lexical site onto the realpath family broke four tests and was caught only by the matrix. Every migrated site was triaged individually. The six installer-migrations tree-walks and the six capability-lifecycle gates take the LEXICAL family because their operands are already realpath-resolved and they deliberately treat the final component as a link; boundary sites take realpath. TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken. `installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT `target === root` by contract, while the canonical comparison ACCEPTS it. Swapped naively, a migration could `rmdir` the user's config root and a third-party descriptor could write at configHome itself. Both keep `=== root` as an explicit additional arm alongside the predicate call — the predicate decides containment, the call site keeps its own extra condition (ADR-4650 decision 6). ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was byte-identical to `isContainedIn` and said so in its own docstring. `isContainedIn` is now exported for callers that have already resolved both operands and need only the comparison, with a doc note that a caller which has NOT resolved them must use a full predicate instead. THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable reason. Two justifications: (a) not a containment decision — an ancestor-walk loop condition, sub-repo grouping, worktree identity matching, declared-path coverage; (b) it IS containment but the canonical predicate is unreachable — `capability-validator.cjs` is a committed pre-build `.cjs` and the compiled `security.cjs` is untracked build output, so requiring it would break a fresh clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure. The marker was renamed from `allow-lexical-prefix-match` mid-phase because that name asserted only (a) and would have stated something false at the (b) sites. A marker suppresses BEFORE the violation counter increments, so a file whose every occurrence is marked still reports `staleAllowlistEntry` — otherwise a drained entry lingers and silently re-permits the site later. DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/` file made `npm run lint` fail with the rule's full guidance message; removing it returned the tree to clean. Both halves recorded — red alone proves nothing, since a rule red for an unrelated reason looks identical. DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two distinct rejection messages and two manual realpath calls; routing it through `tryWithinRoot` collapses them to one message, and a missing module now surfaces as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The security property is preserved and slightly strengthened — the candidate is realpathed and containment re-checked, and the dangling-symlink oracle closure comes along with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4654): record the containment ratchet in CONTEXT.md and the security model Both entries previously described the seam without the thing that keeps it a seam. They now state what the rule bans, and — more usefully for whoever reads this next — what it deliberately does NOT attempt: deciding per path.join call whether an argument came from user input. That question is not locally decidable, and an approximation across ~2000 join sites would earn an exemption list of hundreds, which is the opposite of a ratchet. Also records the marker's two legitimate justifications and that its reason is mandatory, so the escape stays reviewable rather than becoming a mute button. Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json regenerated and confirmed byte-identical rather than assumed — which also confirms eslint-rules/ is not a shipped path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): close review findings and the two matrix failures MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my evidence for collapsing it was wrong. I searched tests/ for the literal string "resolves outside its install root", found nothing, and reported that no test asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module pointing OUTSIDE the install root is not loaded" — the test guarding the exact property I claimed was preserved. defaultRequireFromInstallRoot now does both checks again with both messages byte-identical, each routed through the canonical predicate, which is better than the original since that hand-rolled both comparisons. MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments, so a suppression marker inside a plan body drifts the baseline exactly as an edit does. Measured: with markers in place, two of the four still differed from their committed checksums. The four shipped bodies are now byte-identical to next, and the rule's config excludes those four paths BY NAME rather than by a directory wildcard, so a NEW migration is still covered. Six containment comparisons stay un-ratcheted there; that gap is recorded in the rule's Known gaps, in CONTEXT.md and in the security model rather than left implicit. Justification (c) is removed from the marker's documented reasons, because a marker was proven unable to express it. ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT shape while permitting the incorrect one: startsWith(root) with no separator is the genuinely unsafe form, since it accepts a sibling such as root-evil, and my own test blessed it as valid. Flagging every bare startsWith would swamp the rule, so that stays a STATED gap rather than a silent one. Closed for real: the template-literal spelling, which the census never saw because it only inspected plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated. A separator reached through a const alias is now resolved via scope analysis. And isContainedIn, exported in Phase 3, was missing from the discarded-result set, so a bare no-op call went unflagged on the one function the epic funnels through. SECURITY REVIEW — the marker could over-suppress two ways: a block comment worked identically to a line comment, and one marker silently covered every violation sharing its line. It now requires a Line comment positioned after the flagged node ends, so it anchors to the node it trails. Four sites had dropped an unreachable-but-deliberate equality rejection against the root; each is restored as the call site's own arm. eslint.config.mjs still documented the OLD marker token, which my rename missed — it would have sent the next author in circles. A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a stale eslint cache while twelve real violations existed. Every lint check here now clears the cache first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): anchor a suppression marker to the violation it actually trails The matrix caught this; my own test caught it, on its first execution. The case "two violations on one line: trailing marker suppresses only the one it trails" expected 1 error and got 0 — both were suppressed. ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose range started at or after the node's end. A trailing marker at the END of a line sits after EVERY node on that line, so that condition held for all of them. "After the node" does not identify WHICH node the marker trails. The fix reads as correct and is not. FIX: deferred reporting. Violations accumulate during traversal instead of being reported immediately; at Program:exit each marker claims exactly ONE pending violation — the one on its line whose end is nearest before the marker begins — and every unclaimed violation is then counted and reported. One marker, one suppression. An earlier violation sharing the line is still reported, which is the property the security review asked for and the previous attempt only appeared to deliver. The counter now increments at flush time rather than during traversal, so a suppressed occurrence still does not keep an allowlist entry alive. AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test` is hard-blocked here, so this rule's test file could only ever be executed on the remote matrix — which is why a broken anchoring shipped into a run. ESLint's programmatic Linter API is not a test runner, and exercising the rule through it verifies every case locally in seconds. All 24 now pass locally, including the two-on-one-line case that failed remotely. That loop should have been built before the rule was first sent to the matrix rather than after it failed twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs now carries both, with the how-to quadrant skipped for a stated reason rather than an empty field. The audience for this deliverable is a contributor who trips the rule, and the task-oriented guidance reaches them in the ESLint message itself — which names the correct predicate, says how to choose between the realpath and lexical families, cites the Phase 3 regression caused by choosing wrong, and gives the marker syntax. A docs/how-to page would be a second, driftable copy read by nobody at the moment of failure. enablementSequence is recorded as what it actually is: a VERIFICATION sequence, not an enablement one. The rule is never off, so there is no off-to-on transition to describe. scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not evaluate against the mandated pr:0 placeholder. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d1d9ee85a5 |
feat(#4142): thread convention through the completion-path membership seam (#3644)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
9c20b7b40a |
fix(#4652): use the validated path, collapse the duplication, correct two false claims
Seven findings from the two-axis review, all fixed in place. THE ONE THAT MATTERS: cmdTodoComplete validated sourcePath and targetPath and then ran every fs call against the RAW strings — existsSync, statSync, readFileSync, platformWriteSync, unlinkSync, and the dry-run path payload — never sourceCheck.resolved / targetCheck.resolved. That is the exact "validate one path, use another" shape ADR-4650 names as the defect this epic exists to prevent, and it is the same bug this phase had just fixed in check-command-router. Committed inside the fix for it. All I/O now uses the resolved paths; user-facing messages still echo the raw filename, never a resolved absolute path. A VACUOUS TEST, and the false doc claim it was propping up. The test "[RED #4327] an absolute path outside the project is rejected" would have passed with ZERO containment logic: path.join(pendingDir, '/abs/outside/x') yields <pendingDir>/abs/outside/x — Node does not let a later absolute segment escape — so the name is FOLDED under the root, passes containment, and simply 404s. The test only ever observed "Todo not found". It now asserts what is actually true and actually valuable: an absolute name is neutralized, and the real outside file is not read, not moved, and still present afterward. docs/CLI-TOOLS.md claimed such a path "is rejected as a usage error", which was false; it now describes the fold-under-root behavior. Traversal and embedded separators ARE rejected, and those claims stand. DUPLICATION THIS EPIC EXISTS TO REMOVE. resolvePath already did isAbsolute-or-join + validatePath + reject; cmdGapAnalysisPlanPost and cmdCheckPredicate each re-inlined the identical triplet in the same file. Both now call resolvePath. Cost, stated rather than hidden: its generic message replaces the two sites' distinct "phase-dir escapes…" wording. The message still names the offending input, and one predicate with one message is the point. SYMLINK COVERAGE was required by #4652's "Done when" and was missing. Added for both the todos root and --phase-dir, skipping cleanly on EPERM so the Windows lanes do not fail where unprivileged symlink creation is disallowed. Both fast-check properties were UNSEEDED. Seeded now. The changeset named "check decision-coverage-plan" as a boundary; that is a caller of the shared resolvePath, which the body never mentioned. Corrected. DISCLOSED, not hidden: ctx.phaseDir is now always the resolved ABSOLUTE path, so ${PHASE_DIR} interpolation and the "not found in <targetDir>" message show an absolute value where a relative --phase-dir previously produced a relative one. That is an observable output change. A test pins it and docs/reference/gate-predicates.md states it. Also regenerated scripts/lib/platform-conformance-tier.generated.cjs and its macos twin — the new tests changed check-predicate.test.cjs's tier classification. Caught by npm run lint:ci locally rather than by a bench run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a2331c01f1 |
fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites Manually traced against the current tree: the validating regex at code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical behavior for 1- and 2-segment ids, character class unchanged): code-review.md, code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating sites, plus their comment/error-message text), execute-plan.md's plan-filename extraction, and plan-phase.md's --research-phase flag capture. Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor, which was matching an unrelated earlier `--research-phase` occurrence (line 77's generic-value capture) instead of the targeted site (line 131). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/ Adds findSingleSegmentPhaseRegexDrift, banning the bounded `[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any phase-carrying line across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned the same way as the existing shell-arith rule. Wired into scanAll; confirmed zero violations against the real tree post-#4568 fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4568): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file The emitted-attribution gate also flags 4 files growing: code-review-fix.md (+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9 bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each site's validation regex widened from a bounded single-optional-dotted-segment shape to the unbounded form, and the accompanying comment/error-message text grew by a few characters to mention the new 3-segment example. Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568) Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4568): backfill changeset pr number to 4646 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4d65c248e5 |
fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector Tests only, committed ahead of the implementation so the RED run is real. - tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative cases proving seam calls and path-call-plus-slash-literal are not platform signals; positive pins that genuine platform content, seam-bypassing spawns, chmod and symlink still classify in; macOS signal set and generated list unchanged. - tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest rows and test-conformance still has 3 windows + 1 macOS. - tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint test does, and resolveSelection rejects the retired windows scope. Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): delete the second Windows selector and narrow the conformance tier Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped conformance tier on real Windows/macOS — was not met. Measured on PR #4640 (run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7 changed test files running on a real Windows runner twice. Two selectors, only one in the epic's scope. The test job's three scope:windows shards predate the epic (#494, sharded #3057) and gate on product_changed, not full_matrix, so they fire on every product PR whatever Phase 3's classifier decides. They are deleted; test-conformance becomes the sole Windows selector, as it already was for macOS. Non-Linux jobs 7 -> 4. Gating the lane instead was rejected as provably redundant: for a test file reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file), and that same predicate sets full_matrix, which turns test-conformance on. Every file a gated lane would run is already covered in the same run. The lane's one non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix instead of feeding a parallel lane. Two detectors matched the repo's own test idiom rather than any platform signal and carried 226 of the tier's sole-signal membership against 41 for the other eight: process-seam-subprocess (335 files, 118 unique) matches the tests/helpers.cjs entry points nearly every CLI test uses, and going through the seam is the opposite of a platform signal since shell-command-projection takes platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs only a path call anywhere plus a slash literal anywhere, and that class is already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed. Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured. Adds the size gate Phase 2 never had, as a ratio against a live denominator so it cannot stop binding as the suite grows. 292 files leave real-OS Windows execution. The drop-out set was audited: 14 have a platform-suggestive filename and all 14 are static source-text analyses or seam-mediated CLI tests. raw-child-process was investigated as a suspected false negative and left unchanged -- relaxing it adds 13 files, all false positives. macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated macos-conformance-tier.generated.cjs is byte-identical at 196 files. Fixes #4641 Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): register the new ADR path in the docs-guard exempt baseline tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md in a comment justifying the retired windows scope; lint-docs-guard-registration tracks that reference set, so the baseline needs the new path. Verified the exemption still holds: the path is prose, not a filesystem read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): make the escalation tier-backed and drop every hardcoded count Three follow-ups from measuring the first pass rather than trusting it. The windows-hint escalation now requires tier membership as well as the filename hint. Setting full_matrix runs test-conformance, which runs only the tier; escalating on a test that is NOT in the tier costs four jobs and still never runs that test on Windows. Measured over the 16 RULES entries the narrowed predicate fires on exactly the same rules today, so this is correct-by-construction rather than a behavior change. The broader variant -- escalate on any tier member a rule pulls in, ignoring the hint -- was measured at 14/16 rules and rejected as over-broad. Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as soon as the rebase pulled in one new test file from #4253, which is the whole argument against them: the ceilings are ratios against a live denominator, the committed lists are pinned by comparison against a fresh classification of the live tree, and the three named probe files now assert on their SIGNAL rather than on membership in a literal list -- asserting by filename is the exact error this PR fixes in the classifier. Regenerates both lists against the rebased tree. Same-tree figures are now 547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none added; macOS is unchanged at 197 with a zero-line diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke An isolated adversarial review found a real false negative. Removing the blanket process-seam-subprocess detector also removed the only coverage for tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook spawns options.interpreter via real spawnSync, so runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary executing a shell script extracted from workflow markdown. The seam argument holds for src/shell-command-projection.cts, which takes platform as an injected parameter; it does NOT hold for the test helpers, which spawn real binaries. Conflating the two is what made the blanket detector look purely noisy -- it was 99% noise wrapping a real signal. Adds a narrow shell-interpreter-spawn category keyed on a real interpreter option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33% ceiling. All 9 confirmed by reading the matching source line, zero comment or fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured and rejected -- each adds 9 files but misses the counterexample entirely. Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD strictly rather than carrying the Windows report-only carve-out. The carve-out now lives solely on test-conformance, whose matrix does include windows. Repairs seven pre-existing assertions the category removal invalidated, preserving each case's purpose rather than deleting coverage, and converts the last hardcoded tier bounds to live-derived ratios -- including the macOS sanity range that was still a magic [100, 350]. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): keep the confinement test on a real OS via a documented allowlist A security review found tests/external-descriptor-confinement.test.cjs had dropped out of the Windows tier. It must stay in, and no content signal can express why: it exercises isPathConfined (src/external-descriptor-trust.cts), which uses the AMBIENT path module -- path.resolve(root, target) and path.sep -- with no injection. Its win32 semantics (drive letters, UNC, separator) are only reachable by actually running on Windows, and it is a security-relevant write-confinement gate. A content classifier cannot see 'this module reads the ambient path module', so no regex belongs here. Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows tier only. A Map rather than a list so an entry without a reason is impossible by construction, and tests assert every entry names a file that exists on disk so a stale entry fails loudly instead of rotting. This is the centrally- enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's portability-vocab.cjs already models -- deliberately not a heuristic. Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is untouched and byte-identical: the win32 concern does not apply to a POSIX runner, and a test asserts the allowlist does not leak into that tier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): inject the path impl into isPathConfined and correct the ADR count Two review findings, both fixed rather than dispositioned. A security review found tests/external-descriptor-confinement.test.cjs had left real-OS execution. The allowlist pinned it back, but that only restored INCIDENTAL coverage: isPathConfined used the ambient path module, and its test carried POSIX-only literals, so a win32 confinement escape was unverified on every platform including Windows. isPathConfined now takes an optional third parameter carrying the path implementation, defaulting to the ambient module. Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change is purely additive and every existing two-argument caller is byte-identical. Tests now inject path.win32 and path.posix, covering a different drive letter, a cross-drive absolute, backslash and forward-slash traversal, UNC, and the startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators. Proved load-bearing: dropping the + p.sep from the prefix check fails exactly the two boundary cases and nothing else. Callers' suites 149/149. The spec review caught an off-by-one: the ADR narrated a 264-file tier while the committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264 -> 265 (28.5%). Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated classify() as zeroing windows_tests, a key this change removes -- kept as historical narration but labelled as such. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#131): make the unwritable-HOME test actually test something Found by sweeping for the root-bypass class after fixing commit-files-deletion. This one is the silent variant, and it was broken twice over. First, the condition: the test made a fake HOME unwritable with chmod 0o500. The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed writable and the hostile condition never existed. Replaced with a HOME whose PARENT is a regular file, so every write under it fails ENOTDIR at the VFS layer for every uid -- no permission check is involved at all. Second, and more fundamental: the probe was npm --version, which on npm 11.19.0 performs zero filesystem I/O against HOME. Proven rather than assumed -- neutralizing runNpm()'s isolation turned the sibling test red while this one stayed green, so its assertion could never detect the regression it guards, on any uid, with or without the condition fix. npm config get cache was tried next and proved vacuous the same way (it only string-resolves the path). The probe is now npm cache verify, which really does mkdir _cacache under HOME. Re-proved load-bearing after the change: with isolation neutralized the test now fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was restored and verified diff-clean; suite 13/13. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the net drop-out figure in ADR-4641 The Consequences section still said 292 files leave real-OS Windows execution. That was the count before the narrow shell-interpreter-spawn replacement restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary categories rather than only raw-child-process, and clarifies that the 14-file filename audit was against the 292 initially dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the rejected concentration ceiling and its measurement Applying Goodhart's own question to the new ceiling -- how would you make this metric look good without improving what it represents -- surfaces a real weakness: a ratio can be satisfied by inflating the denominator, so adding OS-agnostic tests loosens it without narrowing the tier. The obvious companion gate was a sole-signal concentration ceiling, since the original defect was one detector carrying half the tier. Measured and rejected: peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the original defect; any threshold below it fails on a legitimate category. The discriminator is whether a signal is platform-meaningful, which no threshold encodes. Weakness disclosed rather than covered by a gate that does not bind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4641): add the changeset fragment for the confinement-check change changeset-lint failed on PR #4643: the PR touches user-facing paths and carried no fragment. The earlier no-changeset call matched #4604's CI-only precedent and was correct then; it was not revisited once the PR grew a src/ change, which is my miss. The fragment describes the real user-visible improvement: the external-descriptor write-confinement check's Windows semantics are now verified deterministically rather than only when the suite happened to run on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the tier count in TESTING-SUITES.md Said the tier narrowed from 546 to 254. The final committed list is 265 of 931 eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored 9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in the ADR, in a live reference page rather than a dated record, so it states the current truth rather than carrying an amendment note. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the measured aggregate from real CI job lists Epic #4589's closeout asserted its reduction from a static count; #4641's acceptance criterion asks for a figure read off a real run. Recorded here: test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4. Also states the caveat that a PR's total CHECK count is not a clean before/after comparison, since many gates are path-scoped and this change touches a broader path set -- the like-for-like figure is the test.yml job count. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): compare job totals the same way on both sides The measured-aggregate table put #4640's COMPLETED run total (21) against this run's count at matrix-expansion time (15). Those are not the same measurement: the completed total includes the post-test Coverage gate and baseline-publisher jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux jobs 7 -> 4, was correct and is unchanged. Called out in the table rather than silently corrected -- comparing two differently-derived numbers is exactly the error class this ADR is about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record measured conformance wall-clock and date the stale counterfactual Adds the per-job durations from both runs. The honest read is that this is a correctness win more than a speed one: file count fell 52% but wall-clock only 9-29%, because what was removed were the cheap static tests and what remains is concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects a future narrowing to buy time proportional to file count. The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about -- pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its tier was UNCHANGED at 197 files, which fixes that as runner variance and is noted as a caution against reading a single duration as signal. Also dates the symlink-keyword counterfactual, which cited a 254-file tier from before the replacement category and allowlist took it to its final 265. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero next gained #4644 mid-flight, so every absolute count shifted. Re-measured on the tree this actually ships against (932 eligible): 548 -> 257 by detector removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed, 9 restored. macOS 198, unchanged by this PR. The percentages did not move across three rebases (58.8% -> 28.5%), which is the whole argument for expressing the ceilings as ratios rather than counts -- noted in the ADR since it is now evidence rather than assertion. Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own win32 test cases introduced the literal win32 into the pinned file, so it classifies in on content via win32-darwin-literal. The entry stays and the reason is written down, because the file's real-OS need is a property of the code under test (isPathConfined reads the ambient path module), not of the test's text -- the text that currently saves it is incidental and could be refactored away silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
db4d8a9bae |
fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic $((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is valid shell-arithmetic syntax at all, and the failed expansion aborts the rest of the snippet in a non-interactive shell. safe_resume_gate runs unconditionally before trusting STATE.md or dispatching any executor, so execute-phase failed at its own gate before the first executor on any decimal phase, regardless of workflow.tdd_mode. Regression from #4194. Fixes all 4 sites: safe_resume_gate and the TDD gate in workflows/execute-phase.md, the completion-signal spot-check fallback in workflows/execute-phase/steps/completion-reconciliation.md, and the executor gate validation example in references/tdd.md. Each now zero-strips only the leading integer segment into a *_INT variable (via %%.* / # parameter expansion — always valid shell syntax regardless of what follows) and keeps the remainder as an escaped-dot string for the anchored commit- scope regex, exactly as issue #4619 verified in both bash and zsh. A plain integer phase (12, 01) computes byte-identically to before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug Behavioral coverage via real bash execution: the old $((10#01.1)) form throws (characterizes the bug, matching the issue's own reproduction); the new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):, feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring issue #4619's own verified table exactly. Updates safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions (one per site) to the new fixed text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic With #4619's fix in place, the guard's original "ban $((10#... outright, match any occurrence" was too blunt: it flagged a comment merely mentioning the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an already-%%.*-stripped integer, and the always-safe plan-id arithmetic (plan ids are plain integers, never decimal). Refines the detector to skip full-line comments and to only flag a captured variable/placeholder name that contains "phase" and does NOT end in _INT/_int — the naming convention the #4619 fix establishes at all four sites for "already reduced to a safe integer." A plan-id variable was never phase-number arithmetic in the first place and is excluded on the same basis. This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new exemptions") and D7 ("a decimal and N-segment phase id survive an end-to-end execute-phase selection without error") for real — the guard now reports zero violations across all five .cts/.md rules. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): cover the plain-padded-integer near-miss matrix too Review found the anchored-ERE near-miss coverage only exercised the decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also verifies the plain padded-integer case (01 -> PHASE_N=1) against its own near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the missing assertion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote only 2 backslash characters in JS source, which single-quoted-string parsing collapses to 1 real backslash at runtime -- but the workflow/reference files actually contain 2 raw backslash bytes at that position (needed so bash's ${var//pattern/replacement} produces the correct single-backslash output). Write 4 backslash characters in the JS source at all 4 occurrences so the runtime string matches the files' real bytes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): refresh the committed compact-content benchmark baseline The new PHASE_INT/PHASE_FRAC arithmetic lines added to gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): note the safe_resume_gate arithmetic growth in the test header The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846 bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading integer segment via base-10 arithmetic instead of forcing the whole value through $((10#...)) and hitting a hard shell syntax error on the first dot. A blank line previously separated the Emitted-Drift-Ack-Growth trailer from the Co-Authored-By trailer below it, which splits git's trailer-block detection: only the last contiguous non-blank run of Key: Value lines at the end of a commit message is recognized as trailers, so the growth ack was silently read as ordinary body text and the differential-attribution gate failed with the growth unacknowledged. Joining the two trailers into one contiguous block fixes it. Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4208): replace chmod-based restore-failure injection with a root-proof git shim `tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated an unwritable index via a `post-index-change` hook running `chmod a-w` on the git dir. That relies on the OS enforcing the *owner's own* permission bits against itself, which uid 0 (a routine identity inside this repo's Docker-based gsd-test benches) does not: every DAC check short-circuits true for root, so the write the chmod meant to block silently succeeds, the restore comes back clean, and the disclosure/rollback behavior under test never actually gets exercised. This is CLAUDE.md's own named anti-pattern for I/O-failure injection ("Cross-platform test IO-failure injection" — chmod tricks fail under root Docker/CI). It is confirmed as the actual root cause here, not a production defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure logic (added by #4253, merged just before this run) was hand-traced and manually reproduced end to end on an unprivileged workstation against a freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces exactly the `staging_failed` + "could not be restored" / "could NOT be restored during rollback" results both tests assert. The other `post-index-change`-based tests in this file (a `sleep` to force a timeout; a real `update-index` to flip a restored entry's mode) are unaffected because neither depends on a permission check — consistent with only the two chmod-based tests failing on the real remote run. Replaces the chmod fixture with a fake `git` placed ahead of the real one on PATH that fails only `update-index --add --cacheinfo` — the one call the restore makes — unconditionally, regardless of privilege level. Every other git invocation execs straight through to the real binary, so the rest of each scenario (`rm --cached`, the restore's own `ls-files` verification, etc.) is exercised exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): backfill changeset pr number to 4644 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI Passing the script as a `-c "<script>"` argv element made it subject to Windows' CreateProcess command-line argument encoding, which silently dropped the escaped-dot backslashes before bash ever saw them (observed on PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the same script via stdin instead removes argv entirely from the transport, so there is nothing for Windows to re-encode. POSIX behavior is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4cc2a466b5 |
fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec `cmdCommit`'s `--files` list can stage an addition but never a deletion: the #2014 guard skips a missing explicit entry because the filesystem cannot tell "moved away" from "not written yet". A caller that moves a file therefore had two forms, both wrong — a directory entry records the move but also commits every unrelated file in that directory (a concurrent session's in-flight todo, in the unattended execute-phase sweep), and a file entry leaves the old path's deletion dangling with the todo tracked at both paths. `--files-removed <paths>` is the caller-declared delete intent. Each entry names a file, or a directory whose tracked-but-absent files are the removals; those paths are staged with `git rm --cached` and join the commit pathspec. `--files` keeps its skip-if-missing contract untouched. A file entry still present on disk fails the commit closed with the existing staging-failure rollback; a never-tracked path is a no-op. `--files-removed` alone is a declared scope, not the unscoped .planning/ sweep. The dispatcher previously folded every non-flag token after `--files` into that list, so a second list flag could not exist; each list now runs from its flag to the next `--` token. The execute-phase todo sweep names the moved todos on both sides from CLOSED[@], and cleanup's archive commit moves .planning/phases/ and .planning/quick/ under --files-removed. Fixes #4208 Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries * chore(#4208): set changeset fragment pr to 4253 * fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard Three CI failures, all consequences of this PR's own change. 1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The three-line rationale comment plus the four-line array-building block added 318 bytes to a file that had only 141 of headroom on next. Move the rationale to docs/CLI-TOOLS.md -- which this PR already extends with the --files-removed contract, and which is where the ADR-857 gate wants call-site detail to live rather than in the host workflow -- and fold the array build onto one line. 93,577 -> 93,372. 2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415 guarantee to its old MECHANISM: it regex-matched the literal .planning/todos/{completed,pending}/ directory pathspecs in the commit --files list. This PR deliberately replaced those with named files (a directory entry also committed an unrelated todo a concurrent session dropped in mid-close), so the guard failed on a change it should have accepted. Re-point it at the new mechanism without weakening it: assert the ADDED array reaches --files, the REMOVED array reaches --files-removed, STATE.md is still committed, and -- newly -- that the two arrays are built from $COMPLETED_DIR and $PENDING_DIR respectively. Verified by negative control: deleting --files-removed "${REMOVED[@]}" from the workflow still fails the test, so the #2415 regression remains caught. Note for the merge queue: #4233 also grows execute-phase.md (+114). The two are additive -- different regions, no textual conflict -- so with both landed the file reaches ~93,486, over the 93,400 margin though under the 93,600 hard ceiling. Whichever merges second will need to reclaim ~86 bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv * fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin Rebasing onto next surfaced the byte-gate collision flagged earlier on this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354), so this PR's +113 landed at 93,467 against the <= 93,400 margin in tests/claude-orchestration.test.cjs. Compact the close_phase_todos step this PR already edits -- drop the PHASE_NUM indirection, fold the normaliser and the match guard, print the closed list with one printf, shorten the step's prose -- without touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no longer spends any of next's 46 bytes of headroom. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback Review of #4253 found three Majors with one root cause: the removal side judged presence by fs.lstatSync alone, where the addition side already reads `git ls-files -v` state. Absence from the worktree is not removal: - a submodule gitlink (mode 160000) whose directory was deleted by hand lists like a file and was `rm --cached` with no .gitmodules cleanup; - a skip-worktree path is never materialised by a cone-mode sparse checkout, so a directory entry over a sparse-excluded tree dropped that whole tree from the index; - an assume-unchanged path's worktree state is not something git itself consults; - an intent-to-add entry (`git add -N`) renders as a plain cached entry on the empty blob, yet nothing tracked exists to remove and no rollback can restore the flag. The index listing now carries each entry's `ls-files -v -s` tag, mode and stage. Only a plain cached (H), stage-0, non-gitlink entry is a removal candidate; every other state is left alone under a directory entry (exactly like a present file) and fails closed when named directly, with the state in the error. "Named directly" is decided on RESOLVED paths, not strings -- realpath of the longest existing prefix with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a symlinked spelling of the tree (macOS `/var` -> `/private/var`, where `process.cwd()` is the real path and the caller's absolute path is not -- CI on this round's first push) all resolve to the same entry, where a string compare against git's cwd-relative output silently took the directory polarity (pre-push review, driven; the symlink case is driven with an aliased fixture directory). The enumeration's domain is what `ls-files -v -s` can emit for an index entry, stated at the classifier. The third Major -- on an unborn HEAD a successful `rm --cached` was never rolled back when a later entry failed -- is fixed differently from the review's suggestion. Pushing the path into stagedPaths would put it on the commit pathspec, which a root commit refuses ("pathspec did not match", driven), and `git reset -- <path>` cannot restore an entry with no HEAD anyway. Instead every index entry this call removes is recorded (mode, blob) before the `rm` and put back with `update-index --cacheinfo` on rollback. That also restores a caller-pre-staged blob at a removed path exactly, where a reset would have silently replaced it with HEAD's version. The rollback is best-effort, as the addition-side reset already was, and the docs say so. Eight tests: gitlink under a directory entry, named directly, and named by absolute path; skip-worktree both forms; intent-to-add both forms; assume-unchanged named; unborn-HEAD partial failure restores the removal; pre-staged blob survives the rollback. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step Review nit on #4253: inserting the --files-removed rationale between the original bash block and its closing fence left an empty ```bash``` pair before </step>. Harmless at runtime, a formatting artifact of this PR's own diff; removed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): a boolean flag inside a commit path list no longer ends the list Review minor on #4253: collectList stopped at the next `--` token, so a positional wedged between a boolean flag and the next list flag (`--files a --amend b --files-removed c`) was claimed by neither list and silently dropped -- a regression in shape against the old slice-to-end parse, which filtered `--` tokens and kept `b`. No current call site interleaves that way, but the gap was real. A list now runs to the next LIST flag (`--files` / `--files-removed`) and skips boolean flags on the way, and a REPEATED list flag merges its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse did -- a first cut stopped at the repeat and dropped `b`, the same silent-drop shape one level over (pre-post comment audit). The only change #4208 makes to parsing is that a second list flag can exist. Tests: STATE.md wedged between --no-verify and --files-removed lands in the commit; both runs of a repeated --files reach it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * test(#4208): drive the reappearance window with a post-index-change hook Review nit on #4253: the defensive re-check for a file recreated between the absence test and `git rm --cached` -- the concurrent-session race this PR's own changeset names -- had no test. git fires post-index-change the moment `rm --cached` writes the index, so a hook that copies the file back exactly then exercises the window deterministically. The call reports staging_failed / "reappeared on disk", commits nothing, and the rollback restores the removed entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm * fix(#4208): restore a staged removal when the call records nothing A `git rm --cached` that succeeds mutates the index whether or not a commit follows. Only the staging-failure rollback put those entries back, so a call that reached `nothing_to_commit` reported no state change while the removal sat staged -- riding along on the caller's next commit. The review named the unborn-HEAD, removal-only shape. Keying on `headExists` would have fixed half of it: the guard also fires with a real HEAD when the removed path is index-only (added, never committed), because `diff HEAD` reads clean with the path absent on both sides. Both shapes now restore, at both `nothing_to_commit` exits. The failure exits are deliberately left alone -- they report a failure rather than no-change, and the addition side leaves its own staged paths there too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot `cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)` now owns the index-state classification, path canonicalisation and entry recording, returning the pathspec entries and the recorded removals its caller merges. Pure motion: no branch, message or probe changed. Only the two accumulators became local names, and `restoreRemovedEntries` stays with the caller because the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the extracted helper is 277. (Figures corrected after publication: an earlier version of this message said 854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit is 581, so this is above it, not below.) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): property-test the two-list commit parser RULESET.TESTS.property-based-testing asks a parser for at least one property test asserting a domain invariant; `collectList` had only hand-picked examples, one per shape a review round had already broken. Hoisted it to module scope as `collectListFlagValues` and exported it in the file's existing exported-for-tests convention -- a parser reachable only by spawning the CLI can be tested one example at a time and no faster. Three properties over generated argv: every positional lands in exactly the run open at it whatever the flag order or count; no positional after the first list flag is dropped or double-claimed; and with `--files-removed` absent the parse equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run ending at any `--` token, and a repeated list flag that does not merge -- each of which the properties catch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): pin cleanup.md's archive commit to --files-removed execute-phase.md's rewrite is pinned by the #2415 guard in this file; cleanup.md's equivalent was not, so reverting its routing would have been caught by nothing -- the mechanism's unit tests never read this file and pass either way. Asserts the two archived directories are under --files-removed and NOT under --files (where a directory entry sweeps in a concurrent session's in-flight writes), and that the destinations and STATE.md stay on the additive half. Controlled by restoring the pre-#4208 sweep, which fails it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): pin that a symlink to a directory is one tracked path Review of #4253 read the `lstatSync(...).isDirectory()` test as a symlink-following defect. Driving it says the opposite: git tracks the link as a single blob (mode 120000) and does not traverse it, so the tracked paths "under" it live at the real directory and were never named by the caller. Following the link would stage those -- the directory sweep #4208 exists to remove -- while the named entry still sat present on disk. Pinned rather than changed, with the premise driven in the test body. Swapping `lstatSync` for `statSync` -- the prescription as written -- fails it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit The base range added `tests/benchmark-compact-content.test.cjs` and a committed token baseline over the compacted workflows. This PR edits `gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on that entry and on the aggregate. Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is those two entries and nothing else. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): report a removal the call could not put back Round review of this round found the restore itself unchecked: the helper ignored `update-index`'s exit code, so a FAILED restore still reported `nothing_to_commit` -- the same false "no state changed" the restore exists to prevent, surviving one level down on the restore-failure path. It now returns a boolean. The two no-change exits report `staging_failed` naming the paths left staged; the staging-failure rollback still ignores it, deliberately, because it is already reporting a failure and an unwritable index is usually the failure being reported. Driven with a post-index-change hook that makes the git dir unwritable the moment `rm --cached` lands, so the restore cannot take its lock. Reverting both guards fails the test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): disclose a removal the rollback could not restore Round review refuted the reasoning behind leaving the rollback path's restore unchecked. The claim was that this exit is already reporting a failure, so the restore's result adds nothing. The counterexample is the ordinary case: the reported failure is usually a DIFFERENT cause -- a contradictory declaration, a reappeared path -- so a caller reading `failures` sees only that cause and learns nothing about the removal still sitting in its index. The rollback now appends a disclosure entry per un-restored removal, naming the path. The reason and `file` still report the failure that caused the rollback; the disclosure is additive. Also moves the restore-failure test's chmod into a `finally`: `t.after` runs AFTER the parent `afterEach`, so a throw before it left the fixture undeletable. Both driven; reverting the disclosure fails the new test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): decide index state by observation, never by an exit code The restore added two commits earlier keyed both its record decision and its success verdict on git's exit code. An exit code answers "did the command succeed", never "did the index change" -- execGit collapses a spawn timeout to a non-zero exit, and a killed git can already have written the index. Round review drove four failures from that one assumption, in both directions: - a failed `rm` still contributed an entry, so the rollback disclosed a removal that was never staged (stale index.lock); - a timed-out `rm` whose write DID land contributed none, so a real mutation was neither restored nor disclosed; - a timed-out `update-index` whose write landed reported failure, publishing a "could NOT be restored" disclosure that was false; - and the read-back that replaced it omitted `-z`, so core.quotePath rendered `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not restored -- the same quoting defect this PR already fixed for `preStaged`. Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the path: gone means this call owns the removal and records it; still there means nothing was staged; a probe that cannot answer becomes its own failure entry rather than an assumption. The restore verifies the same way, comparing the WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a path-only test accepts an entry that came back as something else. The verdict is three-valued -- `restored` / `not-restored` / `unverified` -- and the unverified wording says the restore could not be VERIFIED rather than that it failed. The rm's own failure is pushed ahead of any probe diagnostic so a timed-out removal keeps `timed_out: true` and its own message as the reported cause. Five regression cases, each negative-controlled against the shape it pins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): treat a declared removal path as a path, not a pathspec An index path handed back to git is parsed as a PATHSPEC, and the removal side handed several back. Three driven harms, all of them the sweep-in this flag exists to remove, arriving through the operand rather than through a directory entry: - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it removed `peer.md` and `stays.md` too, only the declared entry was recorded, so the rollback restored one of three and the other two rode out as staged deletions the result disclosed nowhere; - the same name reached `git commit -- <paths>`, which globbed and committed an undeclared `M peer.md` alongside the declared removal; - and the intent-to-add probe (`diff --cached` over the path) matched a STAGED PEER instead of itself, so an `add -N` entry was misclassified as ordinary content, removed, and restored by `--cacheinfo` -- which cannot restore the intent flag. It came back as a real staged addition. Every operand on this path is now `:(literal)`: the `rm`, both index probes, the intent-to-add probe, the restore read-back, the entry-level `ls-files` / `ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream `ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep whatever pathspec behaviour they have today; that is not this change's to alter. `:(literal)` still resolves a directory to its descendants (driven), so the directory form is unchanged. Closes what an earlier cut of this commit declared as a residual: a filename beginning with `:` is now removable end to end, because the commit pathspec no longer reinterprets it. Also fixes a MINOR from the same review: cleanup.md's contract test checked the destinations' position relative to `--files-removed` but never that `--files` was present at all, so deleting the flag still passed. Un-literalising the seven sites fails three of the new tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * fix(#4208): scope the rollback to the caller's own name space Round review drove a rollback that destroyed the caller's own staged work. Two causes, one of them pre-existing: - `git diff --cached` prints REPO-relative paths whatever the cwd, while `stagedPaths` holds the caller's cwd-relative names. In a project nested inside its repo (`<repo>/sub/.planning/...`) the two name spaces never intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`, and the reset unstaged a caller-staged deletion and modification that this call had never touched. `--relative` makes the two sets comparable, and is a no-op when the project IS the repo root. This governs the `--files` side too and predates this flag. - the rollback's `reset` was the last place a removal-derived name reached git as a bare pathspec; it takes `asPathspec` like every other site. Driven on a nested fixture; dropping `--relative` fails the new test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): gate six fixtures that Windows cannot construct CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two primitives the new fixtures rely on do not exist on Windows, both driven on a real Windows host rather than inferred: - a filename containing `*` or `:` cannot be created at all (`IOException` / `FileNotFoundException`), which is four of the pathspec fixtures; - `chmod` cannot make a directory unwritable — a write into a ReadOnly directory succeeds — so the two restore-failure fixtures cannot drive the failure they exist to drive. Each is skipped on win32 with its measured reason, in the repo's existing `{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The behaviours they pin are platform-independent; only the fixtures are not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * test(#4208): build git's index-syntax path with forward slashes The remaining Windows red was mine, not the platform's: `git rev-parse :<path>` takes a forward-slash path, and `path.join` yields backslashes there, so git rejected it as an ambiguous argument. The hook in the same test already used the slash form. Not gated — the behaviour it pins is portable; only the argument was not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj * chore(#4208): refresh the compact-content baseline against the rebased base `next` moved the `new-project` split and the aggregate under this PR's execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs --write` so the only leaves differing from the base's copy are the execute-phase split and the aggregate it feeds. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh * chore(#4208): regenerate the macOS conformance tier for this PR's fixtures `next` gained the macOS-specific conformance tier (#4593) after this branch was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target macos`) now selects `tests/commit-files-deletion.test.cjs` on the `chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the chmod-driven failed-restore cases and the symlink-to-directory case). Regenerated with `--target macos --write`; the platform tier was already in sync. The file was modified, not added, which is why the added-files check did not surface it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1e47560e34 |
feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the same 546-file, Windows-oriented conformance-tier list as windows-latest -- built from signals like windows-shell-token/windows-env-var that have nothing to do with macOS. Issue #4593 asked for macOS coverage sized to its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin- specific behavior) instead. Issue #4593 was filed before Phase 5 (#4603) existed and referenced updating test-full's macOS legs -- that job is gone. Corrected the issue's body before any code was touched: the "shrink from full replay" half of the original ask was already done by Phase 5; what remained was narrowing the still-Windows-oriented tier macOS was inheriting. Two design assumptions were measured and rejected before accepting a design (documented in docs/adr/4593-macos-conformance-tier-architecture.md): - Reusing the general tier's signals minus its 3 Windows-specific categories barely narrows anything (546 -> 424, 78% retained) -- most files match multiple signals and only need one to survive exclusion. - A standalone CRLF/autocrlf signal, despite the issue naming "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit 143/930 files. Root cause: CRLF is primarily a Windows checkout concern in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST- PORTABILITY), so the signal was really re-selecting Windows-relevant files already covered by the general tier, not narrowing macOS specifically. Built 5 new, genuinely macOS-specific signals instead: darwin-literal (darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch, case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim from the general tier (genuinely Unix-relevant, not Windows-motivated). Measured against the real tree: 196 of 930 eligible unit-suite files (21%), versus the general tier's 546 (59%) -- a real, evidence-backed narrowing. scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/ classifyMacosTree/renderMacosGeneratedFile and a --target windows (default, unchanged)/--target macos CLI flag, so the same generator produces two independent, gated outputs rather than needing a second script. New committed output: scripts/lib/macos-conformance-tier. generated.cjs. .github/workflows/test.yml's test-conformance job: only the macos-latest leg's file-list source changes; windows-latest is byte-for-byte untouched. New shipped-file ripples handled proactively (19 install-tree fixtures regenerated, bin/install.js registered). An isolated code-review pass found one real defect: the ADR's per- category count table had drifted by 1 (zsh-dispatch, case-sensitivity) because the new test file's own fixture strings joined the tree it classifies after the table was authored -- fixed, with the union total (196, what CI actually gates on) confirmed unaffected. An isolated security-review pass found no qualifying findings. The ADR also records an explicit requirement for any future widening proposal: check whether the motivating regression is already covered by Phase 1's no-rendered-text-length-assert lint rule (#4590) before re-proposing full macOS/Linux parity, since that is exactly what #4421's root cause was (a rendered-text-length assertion, not a real behavioral divergence). Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bcd99696d3 | chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) | ||
|
|
615b74ff45 |
fix(#4460): correct two stale changesets left by an admin-merge race (#4572)
* fix: address orthogonal-review findings on the new work in this PR Isolated code-review + security-review of everything added to this PR since its original review (hono override, check-env.cjs rewrite/revert, new lib file, its test, installer enumeration). Security review: clean, no findings. Code review found: - BLOCKER: .changeset/silly-hens-relax.md described a hono override this PR no longer actually makes -- PR #4560 landed the identical fix on next first, and this branch's own hono commit became a genuine no-op the moment it was rebased onto that updated next (git diff origin/next -- package.json package-lock.json is empty). Deleted the orphaned changeset; next already carries #4560's equivalent one (.changeset/zesty-seals-click.md). - HIGH: .changeset/tame-hens-jump.md's body still described the execNpm-routing approach that was tried and reverted -- stale text from before that revert, would have shipped a release note for code that isn't actually in the diff. Rewritten to describe what actually shipped (self-contained spawnSync, 15s timeout, accurate ENOENT vs. timeout vs. non-zero-exit diagnosis). - LOW: no comment explaining why the spawnSync call has no try/catch (safe -- its documented contract routes failures through the returned result, never a throw -- but worth stating given this file's whole purpose is graceful degradation). Added one. - nit: exitCode 0 + empty stdout fell through to "npm binary not found on PATH", misdescribing a real npm binary that simply printed nothing. Gave it its own message; updated the corresponding test. Manually re-verified describeNpmVersionCheckFailure's branches and the real check:env success path before re-running gsd-test, since this repo blocks local node --test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: rest of the orthogonal-review fixes (previous commit only caught the deletion) Tooling mistake in the previous commit: a git add with the already-staged deleted changeset mixed into the same pathspec list errored out and silently skipped staging the other four files, so only the changeset deletion actually committed. This commit carries the rest of that same change: tame-hens-jump.md's rewritten body, check-env.cjs's no-try/catch comment, npm-version-check-diagnosis.cjs's exitCode-0-empty-stdout fix, and the corresponding test update. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4460): fix changeset pr field to point at this PR, not the original .changeset/tame-hens-jump.md's pr field still said 4552 (the PR its original text was authored under), but this PR (#4572) is what's actually landing the corrected body -- changeset-lint's own DEFECT.CHANGESET-PR-FIELD-DRIFT check caught it: "pr: 4552, expected pr: 4572". Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8731a90be6 |
fix: revert execNpm import in check-env.cjs, keep the diagnosis improvement
The execNpm-routing redesign (previous commit) broke every real CI job:
check-env.cjs runs as its own standalone "Environment check" step BEFORE
`npm ci` / `npm run build:lib` -- a deliberate pre-flight, run before
there is even a node_modules to build with. Its require of
../gsd-core/bin/lib/shell-command-projection.cjs (a tsc-compiled artifact
that plain does not exist at that point in the pipeline) crashed with
MODULE_NOT_FOUND on every platform, immediately, confirmed via the real
CI log. My own local gsd-test run never caught this because it doesn't
replicate that exact pre-build step ordering.
Reverted the cross-module require entirely; check-env.cjs is back to a
self-contained spawnSync(npmCmd, ...) call, no requires reaching into
gsd-core/bin/lib. Kept the two things actually worth keeping from that
detour:
- the 15_000ms timeout (matches execNpm's own default elsewhere in this
repo -- not invented, an existing precedent -- vs. the original 10s
that failed twice under real Windows CI contention);
- computing `timedOut` via `error.code === 'ETIMEDOUT'` inline (the
same canonical, cross-platform-correct predicate that seam uses),
rather than the earlier signal === 'SIGTERM' check, which that seam's
own docstring documents as platform-fragile with a Windows-specific
false-negative risk.
Manually verified check-env.cjs runs correctly with gsd-core/bin/lib
temporarily removed entirely (simulating the real pre-npm-ci CI
ordering) before re-running gsd-test, since this repo blocks local
node --test.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
1cd17cc3d6 |
fix: route check-env.cjs's npm-version check through the canonical execNpm seam
Per /research + /diagnose direction: the timeout fix landed earlier this
session correctly diagnosed the failure (a real spawnSync timeout under
Windows CI contention, not npm being absent) but it recurred on the very
next push -- same chunk, same ~51-file load. Rather than raise the
hand-rolled 10s timeout myself (CLAUDE.md's own rule: fix the cost, not
the tolerance, and never touch a timeout without explicit instruction),
investigated the repo's own precedent first.
Found: this repo already has a canonical OS-shell-projection seam for
exactly this (src/shell-command-projection.cts's execNpm), already used
by dozens of other scripts/*.cjs files (require('../gsd-core/bin/lib/...')
is an extremely well-established pattern), with:
- the same npm.cmd/shell:true Windows handling check-env.cjs was
hand-rolling, but centralized;
- a 15s default timeout (vs. check-env.cjs's 10s) -- not invented here,
an EXISTING value already governing npm subprocess calls elsewhere;
- isSpawnTimeout / result.timedOut, the canonical cross-platform timeout
predicate (error.code === 'ETIMEDOUT'), whose own docstring explicitly
warns that checking signal === 'SIGTERM' (what my first fix did) is
"platform-fragile" with a Windows-specific false-negative risk -- the
exact platform this bug lives on.
check-env.cjs's npm-version check now calls execNpm(['--version']) instead
of hand-rolling spawnSync + npmCmd + shell:true, and
describeNpmVersionCheckFailure now operates on execNpm's SpawnResultOutput
shape (using timedOut, not signal) rather than a raw spawnSync result.
This is a genuine architectural fix, not just a bigger number: it removes
a duplicate, slightly-divergent re-implementation of an existing seam and
inherits whatever that seam's timeout/handling becomes in the future.
Manually verified end-to-end (npm run check:env against the real
environment) and re-verified describeNpmVersionCheckFailure's branches
directly against execNpm's actual return shape before wiring the test
file, since this repo blocks local node --test.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
0aeafc6425 |
fix: report the real cause when check-env.cjs's npm-version check fails
Discovered blocking this PR's own Windows CI (unrelated to this PR's actual diff, fixed inline per this repo's no-defer policy): PR #4552's "full test (windows-latest, 24, shard 2/3)" job failed tests/check-env.test.cjs's npm-version subtest with "npm binary not found on PATH" under chunk 5/9's heavy load (51 concurrent files, 5+ minutes). Root-caused via the CI log: the check's spawnSync call used a 10s timeout, and every failure mode -- ENOENT, a signal-killed timeout, a non-zero exit, a thrown spawn error -- collapsed into that one message (only `res.status === 0 && res.stdout` was checked), so a genuine npm.cmd cold-start timeout under contention was indistinguishable from npm actually being absent. Extracted the reason-selection into describeNpmVersionCheckFailure, a pure function in the new scripts/lib/npm-version-check-diagnosis.cjs (kept out of check-env.cjs itself, which runs its CLI unconditionally on require with no `require.main === module` guard, so the pure logic can be unit-tested without triggering a real environment check). Reports ENOENT, signal-kill, non-zero-exit, and thrown-error cases distinctly. Does NOT raise the 10s timeout itself -- a slow subprocess under contention is a cost to reduce, not a tolerance to widen. Also fixed a stale tsconfig.build.tsbuildinfo incremental-build cache discovered while verifying this change: npm run build:lib was silently omitting gsd-core/bin/lib/markdown-table.cjs (a real, needed compiled module -- src/state-document.cts requires it), which only surfaced via npm run lint:generated-sync's gen-health-docs check failing with Cannot find module. Deleting the cache and rebuilding fresh restored it; docs/INVENTORY-MANIFEST.json needed no net change once the build was genuinely complete. Manually verified describeNpmVersionCheckFailure's five branches directly (ENOENT, signal-kill, non-zero exit, thrown error, defensive default) before wiring the test file, since this repo blocks local node --test. Re-ran node scripts/check-env.cjs directly to confirm the real success path is unaffected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
cad70f4f3e |
fix(#4120): replace shellcheck npm dep with dependency-free downloader (#4121)
* fix(#4120): replace shellcheck npm dep with dependency-free downloader The `shellcheck` devDependency (added in #4109) pulled in decompress@4.2.1 for archive extraction, which carries an unpatched CRITICAL zip-slip vulnerability (GHSA-mp2f-45pm-3cg9, CVSS 9.1) plus two moderate findings. decompress's latest published version IS the vulnerable one -- no patched release exists upstream, so npm audit fix cannot resolve this by upgrading. Removes the shellcheck package entirely and replaces its role with scripts/lib/shellcheck-fetch.cjs: a small downloader using only Node's built-in https/zlib plus a hand-written tar-entry reader, fetching a pinned koalaman/shellcheck release directly from GitHub releases. The reader never uses an archive-supplied name as a filesystem path (the exact defect class decompress had) -- it only returns the matched entry's bytes; the caller writes those bytes to a path it constructs itself. Bounds the download with a 30s-per-hop timeout, consistent with the ShellCheck subprocess's own timeout. Covers linux/darwin on x86_64/aarch64, matching this repo's actual CI (lint-tests runs only on ubuntu-latest) and local dev needs; Windows fails with a clear, honest error rather than silently misbehaving. Adds tests/lint-workflow-shellcheck-fetch.test.cjs covering the tar-parser (unit cases plus a fast-check property test per CLAUDE.md's parser-testing requirement), a security behavioral pin confirming traversal-style entry names are treated as opaque strings never filesystem paths, and boundary coverage for the redirect-following logic's MAX_REDIRECTS limit (limit-1/limit/limit+1, via an injectable transport, no real network I/O). npm audit: 0 vulnerabilities (was 1 critical + 5 moderate). The lint script reproduces the identical result against the current tree: "212 pre-existing finding(s) from baseline, 0 new" -- no behavior regression, no baseline changes needed. * fix(#4120): register shellcheck-fetch.cjs with the installer scripts/lib/shellcheck-fetch.cjs shipped without being added to GSD_SCRIPTS_LIB_FILES in bin/install.js, which would have left it orphaned on uninstall and broken the golden install-tree fixtures for every runtime. Adds the entry and regenerates the 19 affected fixtures via npm run gen:install-tree. * docs(#4120): add changeset for the decompress CVE fix --------- Co-authored-by: sim <sim@local> |
||
|
|
370cfc6680 |
enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local> |
||
|
|
ac0eed1267 | Merge pull request #4015 from open-gsd/fix/3889-instrument-chunk-timeout | ||
|
|
bf8905fcc9 |
fix(#4012): a hang never emits the events I was listening for
The init marker settled it. The diagnostic reported "THE REPORTER LOADED BUT RECORDED NO TEST EVENTS — the events file contains only the reporter's own reporter:init marker", which refutes the reporter-never-loaded hypothesis and leaves exactly one explanation. The runner spawns a child process per test file and surfaces a subtest's test:start / test:pass / test:fail to the parent's reporter only once the child REPORTS that test — which happens when it completes. The fixture hangs forever, so it never completes, so it never reports. I had recorded exactly those three event types: the precise set that a hang guarantees you will never see. The feature could not have worked for the case it was built for. test:enqueue and test:dequeue are emitted by the runner as it queues and begins each file, independent of anything inside finishing. test:dequeue is what actually means "in flight", and it is now the primary signal, with test:start kept as a secondary one. A file is in flight when it has been dequeued and has no terminal event. The four branches now describe states that are all real: the events file absent (reporter never loaded); the init marker alone (the runner dequeued nothing at all — genuinely surprising now rather than the expected outcome); everything dequeued and terminated (the files finished and the process hung afterwards, a handle leak); and one or more dequeued-but-unterminated files, named, which is the case this whole feature exists to report. Verified against the exact shape the real hang produces, by executing analyzeChunkEvents on a synthetic events file: init + enqueue + dequeue with no terminal event reports hangs.test.cjs as in flight, and appending a test:pass clears it. Four more unit tests cover the ordering and multi-file cases with no subprocess, so this logic is now checkable without a runner round-trip — which matters, because every defect in this feature so far was visible only remotely. T1 is untouched and should now pass for the right reason. Verification runs on the remote runner. Refs #4012 |
||
|
|
444137e63a |
fix(#4012): make the artifact say whether the reporter ever loaded
Down to 3 remote failures, all one chain. The explicit no-events reporting is working — the diagnostic now states the events file does not exist, instead of silently printing the generic message. But it then ASSERTED a cause: "the child was killed before the reporter wrote even one event (process/spawn startup stall, not a test hang)". That was a guess dressed as a finding, and the fixture contradicts it: it starts a real test, so test:start should fire in milliseconds against a 2000ms budget. Two hypotheses remained and I could not separate them locally, because the local test runner is hook-blocked here: either the custom reporter never LOADS in the child, or it loads and no event reaches it before the SIGKILL. Rather than guess a third time, the artifact now answers it. The reporter appends a reporter:init line as its first action, before consuming anything, so the file's contents discriminate: absent means the reporter never loaded; init-only means it loaded and saw no test events; init plus events means it works. The diagnostic has a branch for each and, where the cause is genuinely unresolved, names both possibilities instead of picking one. Also passes the reporter as a file:// URL via pathToFileURL. Node documents the --test-reporter value as an import()-style specifier, and a bare absolute path is not a portable one — notably on Windows. That is a correctness fix whichever hypothesis holds, and it is a live candidate for the first. FIXED_OVERHEAD is derived by reducing over the actual argv strings, so the longer URL is accounted automatically. Verified by execution, not assumption: composing the reporter against an EMPTY event stream writes exactly one line, the init marker. That is the whole point of the marker, so it is pinned by a test rather than left to inspection. T1 stays red and untouched. Verification runs on the remote runner. Refs #4012 |
||
|
|
5504724d7a |
fix(#4012): the reporter body must return nully, not an iterable
Fourth failure on this feature, and this one was caused by the previous fix's lint workaround. TypeError [ERR_INVALID_RETURN_VALUE]: Expected nully to be returned from the "body" function but got an instance of Array. When stream.compose is given an async FUNCTION as the body, that function must return nully. The reporter ended with 'return []' under a comment asserting Node "still requires the exported function to return an iterable" — exactly backwards, and the direct cause of 41 failures across every run-tests.cjs invocation. That return existed only to dodge ESLint's require-yield after the previous commit converted async function* to async function. A lint workaround became a runtime crash, and the comment written to justify it stated the opposite of the contract. Both are now corrected to what the runtime actually does. Verified by EXECUTION rather than by reading: composing the real reporter against a fake event stream completes with no error and leaves both handled events durably on disk, with the ignored event type skipped. The failing form was reproduced the same way first, so the diagnosis is not inferred from the stack trace alone. The new unit test requires the reporter directly and asserts the returned value is nully — the assertion that would have caught this before it reached the runner — plus the exact NDJSON written. No subprocess, so this half of the feature is verifiable without a full runner pass, which matters because every defect in this feature so far has only been observable remotely. T1 remains untouched and red. Verification runs on the remote runner. Refs #4012 |
||
|
|
0abd137ec7 |
fix(#4012): the events reporter has to survive SIGKILL
The remote run proved the instrumentation did not work. Timing landed — "chunk 1/1 was killed after 2006ms" — but the in-flight-file naming produced nothing and fell through to the pre-existing generic message. The feature I wrote to diagnose a kill was itself destroyed by the kill. Root cause, confirmed rather than assumed. The reporter yielded strings, which node pipes into the --test-reporter-destination WriteStream. That stream BUFFERS. execFileSync's timeout sends SIGKILL, which is uncatchable and gives nothing a chance to flush, so the events sat in a buffer that died with the child. The parent's own timer reported correctly because it lives in the parent — which is exactly why half the feature looked fine. The reporter now writes each event with fs.appendFileSync, unbuffered and durable at the moment it happens, to a path passed through GSD_RUN_TESTS_EVENTS_FILE. Env vars do not count toward the Windows 32,767-char argv ceiling, so moving the path out of argv also REDUCES FIXED_OVERHEAD; the accounting moved with it rather than being left stale. The destination is now a fixed devNull sink that stays empty by design. Silence was the reason this was invisible for a whole run. Failing to read the events file now says so explicitly, and distinguishes a file that could not be read at all from one that exists but is empty — the generic fallback firing quietly is what let a broken feature look like a working one. A write is unbuffered but not atomic, so a kill can still interleave a partial line; the reader tolerates exactly one unparsable trailing line and reports the complete ones before it. The failing T1 was left red and untouched rather than weakened to pass. Three new unit tests cover the reader directly, with no subprocess, so the parsing half is verifiable without a full runner pass: missing file, existing-but-empty file, and a truncated final line. Also adds ndjson-reporter.cjs to GSD_SCRIPTS_LIB_FILES in bin/install.js — scripts/lib/ ships, and omitting it meant the file would install everywhere and orphan on uninstall. That single omission caused 4 of the 7 remote failures. Verification runs on the remote runner. Refs #4012 |
||
|
|
8497833a15 |
fix(#4012): a killed chunk now names the file that was hanging
The per-chunk timeout fired correctly but reported almost nothing, so every diagnosis cost a CI round-trip. run-tests.cjs logged chunk START only — no timestamp, no duration, no end line — then on a kill printed all ~55 basenames and asked the operator to work out whether output kept flowing (slow) or stopped early (hang). It could not name the in-flight file because the child is spawned with stdio inherit, deliberately, per #3597/#1051. Three additions. Per-chunk elapsed timing on every path, not just failures, so drift toward the cap is visible before it becomes a kill. Every timing number in the investigation behind this had to be reconstructed by hand from GitHub log timestamps. A second, machine-readable reporter running ALONGSIDE the human one, writing NDJSON to its own file. On a kill that file is read back and the files with a test:start and no matching completion are named, with the staleness of the last event, so "stopped 480s ago at X" reads differently from "still emitting at kill". stdio stays inherit and nothing is piped or tee'd — the maxBuffer and live-output risks that shaped the original design are untouched. Ranking of the killed chunk's files by known weight, flagging any absent from tests/test-timings.json, since an unweighted file is an unknown quantity. Two details that are correct rather than lucky. Passing --test-reporter at all replaces node's implicit default, so the human reporter is now named explicitly and reproduces node's own selection (spec on a TTY, tap otherwise) — visible output is unchanged. And the destination path's chunk index is zero-padded to a fixed width because FIXED_OVERHEAD is computed ONCE before chunking; a variable-length path would have silently mis-accounted the Windows 32,767-char argv ceiling and reintroduced #3597. The reporter flags are added to FIXED_OVERHEAD exactly as --test-force-exit is. The multi-reporter pairing and the stream.compose reporter contract were confirmed against Node's v24 documentation, not recalled — the first draft carried them as an unverified assumption and said so. Also corrects a stale comment claiming the 600s cap sits "below the 20m job cap". The lane is sharded 3x at timeout-minutes: 45; the windows shards were at 19m when chunk 1/5 was killed on |
||
|
|
ab69b9ce56 |
enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from |
||
|
|
d24e22b156 |
enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1 ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what remained was the declaration — and the pin that makes it invisible today. The census corrected two documented figures before any code changed. ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier note of mine claiming 23 was wrong and is corrected). And output({error}) is **64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2. That matters because this phase's criterion demands the pin be asserted over the enumerated population rather than sampled; asserting over a stale 60 would leave four sites unpinned while claiming full coverage, which is the shape of failure this epic exists to remove. The issue does not state the fact that shapes the design: output() never touches the exit code. Confirmed by reading it — it writes fd 1 and returns. So a declared outcome for those 64 sites had nowhere to be READ. The mapping was never the work; wiring somewhere for the declaration to land was. The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol cells, each because the module is emitted to three locations and a module-level `let` would let instances disagree, and runMain already maps a code returned by main(). A third cell inherits that solution. output() records DEGRADED for any {error} payload — key-order agnostic, which is exactly why the "42 sites" figure undercounts — and runMain projects the cell only when main() returns nothing, so an explicit return still wins. error() maps its reason through a table over the closed 25-member enum, leaving all 278 call sites untouched; 226 of them pass no reason at all. The version gate lives in error(), NOT in projectOutcome: registered names are version-invariant there, so mapping a reason straight through would make USAGE project to 64 under v1 and break the pin on its first line. projectOutcome is left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included. Proven rather than asserted. v1 is byte-identical across three real CLI paths — config-get plain, config-get --json-errors, and an output({error}) path — matching exit code and exact bytes against the pre-change build. Under GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND -> NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason, because without it a mapping where everything projects to 1 under both versions would satisfy every other assertion and the declaration would be theatre. A1 iterates all 25 enum members and A3 asserts over the measured 64-site population, so a 26th reason or a 65th site fails until it is given a mapping — the drift guard this phase needs, given ADR-2980's own count had drifted +4 unnoticed. Verification runs on the remote runner. Refs #3912 * fix(#3912): the outcome cell must never lower an exit code The remote run caught a fail-open that this phase introduced, in the phase whose entire purpose is removing fail-opens. `state validate --strict` on a missing STATE.md exited **0** where it must exit 1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set non-zero by the command was clobbered down to success. Confirmed live against a fixture, before and after. This refutes a review conclusion recorded earlier in this phase, that the cell was "fail-closed and can never mask a failure as success". It could, and did. Recording that plainly so the assumption is not repeated: the cell's danger was never only that it might add a failure — it was that projecting it unconditionally overwrites whatever decision came before. Projection is now guarded: it may set a code only when none is set, and an already-non-zero exit code always wins. The full precedence — explicit `main()` return, then an existing non-zero exitCode, then the declared outcome — is written at the projection site. A regression test drives a void return with a pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix build. The second failure was my test encoding the wrong contract, not a code defect. It asserted `output({found:false, error: undefined})` records DEGRADED because the KEY is present. `JSON.stringify` drops undefined, so the payload the user receives is `{"found":false}` — carrying no error at all, and calling that degraded would hand back exit 80 under v2 for output that reads as clean. The discriminator is a serializable error VALUE, not key presence. The test now pins `{error: undefined}` as explicitly NOT degraded, and the design doc's wording is tightened to match. Verification runs on the remote runner. Refs #3912 * docs(#3912): the versioned exit contract, and a flag defect the docs found Diataxis pass for Phase 8, plus a real fix that only surfaced because writing the how-to meant running its own examples. The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned projection this phase provides, so it gets an amendment naming #3912 / ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects DEGRADED to 80. The amendment also records the count drift rather than restating a stale figure — the ADR ratified 60 output({error}) sites in 9 modules; the AST-measured population is 64 across the same 9 (frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the enumerated 64. json-errors.md gains the outcome-declaration reference, including the precedence order a review pass got wrong and the suite refuted: an explicit main() return, then an already-set non-zero process.exitCode, then the declared outcome. Projection may only ever set a code, never lower one. A how-to is owed here and is written, not skipped. Under v1 nothing changes, so the audience is an operator opting into v2 and needing to know what the codes mean for a CI gate — a migration, which is how-to shaped. It covers turning v2 on, the code table, why 80 is "ran and reported a condition" rather than a crash, and how to split a gate that treats any non-zero as fatal. No tutorial: there is no new entry point to learn, and under the default contract a reader would be walked through observing nothing. The defect. Running the how-to's own Step 1 example returned $ gsd-tools --exit-contract=v2 state validate --strict Error: Unknown command: --exit-contract=v2 (exit 64) while the same flag trailing the subcommand worked and exited 80. The flag half-worked, by argv position. resolveContractVersion scans argv non-destructively, so the token survived into the dispatcher, which treats argv[2] as the command name. --json-errors had already solved precisely this at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv splice must happen here too, otherwise the dispatcher below sees --json-errors as an unknown command." The later flag never got the same treatment. Fixed rather than documented around: the version is resolved first — which memoizes the cell and makes an invalid value throw early — and then every --exit-contract= token is spliced out of the dispatcher's argv copy. --exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The regression test pins leading position, trailing position, agreement between the two, and a loud failure on v3 rather than a silent fall back to v1. Neither review engine would have caught this: the defect is invisible in the diff, because the diff does not touch argv handling. It surfaced only from running the documentation's own example. Writing a how-to is an execution pass. Verification runs on the remote runner. Refs #3912 * fix(#3912): the flag splice has to run before the run-with-timeout return An isolated review of the previous commit found that the fix did not deliver what it claimed, and that two of its own tests were weak. All three findings reproduced by execution before any change was made. The fix was placed below a return. main() intercepts `run-with-timeout` at gsd-tools.cjs:4436 and returns from there — above both the --json-errors block and the --exit-contract splice added in the previous commit. So the flag still died in leading position for that one command: $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..." Error: Unknown command: run-with-timeout (exit 64, child never ran) The previous commit message and the test's describe-block both claimed position-independence unconditionally. That was an overclaim, not a gap left open, and it is the part worth naming: the fix was verified by hand on the commands I happened to think of, and `run-with-timeout` returns before the code I was verifying. Both global-flag blocks now run above the interception, with a comment naming it so a later edit cannot slide them back down. Moving --json-errors up fixes the identical pre-existing bug for that flag, verified failing beforehand (exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect, same block, and a known-broken twin next to a fixed one is not a resting state. Two tests were not pulling their weight. The invalid-value test was vacuous — it passed against the pre-fix build, because `--exit-contract=v3` already exited 1 there and already printed the resolve error lazily through error() -> getContractVersion. Both its assertions held before the fix, so it pinned nothing. The real discriminator is that the pre-fix build emits BOTH "Unknown command: --exit-contract=v3" and the resolve error, while the fixed build emits only the latter; the test now asserts that absence. The leading-position and leading==trailing tests asserted proxies — "not 64", "no Unknown command", "the two agree" — none of which pin a value, and all of which would survive both positions being identically broken. With a .planning directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0 under v1 in both positions. Those numbers are pinned now. The multi-token case the descending splice loop exists for is covered too, and run-with-timeout has regression tests for both flags. The lesson is narrower than "test more". Hand-verifying the production behavior does not verify that the test would have caught its absence. The pre-fix binary has to be run against the test's own assertions. Investigated and deliberately not changed: splicing before --cwd parsing degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd: <path>", but that is pre-existing — verified on the pre-fix build via --json-errors, which already did it. This change joins the pattern rather than creating it, and both forms exit 64 on malformed input either way. Verification runs on the remote runner. Refs #3912 * chore(#3912): backfill changeset pr numbers to 3983 * test(#3912): pin the reason-table invariant as set equality, not a count A graph-backed review flagged the unchecked lookup in expectedErrorCode3912. Investigated by execution: the drift guard DOES hold — for an unmapped reason under v2 the production error() yields 1 while the table yields undefined, so the assertion fails. Not a correctness defect, and deliberately NOT made tolerant, since a tolerant lookup would destroy the guard. Two real problems remained. The guard asserted the wrong invariant: it counted the TABLE's keys at 25 rather than checking they match the ENUM's values, so a renamed member keeps the count at 25 and slips past, and a 26th member leaves the table at 25 and slips past too. Both were then caught only indirectly, by an undefined mismatch producing 'must exit undefined'. It is now a sorted set equality, so the failure names the specific missing or extra reason. And the comment above it described a '?? FAIL' fallback that does not exist anywhere in the function. It now states what the code actually does, verified by running it rather than by reading it. Refs #3912 --------- Co-authored-by: sim <sim@local> |
||
|
|
2ea5efc151 |
enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local> |
||
|
|
941b62249e |
enhance(#3906): two terminators over one registry, with a versioned exit projection (#3924)
* feat(#3906): two terminators over one registry, with a versioned projection Adds terminateNow (write-then-terminate, for callers that cannot wait for the event loop) beside runMain (drain-then-exit), both projecting through one shared function so they cannot disagree - the parity the ADR makes mandatory. A failed write does not change the exit code: letting it propagate would fail a hook open, which is what the fail-closed branches exist to prevent. The projection is versioned. v1 reproduces today's integers, including keeping a payload-carried degraded result at exit 0 - ADR-2980 ratified that across 60 sites and declined normalizing it on measured blast radius. v2 applies the registry. --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 selects it; an unrecognized version throws rather than silently defaulting. The registry is now emitted beside both copies of the exit module, so it resolves as a sibling in the built tree and in the committed scripts/ copy that must load on an unbuilt clone. * fix(#3906): actually restrict code 2 to terminateNow, and generate the registry's type The claim that terminateNow is the only place 2 can be produced was false: runMain's outcome arm applied no guard, so runMain(()=>'HOOK_DENY') set exitCode 2 through the drain path - and the parity matrix demonstrated it while calling it parity. runMain now refuses any outcome projecting to the hook-protocol code, gated on the code rather than the name so an alias cannot slip past, and the matrix asserts the restriction instead of contradicting it. The ambient type for the generated registry was hand-written with no gate against the generator's actual output - the declared-surface-diverges-from-runtime defect class this epic exists to close, reintroduced inside it. It is now a third generated artifact covered by the same --check. Also converts every test-body try/finally to t.after(). * test(#3906): derive the glossary fixture's dependencies instead of hand-listing them Adding a require to scripts/lib/cli-exit.cjs broke 31 tests in one suite that built its fixture from a hand-written dependency list, so the new sibling was absent and the copied script could not load. copyScriptWithDeps walks the require graph and exists for exactly this class - #3412 paid the same bill when one new require broke 82 tests across two suites. Migrating rather than adding another copyFileSync line keeps the class closed. The other nine suites referencing that path were triaged; none copies-and-spawns, so none needed migrating. * fix(#3906): enumerate the new shipped file, drop a vendor name from shipped data, and fix three test defects install: scripts/lib/exit-code-registry.cjs was missing from GSD_SCRIPTS_LIB_FILES, so it shipped to every install and orphaned on uninstall. The registry gave HOOK_DENY a meaning naming one harness, and that string ships into every runtime's tree - a guard correctly caught it leaking into the hermes and qwen installs. The registry is runtime-neutral infrastructure; the vendor name belongs in the ADR, not in shipped data. Two more fixture harnesses built their trees from hand-listed dependencies and broke on the new require; both migrated to the derived helper, and all 23 copy-and-spawn candidates were enumerated so the class is closed rather than patched. One generator test used a fixture code that collided with a real allocation, so the generator correctly reported a duplicate where the test expected drift. The large-payload test embedded a 256KB literal in the child's argv, exceeding Linux's 128KiB MAX_ARG_STRLEN so the child never started - it now builds the payload inside the child. * chore(#3906): backfill changeset pr number * docs(#3906): document the exit-code contract selector P2 is the first phase of this epic with a user-invocable surface, so the flag and env var owe a reference entry. Records what actually differs between v1 and v2 today (one outcome), that an unrecognized value is rejected rather than silently defaulted, and the fail-safe property that makes switching safe. --------- Co-authored-by: sim <sim@local> |
||
|
|
8edace40d5 |
enhance(#3904): one exit module — generate the scripts-side copy from a single source (#3917)
* test(#3904): failing-first coverage for the drifted scripts-side exit module The scripts/ copy of the CLI exit seam has no json-error arm, so an unexpected throw prints a raw stack where the documented contract promises {ok:false,reason,message}. Adds the consumer-altitude reproduction plus the negative space it must not swallow, the one-cell assertions for json-error mode, and the standalone-load constraint. RED until the generator lands. * enhance(#3904): generate the scripts-side exit module from one source src/cli-exit.cts becomes the single source of truth and scripts/lib/cli-exit.cjs a generated artifact of its compiled output, byte-compared by a --check entry in lint:generated-sync. The two had drifted: only the .cts copy emitted the documented {ok:false,reason,message} envelope on an unexpected throw, so a scripts-side tool printed a raw stack where docs/json-errors.md promises structured output. The generated file is committed and must load on an unbuilt clone (64+ consumers, incl. check-env.cjs), and gsd-core/bin/lib/cli-exit.cjs is gitignored tsc output that doubles as the build sentinel, so it cannot be required from there. The exit module therefore drops its io.cjs import: the json-error-mode accessors move into it and io.cts re-exports them, leaving its export surface unchanged. The flag lives in a Symbol-keyed cell on globalThis because one source emitted to two locations means two module instances, and a module-level flag would give them two independent values. * chore(#3904): changeset for the generated scripts-side exit module * docs(#3904): name which surfaces honor the json-error envelope contract docs/json-errors.md described the structured envelope as what runMain does without saying which copies of runMain actually had the branch — a claim that was silently false for every scripts/-side tool. Also drops a redundant source-grep test whose marker grew the unverified allow-test-rule pool past its ceiling; the behavioral test beside it proves the same property through real module resolution. * test(#3904): compare exit verdicts, not stderr bytes, across the two copies The parity test asserted byte-identical stderr, which the plain-text path cannot satisfy: the generated copy carries an 11-line banner, so its stack frames report line numbers offset by exactly that much, and the path normalizer stopped at the colon. Byte-identical stack traces were never the contract - two files at two paths necessarily differ there. Now compares the parsed envelope under json mode, the first line and exit code on the stack path, and exact output for ExitError. * chore(#3904): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
4b69dc346b |
fix(#2725): repoint the pre-commit alias guard at sources git can actually stage, drop nine dead ones (#3273)
* test(#2725): failing-first coverage for the inert pre-commit alias guard Replaces two stale tests that asserted on `sdk/src/query/command-manifest.phase.ts` — a path retired with the SDK boundary (ADR-0174), so both passed forever while guarding nothing. The new matrix drives .githooks/pre-commit through its GIT_OVERRIDE/NPM_OVERRIDE seams and asserts on the real tracked sources the drift checker reads. Red until the guard is repointed off the gitignored build outputs it currently watches. * fix(#2725): repoint the pre-commit alias guard at sources git can actually stage `.githooks/pre-commit` carried ten staged-path guards and not one of them could do its job. Nine anchored on `^sdk/…`, a tree retired by ADR-0174, and invoked npm scripts that no longer exist (`check:state-document-fresh` and eight siblings). Their only reachable behavior was to abort the commit with `Missing script` — which required matching a path that cannot exist, so they were dead twice over. The tenth was the real defect. Its npm target does exist, but it matched `^gsd-core/bin/lib/command-aliases\.cjs$` — a gitignored build output (.gitignore:172). An ignored path never appears in `git diff --cached --name-only`, so the guard was not stale, it was structurally unmatchable: it watched the derived layer instead of the source layer, from the day it was written. Repointed at the nine tracked `src/*.cts` sources `scripts/check-alias-drift.cjs` actually reads. The family table moves to `scripts/lib/alias-drift-families.cjs` so the checker and the hook derive their surface from one list, and a new parity test fails if a family is added to the checker without the hook learning to watch its source — the rot mechanism, not just this instance of it. Matching is now `grep -Fxqf` (fixed strings, whole line): exact path equality, no regex anchors to get wrong as the list grows. Staged paths are collected into a variable before matching, because `grep -q` exits on first match and would SIGPIPE its upstream, which under `set -o pipefail` turns a successful match into a non-zero pipeline status. The two CONTRIBUTING.md recipes pasted copies of the hook bodies inline — a third parallel surface, and one that had already drifted: the pre-commit copy carried the same dead `sdk/` patterns, and the pre-push copy would have overwritten the committed hook with a paraphrase that drops the GIT_OVERRIDE seam its test drives. Both now just point git at the committed files. The pre-push recipe's `'@example-corp\\.com$'` also never matched anything — inside single quotes bash keeps both backslashes. `.githooks/pre-push` was audited for the same rot and has none: it keys on no paths, no-ops unless GSD_BLOCKED_AUTHOR_REGEX is set, and is covered. Unchanged. Hooks stay opt-in. Nothing registers `core.hooksPath` for you, per CONTRIBUTING.md's documented one-time setup. * test(#2725): make the hook/checker parity assertion bidirectional Review finding from the isolated adversarial pass: the parity row only caught the hook UNDER-watching relative to scripts/lib/alias-drift-families.cjs. Drop a family from the module and the hook would keep watching its source with nothing to notice — the same divergence class this change exists to close, just pointed the other way. The new row takes every `src/*-command-router.cts` on disk as the universe and asserts the hook stays silent for the eight routers the drift check does not read. Both directions are now covered by running the real hook, not by comparing two lists in the test. Also corrects a CONTRIBUTING.md overclaim caught by the standards pass: 9 of the 11 watched paths derive from the module, not all 11 — bash cannot require a CJS module, so the hook carries literals and the test is what binds them. * fix(#2725): ship the shared family table and fix the mock that hid its own rows Three defects the remote runner caught that local probing did not. The mock `git` in the regression test emitted its staged-path payload as `printf '%s' "src/command-aliases.cts\n"`. Bash does not expand `\n` inside a double-quoted string and printf does not expand escapes in a `%s` argument, so the mock produced one unterminated line containing a literal backslash-n. No whole-line match could ever succeed, and every row that expects the hook to FIRE failed while every row that expects silence passed — which is exactly the signature the run reported: 9 failures, all of them fire-expecting rows. The payload now goes through a file the mock `cat`s, which is byte-exact and is what makes the CR-terminated and empty-staged-list rows mean what they say. `scripts/lib/alias-drift-families.cjs` was not enumerated in `GSD_SCRIPTS_LIB_FILES` (bin/install.js:377), so the installer never copied it. That is not cosmetic: `scripts/check-alias-drift.cjs` ships, and it now requires this module — an installed tree would have failed with MODULE_NOT_FOUND the first time a consumer ran `check:alias-drift`. Added to the manifest, which is what the #3184 install/uninstall parity tests assert against `readdirSync`. Regenerated the 19 committed install-tree fixtures via `npm run gen:install-tree` to record the new emitted path. The diff is +1 line per fixture and nothing else. `npm run lint:ci` exits 0. The earlier claim that `scripts/` is outside the emitted surface was wrong: `scripts/lib/` is copied into every runtime's install tree, which is why 19 golden-install-tree cases moved. * chore(#2725): backfill changeset pr: 3273 --------- Co-authored-by: sim <sim@local> |
||
|
|
342590c70e |
refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite Covers the 50 input classes in the phase test matrix: scope classification (genuinely-empty vs truncated vs unscoped vs unreadable), the section-end owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c), the milestone.complete refusal with negative proof that no directory moved, the version-token boundary defect, drift-guard behavior, and three fast-check properties over document-shaped generators. Committed alone so the remote runner records the failure before the fix lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * refactor(#3184): milestone windowing routes through one owner Three copies of the milestone section-end walk lived in roadmap-parser.cts — two distinct computeSectionEnd function nodes plus an inline third in getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is now the sole owner and the other two are deleted, not kept in sync by comment. The whole-repo drift guard found what the epic did not: state.cts held three more re-derivations of the same vocabulary — two byte-identical milestone bounding checks carrying a defect neither reported copy has (no boundary after the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning predicate. All three route through the owner now. A composition-level duplicate appeared inside this change's own first pass: getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out of the owner's primitives, and had already diverged on whether to skip a closed milestone heading. sliceMilestoneWindow is the one composition. Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is distinguishable from a genuinely empty milestone — those were output-identical, which is the whole failure class. roadmap analyze emits it (#3165), and milestone complete refuses to archive on anything but COMPLETE rather than pass-all moving every phase directory on disk (#3166). The pass-all degrade is preserved where its premise holds: making the filter deny-all would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. extractCurrentMilestone keeps its signature — 200+ affected symbols across 41 files and 25 process flows — and is a one-line wrapper over the scoped owner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): fence-aware phase detection and one heading-selection owner Review fixes from the two orthogonal passes. The blocker: hasPhaseEntries matched ATX phase headings fence-aware via tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown, so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely empty milestone then classified TRUNCATED and milestone complete refused a legitimate archive — a false positive in the destructive direction, worse than the defect this phase set out to fix. Both that path and getMilestonePhaseFilter own pre-existing bullet scan now run on stripFencedCode, since leaving one meant the owner file gave two different answers to the same question. The selection rule — locate, prefer the non-closed heading, else the first — had been written three more times inside the file whose thesis is single ownership. selectMilestoneHeading owns it; all three sites route through it. The copies were behaviorally identical, so this is de-duplication with no observable change, verified by probing that all three paths select the same heading. roadmap analyze emitting a scope no consumer read left #3165's actual symptom alive, so Route 0 in next.md now treats a non-complete scope as scan-failed rather than as a clean empty scan, and the ADR amendment no longer overstates what shipped. Also: the scope refusal moved above the archive-directory create, so a refusal leaves nothing on disk; the versionOverride comment names all four consumers; COMMANDS.md documents the new guard beside its sibling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#2658): exclude the changelog from the malformed-path scan The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships into that tree, and its #2658 entry quotes both malformed paths while describing the fix that removed them — so the release note documenting the fix trips the fix's own regression test. Red on next before this branch. The installer is correct: a probe over a real --trae --local install found 621 emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope was the defect, not the product. Excluded by exact relative path rather than by loosening the patterns or skipping all markdown — the emitted agent and command markdown is precisely what #2658 was about, so the gate stays strong everywhere it matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): regenerate install-tree fixtures for the shared drift scanner scripts/lib/ ships in the npm package and installer, so extracting the shared tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's install tree. Regenerated via npm run gen:install-tree; the delta is exactly that one path per fixture. The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded), so only the extracted library moves. This matches the existing scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only helper carried in the shipped tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal The remote runner caught two regressions this branch introduced. Both were mine, and neither review pass found them — only running the existing suite did. The version-token boundary. I replaced locateMilestoneHeadings' \b with (?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562 fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a closed v8.0-A sibling (#730), and \b is what allows it while the stricter boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to \b; the state.cts consolidation is now a straight merge with no behavior change, and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and the design doc no longer claim otherwise. The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is about the TRUNCATED window specifically — the heading is found and the section closes before the phase region, so pass-all archives everything. UNREADABLE and UNSCOPED are pre-existing, legitimately handled states, and refusing on them broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed to TRUNCATED; docs corrected to match. One of the new tests was also wrong: its fixture gave the shipped and current milestones' phases the same numeric id, and the filter matches on that id, so it could not have distinguished the two windows. Fixture corrected to exercise what it claims to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): enumerate drift-scan.cjs for uninstall The installer copies scripts/lib/ wholesale, but uninstall removes an explicit set — deliberately, so a user's own helpers in that directory survive. The extracted drift-scan.cjs was copied in and never enumerated, so it outlived uninstall, left the directory non-empty, and the rmdir that follows failed. Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is likewise a lint-only helper that ships there and is enumerated. Verified with a real install-then-uninstall into a temp target: scripts/lib/ held exactly the three GSD files and was gone afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset Found while shipping this phase, and fixed here rather than noted. install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE — the comment at the copy site literally says "and any future lib helpers". uninstall() removes them by hardcoded enumeration, deliberately, so a user's own helpers in those directories survive. A wholesale writer paired with an enumerated remover cannot stay in sync by construction: any file added to either directory ships to every user and is then orphaned in their repo forever, since it survives uninstall, leaves the directory non-empty, and the rmdir that follows fails. Nothing reported this. 31,225 tests were green over it. That is the same divergence class this epic exists to delete, sitting in the installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity assertion that fails the moment the two surfaces disagree. The test compares each directory's real contents against its enumeration and names the offending file plus the constant to add it to. Both enumerations are hoisted to module scope and exported, so the test asserts on the actual arrays rather than pattern-matching the installer's source — no allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on the current tree, correct report when an unenumerated file is injected. scripts/changeset/ turned out to carry the identical defect and is covered too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * chore(#3184): backfill changeset PR number Also narrows the wording to match the shipped behavior: the refusal fires on a truncated window specifically, not on any non-complete scope. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74d7bc8239 |
test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3) Introduces a committed per-file size baseline scheme alongside (not replacing) the existing tier anti-creep tests. Green by construction — the baseline records current sizes, so both schemes pass side by side during migration. - scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper, same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline. - scripts/workflow-size.cjs: single source of truth for LF-normalized byte counting (#683) + workflow enumeration, shared by the guard and the generator so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/) because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed runtime, scripts/ root is not, so this keeps it out of the shipped payload. - scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the snapshot (sorted keys, trailing newline, idempotent). - tests/workflow-size-baseline.json: generated snapshot (88 workflows). - tests/workflow-size-budget.test.cjs: import the shared counter (drops the duplicated local byteCount) and add the per-file baseline describe block. - Tests for the helper, the shared module, and the generator (incl. round-trip and fault-injection cases). Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test. * test(#1074): regenerate workflow baseline after Update-branch merge with next The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090) without regenerating the snapshot, leaving the per-file baseline stale by one file. Re-ran `npm run size:baseline` so the committed baseline matches the merged workflow files. Refs #1074. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
f729101eec |
refactor(scripts): replace process.exit() with ExitError + runMain handler (#739) (#740)
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern. - New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain() which translates a thrown ExitError / returned number into process.exitCode (never process.exit()), flushing output and still firing process.on('exit'). - main()-based entrypoints: throw new ExitError(code) for errors, return <code> for verdicts; invoked via runMain(main). Child exit codes preserved via return. - top-level-only scripts: imperative body extracted into main() so mid-flow aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope. - diff-touches-shipped-paths.cjs: stdin event handling restructured to an async read so the whole flow runs under runMain; uncaughtException/unhandledRejection nets replaced by an in-band catch that preserves EXIT_ERROR=2. Exit codes verified unchanged for every converted script (success/error/help and the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error in part 2 (#738) once gsd-core/bin/** is also clean. Refs #739 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a28dcec981 |
chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.
Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
maxRetries (it owns no loop), so the test asserts the option contract
(recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
the helper, not approximated textually at every call site.
Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
destructured and aliased forms; escape hatch is inline
`// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
error-swallowing or name-colliding local teardown helpers) keep the raw call
with an inline eslint-disable + reason.
Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
- assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
STALE ids (a known offender was fixed but not pruned) — identity, not count,
and a ratchet DOWN toward zero.
- assertTightCeiling: a size/length budget whose ceiling must stay within a
grace band of the high-water mark, so budgets may only tighten, never creep.
Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
the remaining six patterns moved from integer baselines to named-set
allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
named filename sets (closes the swap-a-file-keep-the-count blind spot); a
module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).
Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
current high-water mark and an assertTightCeiling anti-creep check added per
tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
are intentionally left as-is — they are not grandfathered creeping budgets.
No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|