af822a8024f732c167b163f163a0487bdbcfebdf
102 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
585cab41cc |
fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices (#4920)
* fix(#4770): lift the Codex sandbox_mode holds — documented enforcement suffices * docs(#4770): backfill the changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
a841575037 |
test(#4527): migrate planning/review-lane batch to named timeout constants (#4680)
Batch 16 of 17 in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in 9 test files with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. No src/bin file touched, no numeric value changed anywhere. Ground truth via eslint found 48 sites, not the issue's stated 38 — tests/code-review.test.cjs alone has 11, not 1 (a 10-site undercount, the largest single-file drift in this epic). All 11 are migrated. Reuses PROBE_TIMEOUT_MS (17 sites across assumption-delta.test.cjs and code-review.test.cjs), LOOP_HOOK_POINT_CLI_TIMEOUT_MS (1 site), QUICK_SPAWN_TIMEOUT_MS (1 site). Adds a new shared constant, HTTP_REACHABLE_PROBE_TIMEOUT_FIXTURE_MS, promoted because two files (reviewer-manifest-body.test.cjs, reviewer-trust-disclosure.test.cjs) independently arrived at the same probe-fixture value across 3 sites. File-local constants elsewhere for values not shared across files: FALLOW_AUDIT_TIMEOUT_MS (code-review-pipeline-regression.test.cjs); a 5-constant set covering runBashScript's own override/forced-timeout/ boundary-triple tests (plan-phase-stall-detection.test.cjs); 9 lane-specific NATIVE_TIMEOUT_MS constants, one per shipped reviewer CLI tool, even where 4 lanes coincidentally share a value (review-lane-invocation.test.cjs); and a 6-constant set covering reviewer-manifest-body.test.cjs's own probe-kind fixtures and its separate, unrelated boundary/invalid set that coincidentally overlaps in shape (not value) with plan-phase-stall-detection.test.cjs's triple. Fixes applied inline from review: GENERATOR_SCRIPT_TIMEOUT_MS was misapplied to 3 sites in plan-review-convergence.test.cjs whose actual shape (nested bash -> node -> gsd-tools.cjs chain) doesn't match that constant's documented direct-spawn class, confirmed against its own cited precedent files — replaced with a new file-local constant naming the correct class, same value. Two stray un-renamed literals in review-lane-invocation.test.cjs (missed by eslint's object-literal-only detection) were investigated and left as disclosed literals with an explanatory comment rather than force-fit onto an unrelated lane's constant, since they belong to a distinct config-override code path. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
56b706a07a |
test(#4524): migrate task-content resolution batch to named timeout constants (#4675)
Batch 13 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeoutMs object-literal property in tests/task-content-resolution.test.cjs, tests/task-command-router-resolve-content.test.cjs, and tests/task-content-resolver-grammar-parity.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 3 files from the rule's allowlist. Every one of the 13 sites describes the same field -- a task-content- resolver manifest's invoke.timeoutMs -- as fixture/validation data; none is a real subprocess spawn timeout, verified by tracing each site to a pure function, a garbage-shape rejection path, or a fully-injected fake exec function. Adds one new shared constant, TASK_RESOLVER_INVOKE_TIMEOUT_MS, used by 2 files in this batch (crossing the promotion bar). Adds file-local constants for a value used by only 1 file, plus three deliberately-invalid values (zero, negative, non-integer) inside one findResolver garbage-shapes test proving the validator rejects a malformed manifest regardless of which way its timeout is invalid. Also fixes a review-caught defect outside the mechanical rule's own scope: a bare-literal duplicate of the new shared constant inside a ResolverTimeoutError assertion (a call argument, not an object-literal property, so the lint rule never flagged it) -- renamed in this same PR per the no-deferrals rule. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c99d7bb2be |
test(#4522): migrate core CLI/domain state batch to named timeout constants (#4662)
Batch 11 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/state-document.test.cjs, tests/phase.test.cjs, tests/commands.test.cjs, tests/pattern.test.cjs, tests/adr-612-bracket-coherence.test.cjs, tests/adr-612-bracket-read-tolerance.test.cjs, tests/milestone-lock.test.cjs, tests/init.test.cjs, tests/state-todos-render.test.cjs, tests/quick-batch.test.cjs, tests/graphify.test.cjs, and tests/effort-surface-axis.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 12 files from the rule's allowlist. Ground truth via eslint found 25 sites, not the issue's stated 24 (phase.test.cjs has 5, not 4) -- disclosed in the PR body. Reuses PROBE_TIMEOUT_MS, GIT_TIMEOUT_MS, and LOOP_HOOK_POINT_CLI_TIMEOUT_MS across 8 files. Adds two new shared constants to tests/helpers/timeouts.cjs (each independently arrived at by 2 files in this batch, crossing the promotion bar): PATHOLOGICAL_INPUT_TEST_TIMEOUT_MS (node:test's own per-test timeout option, not a subprocess bound) and GSD_TOOLS_CLI_MODERATE_TIMEOUT_MS (a single gsd-tools.cjs CLI subcommand spawn, distinct tier from PROBE_TIMEOUT_MS/LOOP_HOOK_POINT_CLI_TIMEOUT_MS). Adds 3 file-local constants for values used by only 1 file in this batch. No src/bin file touched, no numeric value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1f84f45ed5 |
fix(#4259): fold shell continuations before the T6 docs-parity site scan (#4423)
* fix(#4259): fold shell continuations before the T6 docs-parity site scan The scan is $-anchored with [^\n]* on both sides of --grep=, so `git log` and `--grep=` had to share a physical line. A backslash-continued derivation — the natural way to write a git log carrying a long ERE — produced zero hits and T6 passed on it. Both generations of the assertion were defeated. The current anti-revert ban let a wrapped site through outright; at v1.12.0, where T6 instead asserted pattern conformance, a wrapped site was silently exempted from the very checks written to catch the macOS \b-no-op class, so it could have carried exactly the malformed pattern T6 exists to reject. A real candidate implementation for #3926 wrapped its derivation, passed T6, and was caught only by later manual review. Fold the continuations before matching rather than widening the regex: the assertion's message and its PHASE_SCOPE_NUM filter both assume one site is one string, and a [\s\S]*? would run the scan across unrelated statements. Correcting the input repairs everything built on the scan at once. The fold uses [ \t]* after the newline rather than \s* — the shell's own rule, and it cannot swallow a blank line and glue two unrelated statements. The scan is hoisted to findGrepSites so the controls can drive it directly: the continued form is caught, the same-line form still is, a benign --grep stays clean, a wrapped unrelated assignment is not glued, a continuation does not cross a blank line, and the live workflow files still report nothing — so this lands without editing any workflow to appease it. * fix(#4259): honor shell continuation boundaries * test(#4259): use canonical fenced-block scanner * test(#4259): centralize shell continuation scanning --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
2eef8ada4f |
test(#4520): migrate generators/doc-gates/attribution batch to named timeout constants
Batch 9 of 17 in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in 14 files with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 14 files from the rule's allowlist. The issue's own guess ("all run a scripts/*.cjs generator or lint script once, BUILD_TIMEOUT_MS class") needed two corrections found by reading every site directly. First, BUILD_TIMEOUT_MS's own doc comment scopes it specifically to scripts/build-hooks.js, which none of this batch's generator/lint-script sites run — a new shared constant, GENERATOR_SCRIPT_TIMEOUT_MS, covers the class instead. Second, no-pending-3212-markers.test.cjs's single site spawns `git ls-files` directly, not a scripts/*.cjs script at all — routed to a second new shared constant, REAL_REPO_GIT_TIMEOUT_MS, promoted once emitted-attribution.test.cjs's own git-plumbing sites were found sharing the same class and value. Isolated Standards-axis review caught a further misclassification: one of REAL_REPO_GIT_TIMEOUT_MS's three emitted-attribution.test.cjs sites actually builds a fresh throwaway temp repo (createTempDir + git init), contradicting that constant's own real-repo-tree-only scope. Fixed with a new file-local FRESH_FIXTURE_GIT_TIMEOUT_MS holding the exact pre-existing value under an honest name, rather than reusing the shared GIT_FIXTURE_TIMEOUT_MS (which would have doubled the bound). emitted-attribution.test.cjs also gets two more file-local constants: HEAVY_REAL_TREE_TEST_TIMEOUT_MS (node:test's own per-test timeout option, not a spawn bound) and BUILD_HOOKS_UNDER_LOAD_TIMEOUT_MS (the same build-hooks.js script as the shared norm, at 4x its bound inside the suite's heaviest test). emitted-ack-trailer.test.cjs gets IMPOSSIBLY_SHORT_GIT_TIMEOUT_MS — the one value in this migration that is deliberately tiny (20ms), used to force a timeout in a negative test, not generous headroom. No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
150acd78c1 |
test(#4519): migrate security-scanner batch to named timeout constants
Batch 8 of 17 in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/secret-scan-lint.security.test.cjs, tests/security-scan.security.test.cjs, tests/security-prompt-injection.security.test.cjs, tests/prompt-injection-scan.security.test.cjs, tests/read-injection-scanner.security.test.cjs, tests/read-injection-scanner.property.test.cjs, and tests/security.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 7 files from the rule's allowlist. The issue guessed this batch "most likely needs its own named SCAN_TIMEOUT_MS." Reading every one of the 19 call sites directly found a more specific picture: 8 sites across 3 files scan exactly one small temp fixture file and match the existing QUICK_SPAWN_TIMEOUT_MS class exactly (reused, no new constant). Two new shared constants cover genuinely distinct classes that happen to coincide in value: SCAN_USAGE_ERROR_TIMEOUT_MS (a bash scan script given missing arguments) and MALFORMED_INPUT_HOOK_TIMEOUT_MS (a Node hook fed malformed JSON) -- kept as separate names per this migration's standing rule that numeric coincidence is never identity. Three file-local constants cover a real multi-file directory scan, a property-fuzzing safety net, and a path-traversal hook test, each with its own pre-existing rationale preserved. No bound is lowered or raised anywhere in this batch, honoring the issue's explicit caution that security-scan timing margins deserve extra scrutiny. No src/bin file touched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6d6e3eea73 |
test(#4518): migrate loop/hook-point e2e batch to named timeout constants
Batch 7 of 17 in the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/loop-render-hooks.test.cjs, tests/loop-walk.qa.test.cjs, tests/loop-hooks-empty-points-e2e.test.cjs, tests/loop-hooks-ship-pre-e2e.test.cjs, tests/loop-hooks-verify-post-e2e.test.cjs, tests/check-gap-analysis-plan-post-e2e.test.cjs, tests/check-tdd-review-checkpoint-e2e.test.cjs, tests/execute-wave-post-gate-pipeline-e2e.test.cjs, tests/plan-pre-hook-e2e.test.cjs, and tests/qa/tdd-walk.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 10 files from the rule's allowlist. Promotes a new shared class norm to tests/helpers/timeouts.cjs, LOOP_HOOK_POINT_CLI_TIMEOUT_MS: 7 files independently arrived at the same value for a single gsd-tools.cjs CLI subcommand invocation with no confirmed subprocess fan-out. Reuses the existing PROBE_TIMEOUT_MS for loop-render-hooks.test.cjs's 11 sites (same class, exact value match). Adds 4 file-local constants for values that share a class with the new norm or an existing one but diverge in pre-existing value, or that numerically coincide with an unrelated existing constant without matching its actual operation. Isolated Standards-axis review caught that the new shared constant's doc comment exhaustively enumerated 3 verb families while a 7th genuine site (an `init new-project` invocation) also correctly belonged to the class; fixed by rewording the comment to state the class definition (call shape) first and list all 4 representative verbs, explicitly illustrative rather than exhaustive. No value changed, no site's classification changed. No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93ba63aeff |
fix(#4482): strip Copilot notes from OpenCode artifacts (#4532)
* fix(#4482): strip Copilot notes from OpenCode artifacts * chore: add changeset for #4532 * fix(#4482): filter runtime notes across emitted surfaces * fix(#4482): make runtime note filtering runtime-neutral * test(#4482): keep ack fixture outside registered provenance --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
557a0b2876 |
test(#4516): migrate per-runtime install/upgrade adapters to named timeout constants (#4582)
Batch 5 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property across tests/augment-upgrades.test.cjs, tests/shared-hooks-dir-resolution.test.cjs, tests/antigravity-upgrades.test.cjs, tests/cursor-hooks.test.cjs, tests/cursor-hook-workspace-roots.test.cjs, tests/gemini-runtime-removed.test.cjs, tests/kilo-upgrades.test.cjs, tests/kimi-upgrades.test.cjs, tests/kimi-variant-disambiguation.test.cjs, tests/opencode-plugin-adapter.test.cjs, tests/windsurf-hooks-bridge.test.cjs, tests/effort-sync-installed-runtime.test.cjs, and tests/hooks-commonjs-marker.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 13 files from the rule's allowlist. Adds one new shared constant to tests/helpers/timeouts.cjs for spawning a single already-staged hook script directly (a heavier class than the existing quick-spawn norm), shared across three files in this batch. One new file-local constant covers a property-test driver process distinct from any existing class. Every other site reuses an existing shared norm. No src/bin file touched, no numeric timeout value changed anywhere. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
37b965c0d1 |
enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills (src/init.cts) now selects between a canonical agents/<name>.md and a token-minimized agents/<name>.compact.md sibling based on workflow.compact_content, resolved in code (a real function call with a real exit code) rather than a prose config-get gate — the same precedent stream 1's spine/detail split established for a load-bearing seam, applied here because this seam already runs through TypeScript instead of an eager @-include. A missing compact sibling falls back to the canonical persona and discloses the fallback in the served payload itself (a leading HTML-comment provenance line), so the Done-when contract — compact when on, canonical when off, never silent or empty — holds even for an agent nobody has compacted yet. Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md), each an independent, complete rewrite (not an extraction — nothing is "moved" the way spine/detail moves text) that preserves frontmatter, every @-include, every output-format contract, and every guardrail verbatim while cutting restatement and verbose framing. Verified mechanically: every pair registers (a canonical sibling exists), every compact file is strictly smaller, and the full @-include set matches canonical's — including which references are standalone eager-load lines versus inline prose mentions, since demoting one to inline changes what the host actually substitutes. Traced the install path before writing any code (.gsd/phase/.../40-design.md): stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no stem filtering under the default full profile, so the new .compact.md files install for free with zero installer changes — matching issue #4407's stated scope. A tiered agent profile that doesn't stage a compact sibling degrades through the same fallback-with-provenance path already required for an unauthored one, so no installer change is needed there either. Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export (deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are reached by a generic code construction rather than a literal path in prose, and checkReachability's markdown-search shape has nothing to find there). Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new "#4407 compact payload selection" describe block spawns gsd_run agent-skills against real compact/canonical fixture pairs and asserts on the served payload, which can only pass if the seam genuinely wires through. Fixed a pre-existing test whose agents/*.md glob incidentally matched the new .compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's roster, both real, unrelated-to-content defects the new files' mere existence surfaced. Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token benchmark baseline (npm run benchmark:compact-content-variants --write). Closes #4407. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): apply orthogonal review findings from the compact-payload seam Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath) to collapse the duplicated read-and-empty-check shape between the compact and canonical branches in cmdAgentSkills, and updated the adjacent comment enumerating flat JSON extras to name agent_payload_variant alongside source/degraded (added by the prior commit, comment left stale). Security review and the Spec axis found no defects requiring a code change; their non-blocking observations (a pre-existing, unmodified path-construction pattern; the reasoned, documented substitution of a behavioral test for the literal reachability check) are recorded in .gsd/phase/enhance-4407-agent-skill-seam/60-review.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents Root-caused via a real gsd-test run (93 failures) rather than guessing which tests glob agents/ naively. Two classes of defect, both genuine: 1. Identity-roster confusion (11 files/areas): many tests and one production script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f => f.endsWith('.md'))`, which incidentally matched the new .compact.md variant siblings too — a compact file is a rendering of an EXISTING agent identity, not a new one. Fixed at the shared root (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests already consolidated on) and at each independent glob that didn't use it: agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix before checking XL/LARGE membership, so a compact file inherits its canonical sibling's tier instead of silently falling through to DEFAULT), agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual script, not just its test), codex-config.test.cjs (confirmed directly against generateCodexAgentToml that a compact role's derived sandbox_mode is byte-identical to its canonical sibling's before excluding it — not assumed), and copilot-install.test.cjs (two counts that legitimately DO need both files — an installed-file count and a full-conversion smoke test — fixed to expect 70, not stay pinned to 35). no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix: two compact files reproduce descriptive prose already allowlisted at their canonical file's line number; added matching entries at the compact files' own line numbers rather than excluding them from the scan (a genuine bare gsd-tools command-position bug in a compact file would be as real a defect as in canonical). 2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree run): six agents' compact renditions (gsd-debugger, gsd-executor, gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction — confirmed structural, not a compaction-quality gap: each is dominated by content this phase's own rules require verbatim (the ~2.6 KB gsd_run bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in every agent that calls gsd_run, output-format contracts, guardrails). ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing spot in cmdAgentSkills's single-file synchronous read. Removed these 6 compact files rather than ship an over-cap file or invent a multi-part read mechanism out of scope for this phase; recorded by name with the reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and 50-test-matrix.md, per #4407's own "or explicitly recorded as not worth covering" allowance. Their canonical personas are served correctly today via the fallback-with-disclosed-provenance path this phase's own Done-when #2 already requires — 29 of 35 agents now have a compact variant. Also fixes an unrelated, genuinely pre-existing defect this gsd-test run surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md | wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed: two meaning-preserving trims in the <success_criteria> block (a repeated parenthetical replaced with a same-exception reference; one redundant qualifier dropped) bring it to 40,940 bytes. Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6 now-orphaned roster rows removed alongside them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): make .compact.md-aware roster checks resilient to partial coverage Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception), breaking once 6 stems legitimately have none. - tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was picking up the "### Compact Payload Variants" subsection's rows as phantom/uncounted entries in the primary/advanced/inventory-only classification this test validates — a compact row documents an existing agent's alternate rendition and never gets its own AGENTS.md heading, so it was never meant to participate in that classification. Excluded at the parser, not per-assertion. - tests/copilot-install.test.cjs: the derived expected-file-list generator assumed every listAgentFiles() stem has a .compact.md source sibling; checks disk per stem now instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4407): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
30468f16fe |
test(#4515): migrate installer-runs batch to named timeout constants (#4575)
* test(#4515): migrate installer-runs batch to named timeout constants Batch 4 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property across tests/install.test.cjs, tests/install-minimal-hooks.test.cjs, tests/fragment-single-edit-propagation.install.test.cjs, tests/install-regressions.test.cjs, tests/install-runtime-artifacts.test.cjs, tests/npm-integrity-gate.test.cjs, tests/faulty-deps.test.cjs, tests/release-tarball-smoke.install.test.cjs, tests/install-write-confinement.test.cjs, tests/plugin-manifest.test.cjs, and tests/packaging-shipped-scripts-require-only-shipped.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 11 files from the rule's allowlist. Adds one new shared constant to tests/helpers/timeouts.cjs for a fixture-JSON value (in seconds, not ms) mimicking a Claude Code settings.json hook-entry's own timeout field, shared across two files in this batch. Every other new constant is file-local, each carrying a comment explaining why its call site is a distinct operational class from the existing shared norms even where its digits numerically coincide with one. No src/bin file touched, no numeric timeout value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: combine gsd-worktree-path-guard's git rev-parse calls to cut subprocess count under CI load Discovered while verifying PR #4575 (a full test (windows-latest) failure in an unrelated test, tests/kilo-upgrades.test.cjs's worktree-path-guard rejection test). Root cause: this hook's up-to-4 sequential git spawns (git-dir, branch, worktree-toplevel, file-toplevel), invoked inside a native-plugin's own 8000ms-capped subprocess wrapper, left zero margin for node/git startup overhead — the guard is deliberately fail-open on any timeout, so CI-load-induced latency in this chain silently downgrades a security block into an allow. Combines the first three git rev-parse calls (git-dir, branch via --abbrev-ref, worktree-toplevel) into ONE spawn instead of three, using git's documented multi-flag rev-parse output (one line per flag, in order) — verified empirically including the detached-HEAD edge case. No timeout value changed; this reduces subprocess COUNT, not any tolerance. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: add changeset fragment for the worktree-path-guard fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ab0405ad22 |
test(#4514): migrate git-adjacent workflow checks to named timeout constants
Batch 3 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/ci-rebase-check.test.cjs, tests/gsd-validate-commit-crash-policy.test.cjs, tests/pr-branch-planning-filter.test.cjs, tests/reapply-verify-hunks.test.cjs, tests/ship-notes-wedged-pr.test.cjs, tests/slug-derivation-drift-guard.test.cjs, and tests/worktree-safety.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 7 files from the rule's allowlist. Mints a new shared class norm, QUICK_SPAWN_TIMEOUT_MS (10000ms), in tests/helpers/timeouts.cjs: 5 sites across 4 of this batch's files had independently arrived at the same value for the same shape (a cheap, trivial subprocess/hook invocation with no real git/network/fan-out work). No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ff5c5782ee |
test(#4512): migrate process-seam batch to named timeout constants
Batch 1 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property in tests/process-seam.test.cjs, tests/helpers-process-isolation.test.cjs, tests/run-with-timeout.test.cjs, and tests/helpers.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes the 4 files from the rule's allowlist. Adds SEAM_DEFAULT_TIMEOUT_MS to tests/helpers/timeouts.cjs (shared across 2 batch files, mirroring process-seam.cjs's own un-exported default). File-local constants elsewhere for values not shared across files or not a bench-derived class norm. No src/bin file touched, no numeric value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a27cb6b2fa |
enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b (gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4 (gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user: two independent, complete files per covered path (canonical + .compact.md sibling), with the gate picking which one gets Read at the call site. This is a different shape from Phase 5's spine+detail partition, and is safe here specifically because these files are already reached only by a runtime Read — a missed Read already means zero overlay content today, with or without workflow.compact_content, so selecting between two independently-complete files introduces no new failure mode (documented in gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section). Disposition, after inspecting every candidate rather than trusting a byte-size threshold (same rigor Phase 5 applied to review.md): - Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing reference doc emitted verbatim, not orchestrator instruction). The other 9 size-threshold candidates are dominated by fail-closed guards, exact CLI invocations, or output-format contracts (AskUserQuestion blocks) — recorded not-worth-compacting, same reasoning as Phase 5's review.md. - Stream 4: a ground-truth reachability audit replaced the initial size-only candidate list. Two files (summary.md, user-setup.md) got compact variants; a third (spec.md) was drafted, then dropped after discovering its only two call sites are eager @-includes, not a runtime Read — stream-1 material hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md itself has 3 eager call sites and only 1 genuine runtime-Read call site (execute-plan.md); only that one was wired, so the compact variant's savings apply to the sequential single-plan execution path only. - Discovered while auditing reachability: 12 gsd-core/templates/** files with zero references anywhere in workflow/agent/command prose, compiled source, or tests — dead scaffolding predating this phase. Deleted in this same PR per this repo's no-defer policy, after re-verifying against a computed path.join(...) pattern (not just a plain-string search) that nearly caused two genuinely load-bearing templates (user-profile.md, dev-preferences.md) to be misclassified as dead. New checker (tests/helpers/compact-content-variant.cjs): registration, reachability, protected-content-preserved, size-smaller — replacing Phase 3/5's disjointness/completeness checks, which assume a partition rather than two deliberately-overlapping documents. The reachability check's own "unprefixed match" guard had a real bug (rejected the repo's own `~/.claude/gsd-core/...` convention), caught by running it against the already-wired help/modes/full.compact.md pair rather than only synthetic fixtures — fixed to anchor on the nearest `gsd-core` path segment instead. Template consumer parity (tests/compact-content-template-variant-parity.test.cjs): proves each compact variant's `## File Template` fenced block — the actual output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed against — is byte-identical to the canonical file, then runs the one real deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent, backing `gsd-tools uat classify-coverage`) against content built from that shared contract. Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs) rather than extending the existing spine/detail one — different data shape, and the existing script's own contract deliberately isolates it from a test-only helper's shape changing. Emitted-drift acknowledgement: not needed. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before reaching the ack-lookup branch (same reasoning Phase 5 verified for its own diff). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * enhance(#4406): address code-review findings on the variant-swap gate - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content entries described only the spine+detail mechanism (Phase 5) and were missing this phase's variant-swap mechanism and its benchmark:compact-content-variants script entirely — required since this PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets" rule). Both now describe both mechanisms and which call sites are wired. - Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved: a canonical file with zero <!-- gsd:protected --> blocks must be a no-op, not a violation — the only branch of that function the existing fixtures didn't exercise. - Collapsed findCompactFiles/findMarkdownFiles in tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix helper — the two were identical recursive walks differing only in the extension predicate (minor Duplicated-Code finding). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification gsd-test caught this, not static analysis: 10 real failures in tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs, and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install path, which does fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md')) after copying gsd-core/templates/** into the target project, then merges it into both .github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the repo-root bin/install.js — a separately maintained installer bundle outside the src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around that read degrades to a silent skip rather than a crash when the template is missing, which is why this surfaced only once the real E2E install test ran, not from any static check. Re-verified the remaining 11 deleted filenames against bin/install.js specifically (plain substring and quoted-filename search) before trusting that list — all 11 have zero hits there, confirmed dead by the same standard this one file failed. Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md) and the phase design doc from 12 to 11 deleted files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule * docs(#4406): backfill changeset PR numbers Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion CI's own full-test matrix (not gsd-test's matrix, which does not run this check) caught 4 more false-positive dead-template classifications via tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs — a literal, word-boundary basename check across .github/workflows/, gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every file a PR deletes. It has no semantic awareness, so a deleted template's basename colliding with something else entirely still fires: - claude-md.md: gsd-core/templates/README.md had a stale table row claiming /gsd-profile reads this template to generate CLAUDE.md. Verified false (no code reads it anywhere, same search that already covered bin/install.js) — fixed the row to *(inline)*, matching every other command-generated artifact in that table. File stays deleted. - codebase/testing.md: collided with docs/guides/testing.md, an illustrative example row in docs-update.md's sample output table (an unrelated real generated-docs path). Swapped the example topic to "contributing" — the row is illustrative, any topic works. File stays deleted. - codebase/architecture.md, codebase/stack.md: collided with docs/reference/ planning-artifacts.md's directory listing of a user's own generated .planning/codebase/architecture.md and stack.md output — the same semantic mismatch already investigated and dismissed as unrelated earlier in this phase's audit, now caught by a gate instead of judgment. That listing repeats across 5 locale copies of the doc. - continue-here.md: collided with the real .continue-here.md pause-work artifact, referenced across 15+ locale and workflow files. For the last two, the lint's own error message offers "restore the file or update every consumer in the same commit." Rewording 15+ files across languages I cannot verify translation quality for, to shave 2 already-tiny templates that were merely presumed dead, is disproportionate to this PR's actual scope — restored codebase/architecture.md, codebase/stack.md, and continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates section accordingly. Final confirmed-dead set: claude-md.md, codebase/concerns.md, codebase/conventions.md, codebase/integrations.md, codebase/structure.md, codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next node scripts/lint-removed-but-needed.cjs now passes clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes * fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the user asked to be actually fixed, not just re-run past: PR #4497 (landed 2026-09-07, one day before this PR's CI run) isolated tests/codex-config.test.cjs into its own dedicated chunk because its measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe to share a chunk with any other file. That isolation was necessary but not sufficient — even alone, with zero companion-file contention, the file's real Windows execution time sits right at the 600s per-chunk ceiling. Two independent CI runs on two unrelated PRs (this one and #4154) were both killed within ~1.4s of the identical 600000ms mark — not random contention, a deterministic near-miss the isolation fix couldn't address because it never reduced the file's own cost, only removed the risk of a companion file's cost stacking on top of it (which the PR #4497 comment explicitly anticipated: "if a future profiling pass genuinely speeds up codex-config.test.cjs itself, this isolation can be revisited"). The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79 describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760, #3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more), several of which are explicitly documented as "folded" in from separate files that were never actually split back out ("Verified non-duplicate against both the pre-existing target and the other three folded sources"). Split into 4 files by top-level AST statement boundaries (never a naive column-0 regex — an early attempt at that overcounted 79 apparent "describe(" matches when only 21 are genuinely top-level; the rest are nested inside a handful of large folded-in blocks, which a regex can't tell apart from real top-level statements). Verified lossless twice: the split script asserts byte-for-byte reconstruction of every source character, and independently, total test()/describe() call counts match exactly between the original file and the sum across all 4 new files (433/79 both sides). Each new file carries the complete original shared header (imports/helpers) for safety; per-file unused-import warnings from that duplication are resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard form for an intentionally-unused destructured binding — never a bare `{ _foo }`, which would destructure a different, nonexistent property). No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its pinned test in tests/run-tests-harness.test.cjs: the file that keeps the original name (tests/codex-config.test.cjs) is now only ~28% of the original's size and safely isolated in its own chunk as before; the other three new files re-enter normal weight-balanced packing, none individually close to disproportionate. Confirmed no other file hardcodes the hardcoded filename anywhere that would silently stop these tests from running (the CI test-selection scripts determine scope algorithmically, not by literal filename). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b8c70a33c0 |
fix(#4454): surface skipped explicit --files paths instead of silent success (#4538)
* fix(#4454): surface skipped explicit --files paths instead of silent success cmdCommit's staging loop deliberately skips a --files path that no longer exists on disk when the caller passed --files explicitly (#2014 guards against staging an unwanted deletion for a temporarily-absent file), but recorded nothing about the skip. The final result reported unqualified committed: true, so a caller had no way to distinguish "everything in --files landed" from "some paths were silently excluded because they didn't exist" -- a real file deletion meant to be committed (e.g. phase.complete removing .planning/milestone.lock) would silently never land, with git status still showing it unstaged after a "successful" commit. Tracks skipped paths in a skippedFiles array and surfaces them as skipped_files (matching this result family's existing snake_case precedent, timed_out) in the success result AND the nothing_to_commit result (reachable when every named path was missing), included only when non-empty so the common case's payload shape is unchanged. Also extended to the SECOND nothing_to_commit result (reached when git itself reports "nothing to commit" after the nothingToCommit guard was false -- the residual partial-skip + partialCommitRefused window the surrounding comments already document) for the same consistency. The #2014 staging/deletion guard itself is untouched -- purely a visibility fix, exactly as the issue requested. Regression tests cover: existing-plus-missing file (the issue's own repro shape, with the missing path genuinely tracked-then-deleted so the #2014 assertion is meaningful, not vacuous against an never-tracked path); only-existing files (no skipped_files key at all); only-missing files (nothing_to_commit with skipped_files); default mode unaffected; and multiple missing files reported in order. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: recover ack trailers buried by squash-merge concatenation Discovered while verifying an unrelated PR (#4454): tests/emitted-attribution.test.cjs failed citing pr-branch.md's growth as unacknowledged, even though PR #4537 (#4447) had genuinely acked it via an Emitted-Drift-Ack-Growth commit trailer. Root cause: git's %(trailers:...) placeholder (used by readAckTrailers) only recognizes a trailer block that is the TRUE TERMINAL block of a commit message. GitHub's squash-merge commit body concatenates every constituent commit's subject+body in order, then unconditionally appends its own `---------` separator and Co-authored-by trailers. Confirmed directly against a real squash commit: %(trailers) returns ONLY GitHub's own two Co-authored-by lines -- not even the LAST original commit's own trailer survives, since GitHub's appended suffix breaks the backward scan before it ever reaches past `---------` to any original commit's content, last or not. An ack trailer added on any non-final commit of a PR (the common case -- more commits routinely land after an ack, e.g. a lint fix or a changeset) is therefore silently invisible forever to any future comparison against a base predating that squash. Fix: readAckTrailers now runs a second pass (extractSquashBuriedTrailers) alongside the existing whole-message read. For each commit in range, if its raw body contains GitHub's squash-suffix signal, the pre-suffix text (everything before the LAST such signal -- an earlier bullet's own body may legitimately contain a markdown horizontal rule that looks the same, so anchoring on the first occurrence would truncate too early and miss a later bullet's real ack) is split on squash-bullet (`* <subject>`) boundaries, and each chunk is independently trailer-parsed via `git interpret-trailers --parse` -- the same underlying algorithm as %(trailers:...), but runnable against arbitrary text rather than only a real commit object. This finds a trailer buried in ANY bullet, not just the last one. Scoped tightly to avoid reintroducing the false-positive class %(trailers:...) was originally chosen to prevent (a mid-body MENTION of trailer syntax must stay inert): the sub-chunk pass activates only on commits matching the squash-suffix signal, so an ordinary commit whose body happens to contain markdown bullets is completely unaffected, and each chunk still goes through git's own strict per-chunk terminal-block detection. Verified end-to-end against a real squash commit (recovers the buried trailer) and four adversarial fixtures now pinned as regression tests: ordinary bullet prose with no squash suffix (stays empty); a squash-shaped commit where one bullet's body merely mentions trailer syntax mid-paragraph (stays inert, the "row 32" false-positive class, now verified at per-chunk granularity); and an earlier bullet's own markdown horizontal rule not truncating the scan before a later bullet's real ack (the last-match-not-first-match case an isolated review pass caught during this same fix). This overrides one-concern-per-PR per CLAUDE.md's Defects & Warnings policy -- a genuine defect discovered mid-work is fixed inline, not deferred to a separate issue (spawn_task for this was correctly blocked by the no-defer guard). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4454): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e03921c7d8 | enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) | ||
|
|
a0f8f956c4 |
enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior" section lists four checks; the ADR's Decision 5 and its own phase table ("partition rules + the five checks") list five — the same four plus "boundary moves are declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR implements all five. docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule itself, the protected-content list and <!-- gsd:protected --> sentinel syntax (relocated unchanged from gsd-core/references/compact-content-protected-content.md, now deleted — it was never referenced by any runtime workflow Read, only by the predecessor test as documentation, so nothing at runtime regresses, and removing it from gsd-core/references/ also drops it from all 19 installed-project shipped-content trees for a file nothing ever read), and the five checks explained for a human reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped content". tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no registry file, a pair is registered by existing on disk), line normalization (carries forward Phase 2's bare-label-line isTrivial fix and the canonical gsd_run-launcher-preamble exclusion), sentinel extraction, and a Boundary-Move-Declared commit-trailer reader that is a direct structural port of tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader (ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range, same dedupe/conflict rules. tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3 (registration + size cap) run unconditionally against every registered split. Checks 1 (completeness, fires once per split on the PR that introduces a new detail/ path), 4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5 (boundary moves declared) are PR-diff-scoped against the resolved base ref and skip cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a deliberate asymmetry from check 5's trailer reader, which must throw rather than silently pass when ITS range is uncomputable, since that function is answering "did this PR declare its moves" rather than "is there even a diff to look at". Each of the five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair, built against synthetic temp files or real throwaway git repos, per this repo's rule that a guard nobody has seen go red is not yet a guard. Building the real fixtures caught and fixed one real bug before it shipped: check 4's line-presence test was using the trivial-line-filtered normalizer, so a byte-identical spine falsely reported its own protected code-fence line as "deleted" — fixed with a non-filtering membership check. Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md is a fourth workflow sub-file kind that already existed on disk and was already known to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same pattern, same PR, rather than deferred. Also, mechanically required by the new fourth sub-file kind: - scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS alongside modes/steps/templates — a detail/<part>.md inherits its parent's response_language coverage through the same per-file proof, not a parallel one. - tests/workflow-size-budget.test.cjs: explicit regression test locking that detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers — true by construction (measureWorkflows/listWorkflowStems don't recurse), made explicit per the issue's own Done-when item rather than left true-by-omission. - scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates` NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked); docs/INVENTORY.md's table updated to four kinds. Verified: `npm run lint:ci` clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): review findings + a real gsd-test failure in the new guard Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a separate security review ran against the prior commit. Fixed everything each surfaced: - Security (Low, path-traversal existence oracle): checkRegistration's dangling-reference check extracted detail-path-shaped substrings from spine PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot and probed fs.existsSync with no containment check — a spine file containing `../../../etc/detail/passwd.md`-shaped text could make the guard test file existence outside the repo. Added a path.relative-based containment check before the fs.existsSync call; anything that resolves outside repoRoot is now reported as a dangling reference directly, never probed on disk. - Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case CLAUDE.md's TEST RULES require (limit-1/limit/limit+1). - Standards (Property-Based Testing): extractProtectedBlocks (a sentinel parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict parser, bijective-shaped) had no fast-check property test. Added three: a render/parse bijectivity property for the trailer parser (mirroring the exact ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for the same parser, and a well-formed-sentinel-round-trips property for extractProtectedBlocks. Then dispatched gsd-test on the resulting commit. It found a real bug the reviews couldn't have caught (none of them can run inside gsd-test's sandbox): checks 4/5's real-repo assertion failed against plan-phase's own split, reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" — content Phase 2 (#4402) legitimately moved into detail/elaboration.md months before this PR's Boundary-Move-Declared mechanism existed to require a trailer for it. Root cause: `resolveBase()`'s own doc comment already documents that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox container, and its fallback candidate (a bare `next` branch) can resolve to a point in history that predates an already-merged, already-reviewed split — making that split look "newly introduced" from the sandbox's vantage point. Check 1 (completeness) already scopes itself correctly to only genuinely-new detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping, so a stale base made them re-litigate a settled split retroactively. Fixed by having checks 4/5 skip any split name check 1 already counted as newly-split — their own premise ("did an EXISTING split shed/undeclare something") does not apply to a split that is, from the resolved base's vantage point, brand new; that is check 1's domain alone. Verified locally (25/25 tests pass via a direct `node -e` require, since `node --test` is blocked in this repo) and via re-reasoning through the exact real-repo scenario the gsd-test failure showed. Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the prior commit's deletion of gsd-core/references/compact-content-protected-content.md was never reflected there, which is what golden-install-tree.test.cjs's other 19 failures in the same gsd-test run were. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4403): backfill changeset pr number to 4497 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs (weight 17.87, by far the chunk's dominant cost) packed alongside 39 other files. Traced, not assumed: - scripts/run-tests.cjs's own timeout-headroom comment for the OUTER per-shard timeout documents that "adding one test file reshuffled 115 of 268 unit files between shards" — shard/chunk composition is architecturally known to be unstable to single-file additions, which is exactly what this PR's own new tests/compact-content-partition-guard.test.cjs is. - A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own CI), already documents the SAME chunk hitting the SAME 600s backstop with the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of 17.87 — not a stale-table miss") dominating it — the fix then was cutting the Windows per-chunk budget from 60 to 40. That cut clearly was not enough: two documented incidents in two days, at two different budget settings, both centered on one file that alone consumes ~45% of even the reduced Windows budget. - tests/test-timings.json's own header confirms its source data (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the recorded figure" — the packer's weight-balancing is working off data that is both stale (table last regenerated 2026-08-07) and known to underestimate the platform where the failure occurs. Given codex-config.test.cjs is disproportionately heavy AND every companion sharing its chunk is decided by a packing algorithm already documented as reshuffling unpredictably on any new file, tuning the shared budget a third time only moves the marginal line to wherever the next new file happens to land — it does not remove the gamble. Isolating codex-config.test.cjs into its own dedicated single-file chunk, unconditionally and on every platform, removes it at the source: the file never enters the pool packChunks balances, so no other file's packing changes, and no future single-file addition (mine or anyone else's) can silently reintroduce this exact failure by landing in its chunk. Extracted as a small pure function, partitionIsolatedFiles (mirroring this file's existing pattern of pulling packing/analysis logic out of main() for in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents), with 6 new tests in tests/run-tests-harness.test.cjs covering basename matching across path separators, near-miss non-matches, the empty-list case, and the isolated-set contents. Root cause is now closed rather than papered over with a retry: this failure is a property of one specific heavy file's chunk placement, not something that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped up, this isolation can be revisited — this is a packing-side mitigation for a known file's cost, not a claim the cost is irreducible. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93abf3a7b6 |
fix(#4276): read the quoting, not just the digits, before eating an IO number (#4420)
POSIX recognizes an IO number only when the digit run is unquoted, but tokenize()'s redirection branch tested `cur` — the token's characters — and never `curMask`, their quoting. `"2"` and `2` were indistinguishable to it, so a correct `gsd_run query commit ... --files "2">out` had its value consumed as an IO number and scored as an unscoped invocation: the #2269 guard reddening on documentation that was right. The mask was already maintained in the same loop and thrown away on this branch. One conjunct reads it: '0' marks a bare character, so /^0+$/ asks precisely whether every character of the digit run was unquoted. The regression arms come in both directions. The quoted rows (glued, detached, single-quoted, and the partially-quoted 2"3" that a some-character-unquoted test would get wrong) must be scoped; the existing unquoted `--files 2>&1` row is the negative control and must stay unscoped, so an implementation that simply stopped consuming IO numbers altogether fails instead of passing. No live instance exists in any of the six scan roots, so this is latent rather than urgent — a false positive, never a silent miss. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4c60879b5d |
fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources * chore(#4132): record PR number in changeset * test(#4132): cover rejected commands source alias * fix(#4132): reject aliased package fallback * test(#4132): cover rejected agents source alias * test(#4132): cover partially aliased marker provider * fix(#4132): reject partially aliased source providers * test(#4132): cover routed source identity probes * fix(#4132): route installed source identity probes * refactor(#4132): tighten installer source metadata * test(#4132): cover corpus trust boundary attacks * fix(#4132): close installed corpus trust gaps * refactor(#4132): keep installer authority private * fix(#4132): preserve private installer fallback * test(#4132): preserve fixture source authority * fix(#4132): reject overlapping source fallback * fix(#4132): avoid redundant installed corpus reads * refactor(#4132): simplify provider resolution * test(#4132): sync install tree fixtures after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5d804dd287 |
fix(#3709): clear the context-monitor warn sentinel on PreCompact (#3808)
* fix(#3709): clear the context-monitor warn sentinel on PreCompact The monitor's per-session warn sentinel survived a compaction, so once the first CRITICAL of a session had fired, `lastLevel` stayed pinned at 'critical' for the rest of the run. The hook was already wired to PreCompact (#772), but read the event only at the very END, and solely to pick an output envelope. Two documented behaviours died as a result: - "First warning always fires immediately" — the first warning of the post-compaction cycle was debounced instead. - "Severity escalation (WARNING -> CRITICAL) bypasses debounce" — computed as `lastLevel === 'warning'`, which can never be true again, so every later CRITICAL waited out the full five-tool-use debounce, exactly when an immediate warning matters most. `criticalRecorded` was equally sticky: a session that compacted and later truly ran out kept a /gsd:resume-work breadcrumb (#1974) describing the earlier near-miss rather than the exhaustion that ended the run. Reproduced first, with the issue's own literal repro, including the detail that the compaction consumed a debounce slot (callsSinceWarn 0 -> 1). The reset runs BEFORE the metrics read, deliberately: a post-compaction reading is healthy again, so the ENOENT / stale / above-threshold branches would all exit first and never reach it. Returning early also stops the compaction from eating a slot of the cycle it was meant to restart. The event name is now read once through a shared `readEventName()` helper, so this reset and the #2289 output allowlist cannot drift on what counts as "no event name". Seven rows against a real sequence (the defect is state carried ACROSS calls, so they need their own driver — the existing helpers delete the sentinel after each invocation). Reverting the reset turns SIX of them red; the seventh is the non-vacuity row asserting a NON-compaction event must not clear the sentinel, which correctly passes either way. AC4 initially passed with and without the fix — asserting `criticalRecorded === true` is vacuous when the seeded stale sentinel already carries it. It now seeds a `staleProbe` marker that can only survive if the sentinel survives, so its absence is what proves the state was rebuilt. hooks/dist/ is gitignored and regenerated by build:hooks, so no committed dist copy needs syncing. Verified: `npm run lint:ci` exit 0; acceptance criteria 1-6 driven end-to-end against the real hook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): reset ahead of the config gate, and pin the placement itself Codex review of the #3709 fix, before opening the PR. Three findings, all in this change's own new code. 1. `context_warnings: false` prevented the reset. The config early-exit sits ABOVE where the reset was placed, so a session that disabled warnings, compacted, then re-enabled them mid-session resurrected the stale sentinel and the original bug with it. Config is re-read per invocation, so that sequence is supported rather than hypothetical. The reset now runs ahead of the config gate: clearing the sentinel is CLEANUP, not a warning — state that must not outlive a compaction should not outlive it merely because warnings are switched off right now. It cannot emit anything from there, so the disabled contract is untouched. 2. Nothing pinned the "before the metrics read" placement. Every row wrote a fresh metrics file, so the reset could have been moved below the metrics read, the stale check, or the healthy-threshold exit with all seven rows still green — while a REAL PreCompact, which carries no fresh metrics and follows a recovery to healthy usage, silently kept its sentinel. Three rows now pin it: no metrics file at all, usage recovered to healthy, and warnings disabled. Each catches a distinct wrong placement — moving the reset below the config check reds the third; below the metrics read reds all three. 3. The absent-sentinel row proved nothing. `assert.doesNotThrow` was vacuous because the driver caught every child exit, so a hook that exited 1 on the ENOENT unlink would still have passed. The driver now returns the exit code and the row asserts it is 0. Also corrected the `readEventName` comment: it said the event is "read once", which is not literally true — there are two call sites. The point is one DEFINITION of what counts as an event name, so the reset and the #2289 allowlist cannot drift; the comment now says that. Verified: 60 rows in tests/perf-317-context-monitor-fs.test.cjs, 0 fail, with both placement mutations driven to red and reverted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#3709): backfill changeset pr number The fragment shipped with the documented `pr: 0` placeholder, which the changeset lint treats as always-silent, because the PR number does not exist until the PR is opened. Backfilled to 3808 now that it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): a compaction clears the stale reading too, not just the state Review round 1. Major 1 was right and it mattered: clearing only the sentinel traded a warning that never fires for one that fires when it must not. The statusline bridge still holds the PRE-compaction reading, and STALE_SECONDS is 60, so for up to a minute it still reads fresh and still says the context is exhausted. With the sentinel gone, firstWarn is true, so the next PostToolUse emitted a spurious CONTEXT CRITICAL immediately after the compaction that FREED the context — and flipped criticalRecorded, spawning a false context-exhaustion breadcrumb. That is the same breadcrumb inaccuracy #3709 exists to fix, re-entered from the other side. Reproduced before fixing, exactly as the review described. A compaction now invalidates the warning state AND the reading that produced it. Removing the bridge loses nothing: the statusline owns that file and rewrites it on every render, and its absence is already the "no reading yet" state a fresh session starts in, which exits silently. Two things my own verification caught while fixing it: - The first attempt did NOTHING. metricsPath was declared below the PreCompact block, so referencing it hit the temporal dead zone, threw, and the outer catch swallowed it into a silent exit 0. The probe still printed "silent", which looked like success but was the old debounce. metricsPath is now hoisted beside warnPath. - The new Major 1 row was VACUOUS. The driver's `metrics: false` DELETES the bridge, but the defect is a bridge that is still there and still reads fresh, so the row passed on the ENOENT early-exit rather than on the fix. Only the sentinel-only mutation exposed it. The driver grew a `metrics: 'keep'` mode that leaves the stale file in place; both Major 1 rows now red under that mutation. Also from the review: - Minor 1 — the compaction-abort path is now stated in the source rather than left silent, including why a conditional reset (SessionStart source "compact") is out of scope for this fix. - Minor 2 — docs/context-monitor.md completed: PreCompact wiring and the early return under How It Works, a table of all three things the reset clears, the breadcrumb guard, the warnings-disabled interaction, and the never-block property under Safety. - Minor 3 — changeset trimmed from ~1,400 chars of implementation narration to the user-visible change. - Nit 1 — a failed unlink (Windows EPERM/EBUSY) no longer leaves the bug silently intact: the file is neutralised in place instead, with a shape safe for each (an empty sentinel, a timestamp-0 bridge). - Nit 2 — reviewer-process narration removed from shipped test source. The remaining "Codex" mentions are pre-existing and name the RUNTIME. - Nit 3 — the debounce-slot row now asserts the observable consequence (the first post-compaction warning fires) rather than repeating AC1's assertion. - Nit 4 — the file docblock now lists the folded-in blocks and asks the next contributor to extend it. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3709): the unlink-failure fallback truncates to empty, matching deletion The fallback wrote well-formed neutral values, and neither was equivalent to the deletion it stood in for: '{}' parses, so firstWarn was false and the first post-compaction warning was debounced — AC2 undone on exactly the path the fallback exists for — and '{"timestamp":0}' was never stale (the guard is `metrics.timestamp && ...`), so the flow reached emit with remaining === undefined and injected a literal 'Usage at undefined%'. Truncating to '' makes JSON.parse throw on both reads: the sentinel read keeps firstWarn true, the bridge read falls to the outer catch and exits 0 silently (review of #3808, Blocker 1). The branch is now executed for real: an EPERM is injected into the child's fs.unlinkSync via --require preload — method monkeypatching, never chmod 0o000, which root bypasses under Docker/CI (Blocker 2). Both rows proved failing-first against the neutral-value fallback. The boundary trios at WARNING=35 / CRITICAL=25 are completed on the emit path with 34, 26, and 24 (Major 3); 36/35/25 were already pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): the truncation fallback refuses to follow a planted symlink The per-session files live in a shared sticky tmpdir, where an unlink failing EPERM is exactly what another user's planted file produces — and a planted SYMLINK would make the fallback's plain truncating write empty out its TARGET, weaponising the hook against any file its own user can write. Open with O_WRONLY|O_TRUNC|O_NOFOLLOW instead: a symlink fails ELOOP into the same give-up arm. On Windows the constant is absent and '|| 0' keeps the fallback alive there, where the held-handle case it exists for occurs and temp dirs are per-user. Found by Codex review; the new row proved failing-first against the writeFileSync fallback (victim file truncated to zero bytes). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): refuse non-regular files everywhere, not only where O_NOFOLLOW exists Codex round 2: '|| 0' removed the no-follow protection exactly where it cannot be expressed as an open flag — Windows, whose tmpdir is NOT guaranteed per-user (TEMP/TMP overrides, system-temp fallback). An lstat isFile() guard now rejects symlinks and every other non-regular shape on all platforms before the truncating open; O_NOFOLLOW stays, as the lstat->open substitution-race backstop where the platform has it. The symlink row additionally asserts the planted link SURVIVES the call, so a preload match that stops engaging can no longer pass the row vacuously off a successful unlink. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(#3709): tolerate the Windows give-up, still outlaw neutral values Both windows-latest CI lanes fail the two EPERM rows deterministically: the runners hold freshly written files with a share mode that allows DELETE (every real-unlink row passes) but refuses a truncating write-open, so the fallback's give-up arm engages — which is the fallback working as designed, not the defect the rows exist to catch. The rows are now platform-aware: POSIX still requires exact truncation and the behavioural follow-ons; Windows accepts truncated-or-untouched but still rejects the Blocker-1 regression class (a parseable neutral value is never legal anywhere), with the follow-ons gated on the truncation actually landing. Also corrects the hook comment: libuv defines O_NOFOLLOW as 0 on Windows — a no-op, not an absent constant. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): a compaction watermark closes the window bridge deletion only narrowed Round-3 Major 1: the statusline is an uncoordinated process that re-writes the bridge on every render, so a render landing between the PreCompact clear and the compaction's completion re-created the PRE-compaction reading under a CURRENT timestamp — past STALE_SECONDS, into a spurious post-compaction CRITICAL and a false exhaustion breadcrumb: the exact failure the deletion was added to prevent. PreCompact now also writes claude-ctx-<id>-compacted.json ({at}) and the metrics read drops any reading not STRICTLY newer than it — which also covers unstamped/zero timestamps once a compaction happened. Written unlink-then-O_EXCL so a planted file or symlink is never followed; failure degrades to the old narrowing. Docs and changeset now describe the watermark instead of overclaiming for the deletion. Round-3 Major 2: DEBOUNCE_CALLS and STALE_SECONDS get their trios — the gate increments BEFORE comparing, so seeds 3/4/5 pin 4-debounced, 5-emits, 6-emits; ages 59/60/61 pin the strict >. The child's clock is pinned via a --require preload (a wall-clock boundary row would flip on one second of startup delay). timestamp-0's falsy bypass is pinned directly as characterized behaviour. Mutation-proven: dropping O_NOFOLLOW, <= for <, and >= for > each red exactly one row. Minors: the symlink row's comment now names the lstat guard it actually pins, and a preload-blinded-lstat row drives the O_NOFOLLOW substitution -race backstop for real (3); absence assertions use warnRaw so a corrupt leftover cannot pass as deleted (4); the Windows give-up is an explicit t.skip, never a silent if (5); readEventName is total via String(), keeping #2289's side-effects-always-run contract for malformed event names, with a row (6); the PreCompact rationale lives once in docs/context-monitor.md with the code keeping only line-level constraints (9); the changeset is release-note-sized (10). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): the grace window covers the compaction's duration, not just its start Codex on the first watermark cut: the watermark stamps the compaction's START, so a statusline render one second later — still mid-compaction, still the old reading — passed 'strictly newer' and re-fired the false CRITICAL. Readings inside COMPACT_GRACE_SECONDS (60) past the watermark are now dropped: the window covers the compaction's own duration, a healthy reading dropped there behaves identically to an accepted one (it exits above-threshold anyway), and a genuine exhaustion warning is delayed at most one window after a compact. A watermark stamped ahead of the reader's clock is ignored — a clock step backwards or a stray file must degrade to plain staleness, never mute the monitor indefinitely. Both proven failing-first. readEventName is strict about TYPE, not coerced: String() rendered ['PreCompact'] as 'PreCompact' and would run the reset off a malformed payload. typeof: every non-string is 'no event' — silent, side effects intact — with rows for the number, hostile-object, and array-wrapped cases. The lstat-claim preload arm now writes an engagement marker the substitution-race row asserts on, so a match string that silently stops matching can no longer let the row pass off the real lstat guard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: retrigger CI — the previous wave was cancelled by an Actions outage, zero job failures Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3709): drive the compaction rows on the clock, not on a future stamp Round-4 review raised three majors, all in the test scaffolding around the fix rather than in the fix itself. Major 2 (taken first — it is the cheapest and it unblocks Minor 6): call() passed process.env to the child unmodified, so two rows depended on ambient GEMINI_API_KEY. The preserved Gemini fallback is `eventName === "" && !!process.env.GEMINI_API_KEY`, and readEventName returns "" for every malformed name, so with the key set the malformed-event row's `stdout === ''` assertion failed outright — reproduced by running it under GEMINI_API_KEY=x. call() now takes an explicit env, the way the sibling runMonitor helper in this file always has, and both rows pin the variable unset. (The array row survived an ambient key only because its reading was debounced — incidental, not independence, so it is pinned too.) Major 1: the AC2/AC3 rows drove the hook with a bridge stamped 62 seconds in the FUTURE — a shape hooks/gsd-statusline.js cannot produce, since it always stamps Math.floor(Date.now()/1000) on the same clock. They proved "the sentinel was cleared" while their assertion messages claimed the documented immediate-warning behaviour, which is gated behind the grace window and went unexercised. Both rows now run the real sequence on the clock-pinning preload this PR already added for the STALE trio: PreCompact at a fixed instant, then a normally-stamped render one second past the window. Verified non-vacuous — stubbing the sentinel unlink reds both. Major 3: COMPACT_GRACE_SECONDS, the one constant this PR introduces, was the only threshold without a limit-1/limit/limit+1 trio, in a PR that adds full trios for four pre-existing ones. The seeded offsets were +0, +1 and +61; the boundary itself (+60) and limit-1 (+59) were untested. Added, driven by advancing the reader's clock rather than post-dating the reading, so the reading is never ahead of the reader and only the grace gate can drop it. Verified against three mutations — `>` to `>=`, the constant to 59, and the constant to 61 — each of which reds exactly one row of the trio. No production code changed. Verified: 85/85 in this file, lint:ci exit 0, and the two Minor-6 rows now pass under GEMINI_API_KEY=x as well as unset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): harden the watermark read and pin the thresholds it introduces Codex review of the full PR found four majors. All reproduced here against the real hook before fixing. MAJOR — the watermark was write-hardened but read-untrusted. PreCompact already refuses to follow or overwrite a planted object (unlink-then-O_EXCL), but the read was a bare readFileSync, so anything the write side gave up on was followed by every later invocation. In a shared sticky os.tmpdir() that is a mute primitive — a planted recent watermark suppresses monitoring — and a symlink to a FIFO stalls a synchronous read. Measured against the pre-hardening file: a symlink to a planted watermark WAS honored and muted the monitor. The read now uses the same lstat + O_NOFOLLOW pair the sentinel path uses, plus a size bound; symlink, directory and oversized cases are all refused, with a plain-file control proving watermarks still work. MAJOR — the `now + 5` skew tolerance was an unnamed, untested threshold. It is now WATERMARK_SKEW_SECONDS with a +4/+5/+6 trio, verified against two mutations (`<=` to `<`, and the constant to 6), each of which reds one row. This is the same class as round 4's Major 3, one layer up. MAJOR — the malformed-event row shared one session across both subcases, so the hostile-object iteration's `assert.ok(s.warn())` passed off the sentinel the `42` iteration left behind. A regression throwing before the bookkeeping would have kept it green — vacuous for exactly the subcase it exists for. Fresh session per subcase, with an explicit no-sentinel precondition. MAJOR — the stale-reading row's non-vacuity is an artifact of call()'s future stamp: with a production stamp the watermark suppresses the same reading, so the row cannot isolate bridge deletion. The two guards genuinely overlap inside the window, so no end-to-end row can separate them; the comment now says so and points at the direct pin (s.metrics() === null) instead of claiming an isolation it does not have. Docs corrected where measurement contradicted them: the window NARROWS the race rather than covering the compaction's duration, and the delay is not bounded by the window alone — first recovery is watermark+61s with no skew but watermark+66s at the accepted +5s skew. Aborted compactions are muted the same way. The truncation fallback is documented as best-effort, which is what the code and the Windows rows already do. Verified: 89/89 in this file, lint:ci exit 0, symlink/directory/oversize all refused where the pre-hardening file honored them, both new trios mutation-checked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): move the PR's two new exits onto the declared-policy vocabulary #3911 / ADR-3889 migrated this hook off raw process.exit() while this PR was in review, replacing every exit with hooks/lib/hook-exit.js's allow(), which forces each call site to name its crash policy. The PreCompact reset and the watermark gate are added by THIS PR, so they did not exist to be migrated and came through the merge as the only two raw exits left in the file — caught by the new local/require-registered-exit rule. Both are ALLOW: a compaction is never blocked by this hook, which is the policy the rest of the file declares. Caught only in CI, not locally: `npm run lint` runs eslint with --cache, and the cached entry for this file predated the new rule, so a warm local cache reported clean. Re-verified with the cache cleared. allow() terminates rather than throwing, which matters for the watermark call site because it sits inside a try/catch — a throwing helper would unwind into that catch and silently drop the grace-window mute. Verified behaviourally, not by reading: the grace trio, the skew trio and the non-regular-file rows all still pass. Verified: lint:ci exit 0 with a cold eslint cache, full suite exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3709): harden the routine sentinel writes and the read beside them Round 7 ruled that the three routine debounce-accounting writes to the warn sentinel must match the three writes this PR already hardened: leaving the fourth unhardened beside them is the asymmetry that invites the defect back. They now go through one writeSentinel() helper using the compaction watermark's own unlink-then-O_EXCL shape, rather than a second policy — the unlink removes any existing object, and O_EXCL then refuses to create through one, so a write can only land on a fresh regular file this process made. The routine READ beside them was the last bare readFileSync on warnPath, and the same rationale applies to it verbatim; the watermark's read was hardened in round 4 for exactly this reason. Same lstat + O_NOFOLLOW + size bound. Its scope is stated in the test rather than overclaimed: lstat establishes that the sentinel is a plain regular file, not that it is trustworthy, so a cross-owner regular file at the predictable path is still read and is left as a disclosed pre-existing residual. Also fixes an accept-direction regression this PR introduced and six rounds of review missed. readEventName collapsed an ABSENT event name and a MALFORMED one onto the same '', and the preserved Gemini fallback keys off eventName === "", so with GEMINI_API_KEY set a malformed payload began emitting an AfterTool envelope. At the merge-base, data.hook_event_name.trim() threw on a truthy non-string after the side effects and nothing was ever emitted. Measured base-vs-head with a fresh sentinel per run: 42, ['PreCompact'] and {} all went silent -> EMITS, while an absent name and 'PostToolUse' were unchanged. readEventName now returns '' only for an absent name and null for a present-but-non-string one; both call sites compare for equality only, so every well-formed payload behaves identically. Five new rows, each proven fail-first with the mutations attributed separately: reverting the writes reds the write-through and non-regular rows, reverting the read reds the mute and non-regular rows, and reverting the absent/malformed split reds the Gemini row. The changeset's "behaves like a fresh session" is narrowed to name the 60-second suppression window and the best-effort reset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9 * test(#3709): pin both 4096-byte read bounds at their boundaries Round 8 asked for limit-1/limit/limit+1 coverage on the size bound the round-7 sentinel read-hardening introduced (gsd-context-monitor.js:335). The existing refusal row pads to 8192 -- a full 4096 bytes clear of the fence -- so `>` vs `>=`, or an off-by-one in the constant itself, was invisible to it. Covers the sibling bound too. The identical check guards the round-4 WATERMARK read at :278 and its refusal row pads to 8192 in exactly the same way; the review's own rationale (this file already holds WATERMARK_SKEW_SECONDS to a boundary trio, so an uncovered bound is the odd one out) applies to it unchanged. That half is a class sweep of a pre-existing bound and is test-only -- say the word and it comes out without touching the rest. Both trios assert on observable hook output rather than an internal error. Sentinel: an honored {callsSinceWarn:1,lastLevel:'warning'} keeps the debounce arm taken at remaining=30, so nothing is emitted, while a refused one falls back to first-warn defaults and emits. Watermark: honored mutes (stdout empty), refused leaves the warning. The 4097 row is the non-vacuity control for the two accept rows. Payloads are sized by measurement, with Buffer.byteLength asserted to equal the target, not by arithmetic on an assumed prefix width. Proven fail-first in both directions, with the hook restored after: `> 4096` -> `>= 4096` reds both trios (94/96); `> 4096` -> `> 4097` reds both trios (94/96); restored, 96/96. Under both mutations only the two new rows fail -- the pre-existing 8192-padded rows stay green, which is the review's fencepost claim demonstrated rather than assumed. * fix(#3709): correct the changeset's mute-window claim and a superseded comment Both from the pre-push Codex pass on the full PR. The changeset said readings are "suppressed for up to 60 seconds after a compaction starts". That is false at the accepted skew boundary, and this repo's own docs/context-monitor.md already carried the accurate figure: first recovery is watermark+61s with no skew and watermark+66s for a watermark at the +5s skew limit. Measured independently at +64 silent, +65 silent, +66 warning. The changeset now states the window plus the accepted skew, matching the doc rather than contradicting it. A comment in the malformed-event row still described readEventName as returning "" for every malformed name. Round 7 superseded that: a present-but-non-string name returns null and only an ABSENT one returns "", so a malformed payload can no longer reach the Gemini fallback at all. Marked as historical and corrected. The GEMINI_API_KEY pin stays -- the row is about readEventName's typing, not the fallback, and an ambient key would still change what it measures. Codex's three Major findings are not taken, on attribution rather than logic; the reasoning is in the PR reply. In short: the watermark does not exist at the merge-base at all (0 occurrences), so "base emits, HEAD mutes" compares a new feature against its absence rather than showing a regression; and the base sentinel read is a bare readFileSync, which blocks on a planted FIFO exactly as the hardened read would, so the TOCTOU stall is not introduced here. The underlying limits -- watermark provenance, and lstat->open races on a non-symlink substitution -- are real, pre-existing, and already offered to the maintainer as follow-ups. * fix(#3709): read both sentinels through one hardened helper; state the two limits precisely Round 9's Major, with a correction to its premise, and both Minors. The review names "watermark read/write helpers this PR adds" that a call site at :238-250 duplicates inline. There are no such helpers: this PR adds readEventName and writeSentinel, the latter a write-side primitive a read cannot call, and :238-248 is base code the diff never touched. What IS duplicated is the hardened READ. The watermark read (round 4) and the warnPath read (round 7) are the same ten lines twice -- lstat, isFile and a 4096-byte bound, O_RDONLY|O_NOFOLLOW, readSync, close -- differing only in the path variable and the error string, and that is two copies to keep in step by hand. Now one function, readSentinel(target), beside writeSentinel. Refusal throws; both callers already wrapped the read in a try/catch that degrades to "no file", so behaviour is unchanged by construction. Proven rather than assumed: with the helper replaced by a bare readFileSync in a complete scratch tree, exactly the five hardened-read rows in tests/perf-317-context-monitor-fs.test.cjs go red -- round 7's symlinked and non-regular sentinel and its size bound, round 4's non-regular watermark, round 8's watermark size bound -- so the helper carries both call sites' guarantees and the rows pin it. 96/96 with the helper in place. Minor, drop vs delay: the grace-window comment said "dropped" on one line and "delayed" three lines later, and docs/context-monitor.md said "delayed". A genuine exhaustion reading inside the window is skipped, not queued: its warning and its #1974 breadcrumb both fire on the next reading after the window, so both are delayed when a later reading comes and lost when none does -- a session ending inside the window records neither. Comment and docs now say exactly that, and that the loss is accepted over trusting a reading that may be the pre-compaction value under a fresh timestamp. Minor, ordering: the PreCompact unlink and the debounce writeSentinel(warnPath) are two writers with nothing serialising them; a debounce invocation that read pre-compaction state and lands its write after the unlink would resurrect the sentinel the reset removes. The hook relies on the host dispatching a session's hooks one at a time, which Claude Code does and the other runtimes are assumed to. Stated at the reset as an assumption, with the lock-file alternative named and not taken. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3709): write the compaction watermark through writeSentinel Review of #3808, round 10. The PreCompact watermark write was the block writeSentinel was lifted from in round 7, and it kept its own inline copy of unlink-then-O_EXCL a few lines below the helper. Round 9 flagged that write-side duplication; the round-9 reply misread it as the read side and unified only the reads. The write now calls the helper too, so the hook holds one copy of the hardened write, not two. Behaviour is unchanged: same unlink-then-O_EXCL sequence, same flags, same best-effort outer catch. The one difference is that writeSentinel closes the descriptor in a finally, where the inline copy leaked it if writeSync threw before closeSync. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7bQLXAKubb4EFLtCPiLiL * fix(#3709): read the statusline bridge through the same hardening as the sentinels Review of #3808, round 11. `metricsPath` is built one line from `warnPath` and `watermarkPath` — same tmpdir, same predictable `claude-ctx-{sessionId}` shape, same threat model this PR documents at length for its siblings — and it is the only one of the three read on EVERY invocation. It was also the only one still reached by a bare `readFileSync`, so the symlink-follow and the symlink-to-FIFO stall that rounds 4 and 7 closed on the other two stayed reachable here, on the file's highest-traffic path. It now goes through `readSentinel` like the rest. The 4096-byte bound is ample for it: the statusline writes four fixed fields (`gsd-statusline.js`), about 140 bytes with a UUID session id, so no legitimate bridge approaches it. A refusal lands in the same rethrow an unreadable or malformed bridge already did. The comment introducing `readSentinel` claimed the warn sentinel was "the one bare readFileSync". Read as scoped to `warnPath` that was true, but it reads as a claim about the file and it is not one — the bridge kept its own until this round. Corrected rather than left to mislead the next reader. Round 11 Minor: `readSentinel` discarded `fs.readSync`'s return value and assumed the buffer was full, so a file truncated between the `lstat` and the read left a zero-filled tail. It now refuses a short read. Stated plainly because it was measured: this guard has NO observable behavioural delta — deleting it leaves the new row green, because the NUL tail makes `JSON.parse` throw one line later and both paths degrade to "no sentinel". It is a consistency fix in a function whose purpose is refusing to trust what it read, and the test comment says exactly that rather than implying coverage it lacks. Five rows added: the bridge refusing a planted symlink (with an attacker-chosen reading that WOULD warn if followed, so silence is proof), a non-regular bridge, an oversized bridge, the shrink path end to end, and the direction that matters most — a healthy bridge still warns, so the hardening is not a mute. Proven by mutation: reverting the bridge to `readFileSync` reddens two rows. The shrink injection carries an engagement marker for the same reason the lstat-claim one does, learned the same way: the hook rewrites the sentinel later in the invocation, so a size check afterwards passes whether the truncation landed or not. An independent full-PR pass on this round added two more, both taken: `writeSentinel` discarded `fs.writeSync`'s return value, and a short write is permitted by the syscall — so a truncated sentinel could reach disk and every later read would reject it, silently losing the debounce accounting or the watermark this write exists to record. It now loops until the payload is written, as Node's own `writeFileSync` does, with an explicit no-progress guard. Pinned by a row that injects a one-byte first write; reverting the loop reddens it. The directory row's comment claimed it pinned the `lstat` isFile() check. It does not — measured: deleting that condition leaves the row green, because reading a directory fails on its own a line later. The comment now says the row pins the outcome, and names the symlink row as the one that pins isFile(). DISCLOSED, NOT FIXED HERE — a session id long enough to push the derived filenames past NAME_MAX. The bridge is `claude-ctx-{id}.json`; the sentinel and watermark add longer suffixes, so on a 255-byte limit the watermark stops fitting at a 230-character id and the sentinel at 233. Measured base-vs-HEAD at 233+: base is SILENT, HEAD emits the warning, because the bare `writeFileSync` base used threw ENAMETOOLONG out of the warning path while `writeSentinel` degrades best-effort and lets the warning through. That is an accept-direction delta and it is in the delivering direction — base swallowed a warning the user should have seen, which is this issue's own failure class. The underlying limit is a property of the per-session filename scheme, shared by two files that predate this PR, and bounding session ids belongs to whatever writes them (`gsd-statusline.js`), not to the sentinel logic. Happy to fold a length guard in here if you would rather have it in this PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4499933807 |
fix(#3802): resolve the heredoc body before validating the commit subject (#3816)
* fix(#3802): resolve the heredoc body before validating the commit subject With hooks.community: true, gsd-validate-commit.sh blocked EVERY heredoc-form commit with CONVENTIONAL_COMMITS_VIOLATION regardless of the message, including Claude Code's own documented idiom: git commit -m "$(cat <<'EOF' feat(auth): add login flow EOF )" Reproduced before changing anything: conforming heredoc -> exit 2; plain -m "feat(auth): add login flow" -> exit 0. Root cause is the extraction regex `-m[[:space:]]+"([^"]+)"`. Bash `[^"]` matches newlines, so the capture ran from the quote after -m to the FINAL quote at `)"`, swallowing the whole span. `head -1` then returned the literal `$(cat <<'EOF'` as the subject, which can never satisfy Conventional Commits. Fixed by not answering a regex bug with another regex. hooks/lib/git-cmd.js already exists because "a naive regex misses all three" invocation forms, and extractBranchArgument is the established precedent for pulling an argument off a git command line. extractCommitSubject joins it on the same tokenizeShellLike seam — which, checked first, already returns the entire heredoc span as ONE token, leaving only "resolve the body to its first line" as new logic. Because the walk starts at the subcommand, `git -C <path> commit` and env-prefixed invocations now extract correctly too — forms the raw string scan never handled. Deliberately unchanged, and pinned as such: a glued `-mfeat: x` and `--message=...` still yield no message, exactly as the regex left them. The fix stays scoped to the reported defect rather than widening on a true observation. Two things I got wrong and corrected by measuring rather than reasoning: - I expected `git commit -m ""` to be blocked. Checked against the ORIGINAL hook: allowed before, allowed now, identical. The scanner drops the empty token so it takes the null path. My expectation was wrong, not the code. - That exposed a false comment I had just written, claiming the exit-status split prevents silently allowing `-m ""`. It does not. The split IS load-bearing, but for a heredoc whose body's first line is blank, which resolves to an empty subject and is correctly blocked. The comment now names the real case and records that `-m ""` is not it. Tests at both layers: 9 unit rows on extractCommitSubject beside its sibling in tests/worktree-safety.test.cjs, and 5 behavioral rows piping real PreToolUse payloads through the hook in tests/hooks-opt-in.test.cjs. Replacing firstLineOfMessageArg with a plain first-line return reds 8 of them across both files. (A first mutation attempt silently no-opped and reported green — the mutated body is echoed in the transcript for the run that counted.) Out of scope, per the issue: the hooks.commit_types config surface, split off by the maintainer as #3811 and explicitly sequenced after this. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): confine the fix to heredoc resolution, closing four regressions Codex review of the first attempt. It was right, and the finding is one my own rules already name: a true observation is not a licence to widen the diff. The first attempt replaced the shell's `-m` extraction with a token walk. That looked like the better abstraction — this module exists precisely because a naive regex misses invocation forms — but selecting WHICH argument is the message was never the defect, and changing it regressed four forms that upstream allowed, plus opened a bypass: - `git commit -- -m WIP` -- introduces pathspecs; `-m` is a path - `git commit --amend && echo -m WIP` a later command's flag became the message - `git commit -m "" --allow-empty-message` the shared scanner drops empty tokens, so the next flag became the message - `git commit -m WIP` unquoted argument - `-m "WIP notes <<EOF\nfix: smuggled subject"` was ALLOWED — the opener was recognised unanchored, so validation skipped past the real, non-conforming subject. An enforcement bypass, not a misclassification. Now confined to the actual defect. The shell's `-m` capture is restored byte for byte, and only the subject-from-message step is delegated, to a PURE STRING helper `resolveCommitSubject()` that never tokenizes. Verified as a differential against the upstream hook run inside the real tree: the only behaviours that change are the two intended heredoc rows (2 -> 0); all four forms above read identical, and the bypass case blocks. That differential also corrected my own control. An earlier comparison ran the upstream hook from a scratch directory, where its `lib/` could not resolve `../../gsd-core/bin/lib/token-scanner.cjs`, so the classifier failed open and reported exit 0 for everything. That made a real regression look pre-existing. Re-run inside the tree, `<<-"TAG"` (a double-quoted tag nested in the double-quoted argument) is genuinely pre-existing — the capture truncates — and is now recorded as a known limitation rather than silently "fixed". Also fixed from the review: - `<<-` strips leading TABS from body lines; returning the raw line blocked a conforming message. - a non-identifier tag such as `END-MSG` is a valid bash word and was rejected. - an immediately-following terminator is an EMPTY message, not a subject. - a node/library failure now falls back to the previous `head -1` instead of skipping validation, so a broken extractor degrades to old behaviour rather than becoming a new silent-allow path. Tests strengthened per the review: the opener-spelling rows now assert BOTH directions per spelling, since "conforming passes" alone would also pass if the resolver returned an empty subject for a spelling it failed to parse. Added differential rows pinning the five previously-allowed forms, and a row for the bypass. Dropped two rows whose comments claimed the raw scan could not handle `-C`/env-prefix invocations — it could; the claim was wrong. Replacing resolveCommitSubject with a plain first-line return reds 9 rows across both files. (Mutant body echoed in the transcript; an earlier mutation attempt on this branch silently no-opped and reported green.) Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): keep the installed hook runtime-neutral `hooks/lib/git-cmd.js` ships into every runtime, including hermes and qwen, where tests/install.test.cjs enforces that no Claude reference leaks into the installed tree. My JSDoc named the idiom after the runtime that documents it. Reworded to describe the SHAPE rather than the vendor; the runtime is still named in the changeset, which feeds CHANGELOG.md where such references are allowed, and in the tests, which are not installed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#3802): backfill changeset pr number The fragment shipped with the documented `pr: 0` placeholder, which the changeset lint treats as always-silent, because the number does not exist until the PR is opened. Backfilled to 3816 now that it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): close the truncated-capture hole, add the required test artifacts Review round 1. Major 3 was the one that mattered, and it disproved a claim I had stated in falsifiable form — the PR body said only two behaviours change; the differential found five. Major 3 — an embedded `"` truncates the `-m` capture, so the resolver received a PREFIX of the real subject and the length gate measured the wrong string. Before this fix the whole form was blocked outright, so the gate was unreachable; the fix opened the path and then mismeasured it. A new enforcement hole, so it is CLOSED here rather than declared. Closed precisely rather than bluntly. A first attempt refused to resolve any body with no terminator, which also blocked commits whose SUBJECT was intact and whose quote sat further down the body — a false positive of its own. Truncation is only fatal to the line it lands IN, and a captured line is complete exactly when another line follows it, because the capture kept its newline. So an unterminated body whose subject line is followed by more text stays measurable; only a subject line running to the end of a truncated capture falls back to the opener, which fails the format gate exactly as this form did before the fix. Major 1 — fast-check property rows for the new parser, via the shared seeded setup helper rather than requiring fast-check directly, per repo convention: totality (a security property here, since an exception on this path fails OPEN), idempotency, and that the result is always a single line drawn from the input — the third catches a resolver that concatenated or trimmed while satisfying the first two. Major 2 — the 72-char gate is now exercised at {71, 72, 73} on the RESOLVED heredoc subject, with the fixture length asserted so a mis-built fixture cannot silently pass. 92 chars did not show which side of `> 72` the code sits on. Minor 1 — leading blank body lines are skipped, as git's cleanup=whitespace does. A conforming commit written that way was still blocked, which is the same defect class #3802 reports. Nit 1 — a backslash-escaped delimiter (`<<\EOF`) is now the same delimiter rather than failing closed on a delimiter that includes the backslash. Nit 5 — changeset trimmed from 2,208 chars of design note to the user-visible change. Mutation discipline, including a correction to my own: dropping the truncation guard reds the unit rows, and the pre-review naive shape reds the hook-level row too. My first mutant did NOT distinguish the hook row — removing the guard made an empty slice and blocked for an unrelated reason, so the row passed and looked proven. Only mutating to the actual pre-review shape showed it discriminates. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): measure the subject as git does — strip trailing whitespace, split CRLF git's cleanup=whitespace strips whitespace at BOTH ends of a line; the resolver handled only the leading direction, so a 72-char subject with trailing spaces measured 75 and stayed blocked — the defect class #3802 reports, surviving one round further (review of #3816, Major 2). The resolved subject now drops trailing spaces and tabs; the plain non- heredoc path is untouched, keeping the fix confined to heredoc resolution. The length-gate boundary rows gain dirty fixtures: 72+3 trailing spaces passes, 73+1 stays blocked on LENGTH. split('\n') left \r on every body line, so on CRLF input the delimiter never matched: the truncation guard was inert, an empty message resolved to 'EOF\r', and a real 72-char subject measured 73. Split on /\r?\n/ (Minor 3). The three property tests never reached the parser — the pinned-seed fc.string corpus contained no newline and no opener, so every property reduced to f(s) === s (Major 1). The generator now constructs heredoc- shaped input (all opener spellings, <<- tabs, optional terminator, CRLF) and each property asserts a floor on inputs its corpus actually resolved. All new rows proved failing-first against the pre-fix resolver. Also records the unquoted-delimiter expansion limit as one JSDoc sentence (Informational 5). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): close two recognition bypasses, pin the dquoted-delimiter limit Codex whole-PR review found two enforcement bypasses in the resolver: - The opener's path prefix was \S*, which accepted `id;/bin/cat` — the resolver then validated the heredoc BODY while bash runs `id` first and git's real subject is id's OUTPUT. The prefix is now a path-character class; any shell metacharacter fails recognition and the form falls back to the opener line and the format gate. - The blank-line skip used JavaScript trim(), whose Unicode whitespace class skips lines git KEEPS: a NBSP first body line resolved to the SECOND line while git's real subject is the NBSP line (verified against git stripspace — the c2a0 bytes survive). Blank is now git's ASCII space/tab only; a Unicode-blank line is returned and fails the format gate, the same fail-closed direction git takes. Both proven failing-first at resolver AND hook level. Also: the <<"TAG" spelling is recorded as a documented limit — the -m capture stops at the delimiter's own quote so the caller can never deliver it (fail closed; widening the capture would change every embedded-quote case) — with a hook-level row pinning the limit; and the derivation property no longer accepts '' unconditionally, only for heredoc-shaped input, so a conditional constant-'' regression can't satisfy the corpus floor unnoticed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): recognition whitespace is ASCII, and '' answers to the generator Codex round 2: the opener's \s accepted Unicode whitespace bash does not split on — $(<NBSP>/bin/cat was recognized here while bash reads <NBSP>/bin/cat as the executable NAME, so recognition claimed a substitution that does not run cat. Every whitespace position in the recognition is now [ \t], the same ASCII rule as the blank-line skip, proven failing-first. The derivation property's ''-acceptance now consults GENERATION-TIME metadata: the heredoc generator records whether it built an empty message (terminator reachable, all scanned lines ASCII-blank, <<- tab stripping accounted for), and '' is accepted exactly then — a resolver conditionally degrading to '' on non-empty heredocs now fails, closing the residual round-1 permissiveness without re-deriving resolver logic. The changeset no longer overstates the opener spellings: it names the capture-deliverable set and the documented <<"EOF" limit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): nothing after the terminator escapes measurement Round-3 BLOCKER: everything after the heredoc terminator was silently discarded, so `-m "$(cat <<'EOF'\nfeat: ok\nEOF\n) <200 a's>"` — one 200+ char real subject once bash substitutes — measured 8 chars and dodged COMMIT_SUBJECT_TOO_LONG, a hole the base did not have. The canonical idiom's tail is exactly one closing-paren line; any other tail now falls back to the opener line and the format gate, the pre-fix behaviour for the whole form. Proven failing-first at resolver and hook level, including the glued-text and second-substitution variants. Also from round 3: `cat<<'EOF'` (no space) is legal bash and now resolves — the token before << is still literally cat; the env-prefixed and option-terminated spellings join the JSDoc KNOWN LIMIT list instead (fail closed, modelling bash prefix words is cost with no reported user); the changeset states the embedded-quote truncation limit for the message body, not just the <<"EOF" spelling; the dquoted unit and hook rows now cross-reference each other; and the fast-check setup helper's docstring no longer claims property-file exclusivity. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): glued text outside the closing quote must not shrink the measurement Codex on the round-3 guard: bash concatenates -m "$(…)"suffix into ONE argument, but the capture holds only the quoted part — so the resolver measured the heredoc body (8 chars) for a 200+ char real subject, a net-new length-gate bypass the base did not have (base measured the opener and blocked). When the closing quote is followed by anything but whitespace or end-of-command, the hook now skips the resolver and keeps the pre-fix first-line subject: the heredoc form fails the format gate exactly as on base, and the plain single-line form keeps base behavior unchanged — both pinned as differential rows, the glued-suffix row proven failing-first against the unguarded script. The property generator's ''-oracle now models the post-terminator guard it previously predated: expectEmpty requires the FIRST reachable terminator to be followed by the one canonical closing-paren line, so a resolver regressing to '' on a non-canonical tail (e.g. a body line that doubles as an early terminator) fails the derivation property instead of being blessed by stale metadata. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: retrigger CI — the previous wave never started (Actions queue stall) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): only resolve a heredoc whose body bash does not rewrite Round-4 review found two net-new enforcement bypasses: commands the base hook blocked (exit 2) that this branch allowed (exit 0). Both reproduced as a base-vs-head differential against the real hook, not inferred. The predicate "may I resolve this?" was computed from the resolver's input string alone, while two of its determinants live outside that string: 1. WHICH -m quote arm produced the input. Inside -m '...' bash performs no command substitution, so $(cat <<'EOF' is literal text and git's real subject is the opener line. The resolver ran on both arms, so all four delimiter spellings went 2 -> 0 on the sq arm — reachable by the ordinary slip of typing ' for ". The hook now records MSG_QUOTE and gates the resolver on dq; sq keeps head -1, exact base parity. 2. WHETHER the delimiter suppresses expansion. Only <<'D', <<"D" and <<\D do; a bare <<D is expanded by bash before git sees it. Resolving the literal dodged the format gate (feat: $UNSET_VAR reaches git as feat:) and the length gate (feat: ${LONG} reaches it at any length). The opener regex now separates the backslash-quoted and bare alternatives and refuses the bare one — the same fail-closed rule the metacharacter, truncation and post-terminator guards already follow. A test row asserted exit 0 for a bare-delimiter body, so the suite defended the second bypass and the fix could not land without editing a test that read as intentional. That row and its two unit counterparts now assert the block, per RULESET.TESTS.delete-bad-tests. Two unrelated rows used <<-EOF to exercise tab stripping; they move to <<-'EOF' so each tests what it names. Scoping the adjacency guard to the matched arm — required by the fix above — also removes a spurious block (round-4 Minor 1): a double-quoted heredoc whose body mentioned a glued single-quoted token tripped the sq arm. The JSDoc claimed <<"EOF" was unreachable through the caller and that the bare-delimiter gap was pre-existing. Round 4 disproved both; both corrected here, along with the matching changeset sentence. Verified: 7 bypass commands now block at head (was allow), the #3802 fix and plain-form parity are unchanged across 8 control commands, hooks-opt-in 44/44, worktree-safety 401/401, property-test non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): resolve only where the captured text is provably git's subject Codex review of the full PR found two more inputs where the validated text is not the subject git receives, both net-new bypasses (base 2 -> head 0), plus one escalation of round-4 Minor 2. All reproduced here against the real hook and confirmed against real commits before fixing. BLOCKER — the matched -m need not be git's message. The capture is a search over the whole command and the double-quoted arm runs first, so it could select a -m that is not the subject at all. git concatenates multiple -m values and takes the FIRST as the subject, so git commit -m 'WIP first' -m "$(cat <<'EOF' … )" commits the subject `WIP first` while the hook validated the heredoc. Same for an unquoted earlier -m, for a heredoc after `--` (a pathspec, not a message), and for one belonging to a later `&& echo`. The mis-selection is pre-existing; resolving it is what made it a bypass. The hook now resolves only when nothing before the matched -m could have been an earlier message, an end-of-options marker, or another command. BLOCKER — cleanup mode is part of the predicate. The resolver skips leading blank lines and strips trailing whitespace because git's DEFAULT cleanup=whitespace does. Under --cleanup=verbatim git does neither, so a 72-char subject plus three trailing spaces is committed at 75 bytes while the hook measured 72 — COMMIT_SUBJECT_TOO_LONG dodged. This one hides from `git log --pretty=%s`, which strips trailing whitespace in its own output; the raw commit object shows 75 vs 72. Any named mode other than whitespace, in either the --cleanup= or -c commit.cleanup= form, now refuses to resolve. MAJOR — recognition trusted any path ending in /cat, so a planted `../evil/cat` printing `WIP injected` had its heredoc body validated while git's real subject was `WIP injected`. Only a bare `cat` or an absolute path is recognised now. A bare `cat` shadowed on PATH is a documented residual and is not fixable from a string — nor a meaningful boundary, since planting an executable already allows running git directly. The changeset and the JSDoc both asserted that a `"` anywhere in the message blocks. Measured false: a `"` on a later body line resolves fine, because the subject completes before the truncation point; only a `"` in the subject line blocks. The changeset also listed <<"EOF" as covered when it measures 2/2. Both rewritten to claim only what is measured, and the residual false positives are now named. Verified: 4 + 2 + 3 new bypass commands now block, with non-vacuity controls proving the default path still resolves; all round-4 maintainer blockers stay closed; the #3802 fix and plain-form parity unchanged across 7 controls; hooks-opt-in 47/47, worktree-safety 402/402, property non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): scope the cleanup-mode guard to the command outside the message The guard scanned the whole $CMD for `--cleanup=` / `commit.cleanup=`, and the heredoc BODY sits verbatim inside $CMD, so any conforming message that merely MENTIONED the token was refused, fell back to the opener line, and was blocked with CONVENTIONAL_COMMITS_VIOLATION. These are ordinary English in this repository, whose own hooks and docs discuss cleanup modes constantly. Reproduced against the real hook: `fix: document commit.cleanup=strip behavior` blocked, the same message without the token allowed (review of #3816, round 5 — BLOCKER). Scoping to $MSG_PREFIX alone, as prescribed, would have reopened the round-4 length-gate bypass the guard exists for: git accepts the flag on EITHER side of -m, and `git commit -m "<heredoc>" --cleanup=verbatim` is caught today only because the scan is command-wide. Measured, not assumed. The scan now covers MSG_PREFIX + MSG_SUFFIX — the whole command minus the one span that is message text — joined with a space so a token cannot be forged across the seam. Swept the guard class rather than the reported instance. The adjacency guard does not share the defect: an in-body `-m "foo"bar` is refused by the already-documented embedded-quote capture limit (any `"` in the subject line truncates the capture), and an in-body `-m ` without quotes resolves and is allowed. Deliberately untouched. Both directions pinned failing-first: the three false-positive rows red against the unscoped guard, and the trailing-flag row reds against prefix-only scoping. Each mutation was echoed back to prove it landed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XogDtuuuGQEfsWaLSaZCLB * fix(#3802): read commit options the way bash hands them to git Round 6 reported the adjacency guard scanning all of $CMD for a glued `-m "..."`, so a glued -m belonging to a chained-after command refused a heredoc that was never truncated. Glue is a property of the ONE character following the matched span, so that character is now the whole window. Separators and redirections are excluded because bash does not concatenate across them: in `-m "msg"&& echo hi` the argument ends at the quote, so there is no truncated capture to defend against. An independent full-PR pass then found three accept-direction defects this PR had introduced in earlier rounds, each measured against a real commit by reading the raw commit object — `git log --pretty=%s` strips the trailing whitespace that makes the length wrong and hides it: --cle=verbatim git accepts any unambiguous prefix of a long option, so the mode was set by a token that is not the literal --cleanup. 75-byte subject recorded, 72 measured. -am 'WIP first' git reads this as -a -m, so the real subject is `WIP first` and the heredoc is only the second message. The scan looked for a standalone -m. --clean""up=verbatim bash removes quotes before git sees the -""m argument, so a spliced spelling is the same option and matched no literal. The two option-name scans now read their window with quote characters removed, which is what bash does to it, and the cleanup class covers git's abbreviations. The adjacency test deliberately keeps the raw text: it asks about a literal character position, not an option name. Narrowing the cleanup window to git's own command segment was tried and reverted. `;`, `&` and `|` end a command only outside quotes, and this is a substring scan, not a parse: an unconditional trim cut the window short on `--author "a&b"`, and a quote-aware trim still cut it on `--author a\&b`. Each hid a real trailing --cleanup=verbatim and accepted a 75-byte subject. The resulting false positive — a --cleanup carried by a chained command refuses the commit — is documented and pinned instead. Refusing a commit git would take is recoverable; accepting an over-long subject is not. Sixteen rows in tests/hooks-opt-in.test.cjs. Seven mutations, including both reverted narrowings, so no dead end can be reintroduced silently. * fix(#3802): close six accept-direction bypasses in the resolve guards Round 7's FIRST-MESSAGE GUARD Major does not reproduce. Measured against the real hook in a complete tree at the reviewed head: the classifier gate runs before any guard, so `git add -A && git commit …` (git->add stops on a non-commit subcommand) and `cd dir && git commit …` (the first executable is not git) exit 0 without a guard being evaluated. The control is the proof — a subject the bare form blocks with CONVENTIONAL_COMMITS_VIOLATION exits 0 in both chained forms, so the hook never validated them and cannot be over-blocking them. The guard is unchanged; scoping this scan to $MSG_PREFIX alone is what reopened the round-4 trailing-flag bypass. The class was real, though, one shape further out: `FOO=bar; git commit …` IS classified and then refused, because assignment detection is prefix-anchored and the tokenizer does not split operators. Pinned as a counterexample and disclosed rather than generalised away; narrowing it means changing isGitSubcommand, the shared git-commit detector every gating hook uses, and it fails closed. Six accept-direction bypasses are fixed. Each let the hook resolve and ALLOW a commit whose real subject the rules refuse; the three that turn on git's recorded subject were confirmed against the RAW COMMIT OBJECT, since `git log --pretty=%s` strips trailing whitespace and hid two of them: --cleanup=whitespace -m <72+spaces> --cleanup=verbatim git kept 75 bytes -mWIP -m <heredoc> git recorded `WIP` --mes=WIP -m <heredoc> git recorded `WIP` -\m WIP -m <heredoc> git recorded `WIP` git commit --amend --no-edit \n echo -m <heredoc> echo's argument read --squash=HEAD -m <heredoc> `squash! …` Causes: one BASH_REMATCH inspected only the FIRST cleanup directive while git applies the last, so multiplicity now refuses rather than guesses at an argument order a substring scan cannot recover; the option scan required a trailing space or `=`, missing attached values and long-option abbreviations; dequoting removed quotes but not the syntactic backslashes bash also removes; the separator scan omitted newline; and --squash/--fixup have git compose the subject, so the supplied message is not the subject at all. Every fix widens refusal, the direction this file documents as recoverable. The multiplicity count first broke the hook outright: the script runs under `set -euo pipefail` and grep exits 1 when it matches nothing, which is the common case, so every ordinary commit died at exit 1 with no verdict. Guarded, and only caught because the probe runs the real hook rather than the scan. Five new rows, all five proven red against the pre-fix hook, each carrying a non-vacuity assertion that the canonical single-`-m` heredoc still resolves. Changeset corrected on three counts: "all fail-closed" was wrong (persistent commit.cleanup fails OPEN, as do the -C/-c/-F/-t message sources), "global options are all walked through" was too broad, and the chained-before claim now states what is measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9 * fix(#3802): stop the separator and glue classes matching a literal backslash Round 8's Major, with two corrections to its account. `;`, `&` and `|` are metacharacters inside `[[ ]]`, so an inline bracket class must escape each one. POSIX bracket expressions have no escape mechanism of their own, so on bash 3.2 -- the system /bin/bash on macOS, already a supported target here per the `declare -A` ban in tests/install.test.cjs -- those backslashes reach the regex engine and add a literal `\` to the class. bash 4+ consumes them, which is why this is invisible on a modern bash. The hazard is specific to bracket expressions: `\(` outside one is made literal correctly on every version, and the subject validator and the `-m` capture classes were checked and are unaffected. The prescribed fix is not taken, because it does not parse. Inline `[;&|]` is a bash SYNTAX ERROR on 3.2 and on 5.3 alike -- the backslashes exist to get the metacharacters past the `[[ ]]` parser, so removing them leaves an unparseable script. Each class is held in a variable and expanded unquoted on the right of `=~` instead, which is a plain regex on both versions. One root cause, consequences in BOTH directions. The reported half is the separator scan over-blocking. The half not reported is the accept direction, and it is the more serious: the glue class is NEGATED, so on bash 3.2 a backslash-glued suffix fell inside the exclusion and the hook RESOLVED a heredoc it should have declined -- measured exit 0 on 3.2 against the unfixed hook, exit 2 everywhere else, with a letter-glued control refused in all four cells. The reported repro is not actually fixed by this, and the changeset says so. A `\`-newline line continuation carries a literal newline, which the round-7 separator guard refuses on every bash, so that shape stays blocked with or without this change. Narrowing the newline guard is not attempted: telling a continuation from a separator by substring scan is the class that was tried twice in earlier rounds and reverted both times, and an escaped backslash sitting immediately before a real newline is indistinguishable from a continuation. Disclosed as a known fail-closed limit instead. Every new row runs under each bash on the machine. Against the unfixed hook both bash 3.2 rows go red while all four bash 5.3 rows stay green -- written the ordinary way these rows would run under PATH bash, pass against the broken hook, and prove nothing. Two non-vacuity controls per interpreter prove the validator is reached rather than passing everything. All 8 rows of the established differential harness are byte-identical before and after on both versions: no regression, no new refusal. * fix(#3802): remove the $ of a dollar-quote from the option-name scans Independent round-8 review, accept direction. The option-name windows are dequoted so they match "the command as bash hands it to git" -- round 6 removed quote characters, round 7 removed syntactic backslashes. Both passes missed that bash has two further quoting forms whose introducer is a `$`: `$'...'` and `$"..."`. Removing the quote characters alone left that `$` stranded INSIDE the option name, so `-$"m"` dequoted to `-$m` and matched no literal, while bash passed a real `-m` to git. Measured on bash 3.2.57 and 5.3.15 against a real repository: the hook allowed git commit --allow-empty -$"m" WIP -m "$(cat <<'EOF' fix: a perfectly ordinary conforming subject EOF )" with exit 0, and `git cat-file -p HEAD` recorded the subject `WIP`. The comparison that establishes this is HEAD-internal, not a differential: the same command spelled `-m WIP` is refused (exit 2). The merge-base refuses EVERY heredoc form, including a perfectly conforming one, so its exit 2 on this input says nothing about whether any guard fired -- it is the absence of the feature, not a working check. The same miss covered `$'m'`, spliced `--message`, `--cleanup`, `--squash` and `--fixup`. An option NAME finished by a command substitution -- `--clean$(printf up)=verbatim` -- is a different problem and gets its own guard: bash runs a program to complete the name, so the argv git receives is not derivable from this string at all, and resolution is refused rather than guessed. The guard is scoped to the NAME: the class is a `-`-leading token whose characters up to the substitution contain no `=`. A substitution supplying a VALUE -- the ordinary `--author="$(git config user.name)"`, spaced or glued, in either window -- is untouched and still resolves, pinned in both directions. It is a SHAPE, not a segmentation of the command line; segmenting was tried twice in earlier rounds and reverted both times, and that reasoning stands. Both new rows fail against the unfixed tree with their own assertions, proven in a complete worktree at the previous head rather than a hook copied out of its tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): recognise a canonical cat, not any absolute path ending in /cat Independent round-8 review, accept direction. Round 4 restricted heredoc-opener recognition to an absolute path, after a relative `./cat` was measured being trusted to echo its stdin. It stopped at "absolute", so any absolute path ENDING in `/cat` was still trusted -- the same claim the round-4 reasoning had rejected one spelling earlier. Measured on bash 3.2.57 and 5.3.15 against a real commit: with an executable at `/.../fake-cat/cat` printing `WIP injected`, the hook validated the conforming heredoc body and allowed the commit (exit 0) while `git cat-file -p HEAD` recorded the subject `WIP injected`. The same command through `./cat` was already refused, which is the control that shows this is the round-4 class one spelling out rather than a new one. Recognition is now the canonical system locations -- bare `cat`, `/bin/cat`, `/usr/bin/cat` -- which is the only identity claim a string can support. `/usr/local/bin` is deliberately excluded: it is user-writable on ordinary machines, which is the plantable case this guard exists for. Anything else falls back to the opener line and the format gate: fail closed, exactly the pre-fix behaviour for the form. The pre-existing residual is unchanged and still documented: a bare `cat` shadowed earlier on PATH is indistinguishable here, and is not a meaningful boundary -- anyone able to plant an executable on PATH can run `git commit` directly. This hook stays an authoring guard, not a security control. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): an option name carrying a shell expansion is unresolvable Independent review, round 9, accept direction. Four more spellings, and a change of strategy that is the actual point of this commit. Rounds 6, 7 and 8 each tried to EMULATE what bash does to an argument before git sees it -- round 6 removed quote characters, round 7 syntactic backslashes, round 8 the `$` that introduces a dollar-quote -- and each round review found another transform that had been missed. Round 9 found four more. All measured on bash 3.2.57 and 5.3.15 against a real repository, each with the plain spelling of the same command as its control (refused, exit 2) and `git cat-file -p HEAD` for the subject git actually recorded: -$'\155' WIP hook 0, real subject `WIP` ANSI-C octal -> m -$'\x6d' WIP hook 0, real subject `WIP` ANSI-C hex -> m -`printf m` WIP hook 0, real subject `WIP` backtick substitution x= … -${x}m WIP hook 0, real subject `WIP` parameter expansion -? WIP hook 0, real subject `WIP` pathname expansion and the same class through the cleanup guard, where git recorded a 75-character subject the length gate had measured as 72: --cle$'\141'nup=verbatim, --clean`printf up`=verbatim, --cle?nup=verbatim The last two settle it. An option name finished by a PARAMETER expansion depends on a variable's value at run time; one finished by a PATHNAME expansion depends on the contents of the working directory. Neither is derivable from the command string at any level of effort, so emulation cannot be completed -- not "has not been completed yet". A fifth patch in that direction would have the same shape as the previous four. The rule is therefore no longer "normalise it and match the literal". It is: an option NAME carrying a shell expansion or quoting construct is UNRESOLVABLE, and unresolvable refuses. One rule covers every spelling above and every spelling nobody has thought of yet, in the fail-closed direction. The dequoting passes are kept rather than replaced: they still normalise the deterministic removals, so the guards RECOGNISE `--clean""up=` and `-\m` as the options they are instead of merely refusing them, which keeps the existing rows meaningful. Scope is unchanged and still pinned in both directions: the class is a `-`-leading token whose characters up to the construct contain no `=`, so a construct supplying a VALUE -- `--author="$(git config user.name)"`, the backtick spelling, `--date="${NOW}"`, a glob character inside an author string, a pathspec after `--` -- still resolves. Nine such forms are asserted to pass beside the seven that must refuse. The class is bracket-only and holds no backslash, per round 8: a POSIX bracket expression has no escape mechanism, and a backslash written inside one becomes a literal member on bash 3.2. The new rows fail against the previous head with their own assertion message, in a complete worktree with the lib built, not a copied hook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * docs(#3802): disclose and pin the two spellings the round-9 class over-blocks A scoped review of the round-9 class asked one question -- does it refuse a conforming heredoc commit that the previous head accepted -- and found two spellings that it does. Both measured on bash 3.2.57 and 5.3.15, previous head 518d97b64 exit 0, current head exit 2: git commit -S$SIGNING_KEY -m <conforming heredoc> git commit -m <conforming heredoc> -- -*.txt Disclosed and pinned rather than narrowed, for two reasons. Narrowing is not available cheaply. Dropping the bare `$` member reopens `-$xm`: with `xm=m` bash hands git a real `-m`, which is the parameter expansion bypass the round-9 commit exists to close. Skipping tokens after `--` means deciding where git's options end from a substring scan, which is the class this file has already reverted twice for opening accept-direction holes -- a `--` inside a quoted value (`--author "a -- b"`) would truncate the window and hide a real trailing directive. And the limits are narrower than they look, because in both cases the spelling a developer actually reaches for still resolves: -S "$KEY" and --gpg-sign="$KEY" resolve '-*.txt', "-*.txt", ':(exclude)-*.txt' resolve The pathspec one is worth stating precisely: a glob only reaches git AS a pathspec when it is quoted, because an unquoted one is expanded by the shell before git is executed. So the refused spelling is not passing a glob to git at all, and the spellings that do are unaffected. Refusing a commit git would take is the recoverable direction; accepting a non-conforming subject is not. That is the trade this file already makes everywhere else, and it is made explicitly here. Nine rows pin the working spellings beside the three that refuse, so a later narrowing cannot silently drop the cases that must keep working. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): join backslash-newline continuations before the resolve guards Round 9's Major, with a correction to its diagnosis. The cited bracket classes at :223 and :260 no longer exist -- round 8 moved both into SEP_CLASS and GLUE_CLASS, and a lone backslash before -m resolves (exit 0) at the reviewed head on both bash 3.2.57 and 5.3.15. What refuses the repro is the NEWLINE a `\`-continuation carries: round 7's separator guard reads any newline in a window as a command boundary, and `git commit \` newline ` -m "$(cat <<'EOF' …` was refused for that reason. Round 8 disclosed it as a fail-closed limit; round 9 calls the idiom common and the limit a Major, and it is fixed here. It was left as a limit because "is this newline a continuation" looked like the segmentation question this file has reverted twice. It is not: bash's rule is local and character-level. A newline preceded by an ODD run of backslashes is a continuation and bash removes both; an EVEN run (`\\` then newline) is a literal backslash followed by a real newline, which IS a separator. Both scan windows are joined that way immediately after they are cut from the command and before any dequote copy is derived, in three bash-3.2-safe parameter expansions: every `\\` pair is parked on \x01, any backslash-newline that remains is a lone one and is removed, then the pairs are restored. Measured on both bashes, both directions: git commit \<nl> -m <heredoc> 2 -> 0 the fix git commit \\<nl> -m <heredoc> 2 -> 2 literal \ + real separator git commit<nl> -m <heredoc> 2 -> 2 bare newline -m <heredoc>\<nl>suffix 2 -> 2 bash glues it; the glue guard sees it glued git commit … \<nl> --allow-empty<nl>echo -m … 2 -> 2 the REAL newline still separates The prescribed `[\;&|]` is not taken: a backslash written inside a bracket expression becomes a literal member on bash 3.2, which is the round-8 defect from the other side. Rows run under each bash on the machine. The fix row fails against the previous head in a complete worktree with the lib built; the four control rows were measured against that same head and were already refused, so they pin existing behaviour rather than the change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
dce40eeb6e |
fix(#4002): rewrite zcode command @-refs to the zcode runtime home (#4188)
* test(#4002): zcode commands must rewrite at-refs to the zcode home * fix(#4002): add the missing zcode case to the runtime rewrite engine * chore(#4002): add ZCode to the bug-report runtime dropdown and drop the changeset * fix(#4002): attribute zcode command and skill ripples to the rewrite engine * fix(#4002): attribute zcode nested-skill ripples to the rewrite engine * chore(#4002): backfill changeset pr number * fix: bump qs past GHSA-x5fp-wj9c-mxmx (transitive, advisory reddened next) --------- Co-authored-by: sim <sim@local> |
||
|
|
e242d85c92 |
fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh (#4116)
* test(#4109): add zsh regression coverage for gate-check and invoke_reviewers dispatch Extends tests/review-plan-coverage-manifest.test.cjs to cover the two DISPATCH_SLUGS consumption sites the existing #3301 coverage-check tests don't reach: the write_reviews gate-check block (ALL_LANES_SKIPPED / TOTAL_LANE_FAILURE counting) and invoke_reviewers' dispatch + join loops. Both extract the real shipped bash from review.md and run it under bash and zsh, same convention as the existing coverage-check rows. Expected RED under zsh pre-fix, GREEN post-fix. * fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh Bash word-splits an unquoted scalar on IFS by default; zsh does not (no setopt SH_WORD_SPLIT anywhere in these files), so a scalar accumulated from multiple space-separated tokens and then consumed via bare `for x in $VAR` collapses onto one bogus iteration under zsh whenever it holds 2+ tokens. Fixes all four review.md sites reported by #4108/#4109 (invoke_reviewers dispatch + join loops, the gate-check block, and the #3301 coverage-check block), plus the identical pattern found by a repo-wide sweep in complete-milestone.md, code-review.md, pr-branch.md, sync-skills.md, and execute-phase/steps/per-plan-worktree-gate.md. Each site is rewrapped in unquoted command substitution (`$(printf '%s' "$VAR")`), which re-splits identically under both shells regardless of SH_WORD_SPLIT — the same mechanism that already made the accumulator-building loops in these files shell-safe. * test(#4109): warn loudly when the zsh probe fails instead of skipping silently detectShells() drops the zsh test lane whenever a live zsh probe fails (e.g. on a CI runner without zsh installed), and a dropped lane reads identically to a passing one in the suite's own output — exactly the blind spot that let #4109's bug class ship undetected. Both copies of this helper (this file and its byte-identical duplicate in review-build-prompt-optional-sections.test.cjs) now emit a greppable console.warn when the probe fails, so a run without zsh reads as "zsh coverage unknown" rather than "all lanes green". * ci(#4109): gate workflows/*.md's embedded bash blocks in lint:ci Adds scripts/lint-workflow-shellcheck.cjs, wired into npm run lint:ci, so this bug class can't land undetected a third time. Two independent checks run over every ```bash block in gsd-core/workflows/**/*.md: - ShellCheck (new `shellcheck` devDependency, downloads and caches the real koalaman/shellcheck binary) catches the general unquoted-expansion/ word-splitting family in argument position. 212 pre-existing findings across the tree are absorbed into scripts/lint-workflow-shellcheck-baseline.json (matched on {file, code, message}, not line number, so unrelated edits elsewhere in a file can't spuriously un-baseline anything) — fixing all of them is out of scope for this issue; only NEW findings fail the build. - A custom structural check specifically for #4109's own shape: ShellCheck does not flag a bare `for x in $VAR` word-list — it treats that as an intentional idiom under any ruleset (confirmed empirically). This check does, and gates the build on any occurrence, with zero tolerance (no baseline) since every known site was already swept and fixed on this branch. * test(#4109): update pr-branch cherry-pick loop test anchor for the zsh fix extractPickLoop()'s PICK_LOOP_MARKER located the create_pr_branch cherry-pick loop by the literal substring "for HASH in $INCLUDED_COMMITS", which no longer appears verbatim after this issue's fix rewrapped that loop in unquoted command substitution. Updates the anchor (and its error message, now derived from the same constant instead of duplicating stale text) to the new literal form. No behavior change to the extraction logic itself. Emitted-Drift-Ack-Growth: review.md — #4109's fix adds explanatory comments at 4 sites; net code is functionally equivalent, comment expansion grows the file Emitted-Drift-Ack-Growth: complete-milestone.md — #4109's fix adds explanatory comments at 3 sites documenting the zsh word-splitting divergence Emitted-Drift-Ack-Growth: code-review.md — #4109's fix adds an explanatory comment documenting the zsh word-splitting divergence Emitted-Drift-Ack-Growth: pr-branch.md — #4109's fix adds explanatory comments at 3 sites (TRANSIENT_DIRS, INCLUDED_COMMITS, FILTER_PATHS) Emitted-Drift-Ack-Growth: sync-skills.md — #4109's fix adds explanatory comments at 2 sites (CREATE_LIST/UPDATE_LIST, REMOVE_LIST) * test(#4109): cover lint-workflow-shellcheck parsers + add timeout Adds tests/lint-workflow-shellcheck.test.cjs covering the hand-rolled parser/logic functions exported by scripts/lint-workflow-shellcheck.cjs (stripCommandSubstitutions, stripShellComments, substitutePlaceholders, extractForLoops, findBareForLoopSplits, findingKey, partitionAgainstBaseline) that shipped with zero coverage — CLAUDE.md requires at least one fast-check property test for parsers, included here (bare/braced forms always flagged naming the variable; quoted/substituted/literal forms never are). Also bounds runShellcheck's subprocess with a 60s timeout: the `shellcheck` npm package's own API has no timeout option and internally blocks on a synchronous spawnSync, so this reimplements the binary resolve/download step via the package's own exported config/download and calls spawnSync directly with a native timeout. And guards the CLI entry point with `if (require.main === module)`, matching this repo's sibling dual-purpose lint scripts — without it, requiring the module for its exported functions (as the new test file does) also triggered a live ShellCheck run as a side effect. * docs(#4109): add changeset for the zsh word-splitting fix * chore(#4109): backfill changeset PR number (#4116) --------- Co-authored-by: sim <sim@local> |
||
|
|
686c05a740 |
fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128 (#4059)
* fix(#4058): raise MAX_ACK_TRAILERS from 64 to 128 64 had no principled derivation (no git trailer limit, no CI resource bound) and a wide-touching maintenance PR can legitimately accumulate more than 64 distinct emitted-drift acknowledgments, tripping the cap and failing Required tests even though nothing is actually wrong. The cap's underlying purpose is unchanged: parseAckTrailers() still throws rather than truncates on overflow, and still forces pruning of stale acknowledgments at some ceiling. Only the ceiling moves. * chore(#4058): backfill changeset PR number (4059) --------- Co-authored-by: sim <sim@local> |
||
|
|
90fff40f5a |
fix(#3900): overlay walker skips non-regular files — a repo-root socket no longer fails 32 tests (#4062)
* test(#3900): a unix socket under the repo root must not break the overlay (failing first) * fix(#3900): the overlay walker skips non-regular files (sockets, FIFOs, device nodes) The walker classified every not-a-directory entry as a file, so a unix socket anywhere under the repo root (an MCP server, a language-server daemon, a watcher) made copyFileSync throw ENXIO — the reporter measured 32 local test failures across 6 files, all in before() hooks, none about the code under test; a FIFO would block until a writer appeared. The hard-link path does not save it: the overlay sits under tmpfs, so cross-device EXDEV always falls through to the copy. One guard at the classification site (srcStat.isFile() before the copy/link branches) — sockets, FIFOs, and device nodes are not repository content; CI never saw this because a fresh clone has no working-tree daemons, which is exactly why it ate local debugging time. Also adds opts.root (test-only root override, default REPO_ROOT) so the regression tests build the overlay from a tiny synthetic root in milliseconds instead of paying a full repo walk. * chore(#3900): changeset fragment (pr number backfilled after PR creation) * fix(#3900): review fold-ins — sibling cold-fixture guard, changeset format, root doc buildColdInstallTree enumerated the live repo's hooks/ with a NAME-only filter — the same class #3900 documents: a daemon socket or FIFO inside hooks/ would reach cpSync and throw ENXIO (or block). Dirent type guard added. The changeset gains the mandated bold user-visible lead, and opts.root is documented at its definition (mirroring cold-runtime-lib- fixture's repoRoot convention). * chore(#3900): backfill changeset PR number (4062) --------- Co-authored-by: sim <sim@local> |
||
|
|
ab69b9ce56 |
enhance(#3987): guard slug re-derivation and the swallowed-precondition shape — §8.5 was guardable after all (#3999)
* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from |
||
|
|
12f9d1d9a0 |
enhance(#3913): docs, and the guards come down (#3994)
ADR-3889 terminal phase. Generated docs/reference/exit-codes.md from the exit-code declaration with a --check drift arm; deleted the inert soft-error-exit-zero oracle; promoted untyped-success from SMELL to VIOLATION so it can fail a build; pruned all 5 smell-baseline entries. Fixed inline: two mis-scoped oracles (routing-validity, value-hygiene), a second source behind the band table, unescaped declaration strings reaching Markdown, and a pre-existing Windows 8.3 short-name path-comparison defect. Guard ledger corrected from a claimed net -4 to a measured net -1. Closes #3913 |
||
|
|
fa41bfec5c |
enhance(#3942): the emitted-drift ack is PR-lifetime data — move it to a commit trailer (#3954)
* test(#3942): failing-first suite for the emitted-drift ack commit trailer Binds 37 input classes from the phase test matrix to the behavior ADR-3942 specifies, before any of it exists. Stubs return benign empty values rather than throwing, deliberately: several rows assert that something DOES throw (cap overflow, uncomputable commit range), and a throwing stub would turn those green for the wrong reason and destroy the red. The two rows that carry the design's load: - merge-base semantics. The range is $(git merge-base base HEAD)..HEAD, not base..HEAD, because changedPaths comes from `git diff base...HEAD` (three dot). Two-dot would let the ack set and the change set disagree about which commits are this PR's. The fixture forks a topic branch, puts a trailer on each side, and asserts only the topic-side trailer is in range. - fail-closed on an uncomputable range. With fragments a depth-1 checkout passes VACUOUSLY, every fragment reading as brand-new. With trailers the range cannot be computed at all, and returning an empty set would silently disarm the gate, so it must throw. The fixture builds a genuine shallow clone rather than simulating one. Also covers the self-inflicted case: this change's own documentation quotes the trailer syntax, so an example landing at the end of a commit message would arm a live acknowledgment keyed on the literal placeholder text. Keys carrying angle brackets or whitespace are rejected. Authored per the phase artifacts 40-design.md and 50-test-matrix.md. Not yet run on the remote runner — this commit exists to be tested. Refs #3942 * chore(#3942): move the emitted-drift ack to a commit trailer Implements ADR-3942, superseding ADR-2719 section 3 and its #2789 amendment. Sections 1, 2 and 4-7 are retained: the conservation law is unchanged, only the storage of its escape hatch moved off the working tree. An acknowledgment explains one PR's ripple, and the moment that PR merges the ripple is in the base, so it can never clear anything again. It was stored in permanent shared state anyway, and every consequence of that mismatch had to be built and then maintained. The chain is #2789 -> #2914 -> #3078 -> #3842 -> #3823 -> #3875, each fix generating the next defect, ending in a scheduled sweeper whose own first PR could not merge itself. Added parseAckTrailers + renderAckTrailer (pure) and readAckTrailers (IO shell), reading Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: trailers over the merge-base range. tests/emitted-ack-trailer.test.cjs, 37 cases, written failing-first and confirmed red before any of this existed. Changed diffEmitted takes two structurally distinct key-space maps instead of one shared paths map. That closes a latent defect: the spaces were separated by convention only, so a growth key satisfied a hash lookup by naming coincidence. staleAcks now reports which space a key was declared in. REMEDIATION teaches the trailer, per space, with its example rendered through renderAckTrailer so the taught grammar cannot drift from what the parser accepts. Removed the sweep workflow, the guard-no-ack-on-next job, the standalone linter and its lint:ci entry, the fragment directory and its three spent fragments, the legacy single-file union, and the baseAck/spentAcks mechanism -- spentness is now structural, not computed. Two range properties carry the design and are pinned by tests rather than asserted: the range is merge-base scoped, matching git diff base...HEAD, so an already-merged trailer is out of range by construction; and an uncomputable range throws instead of reading as zero acknowledgments, which is the inverse of the fragment guard's vacuous pass. Three deliberate observable changes, each disclosed in the changeset: the unread runtime field is gone, the legacy file is no longer read, and cross-space excusal no longer works. Ten open PRs carry fragments and will meet a modify/delete conflict. Measured before landing and accepted deliberately; the one-line migration is in the PR body. Verified: lint:ci exit 0. Remote runner to follow on this exact sha. Refs #3942 * fix(#3942): silent trailer collapse, lost coverage, and an unbounded cap Six findings from the orthogonal review round, all fixed in place. BLOCKER -- two trailers of the same name on one commit collapsed silently. readAckTrailers built `separator=1d` where git needs `separator=%x1d`: the `separator=` value inside a %(trailers:...) placeholder is itself a pretty-format string, so the bare hex was emitted as two literal characters and the split on \x1d never matched. Two same-name trailers therefore joined into one value with errors empty -- the first reason absorbing the second entry's key. Silent truncation, the exact class MAX_ACK_TRAILERS throws to prevent. Confirmed with od -c against real git output before and after. The failing-first matrix did not catch it because its "both spaces coexist" row uses Hash plus Growth -- different trailer NAMES -- so the value separator was never exercised. Two regression tests now cover same-name trailers directly. Coverage recovered: normalizeAckReason and INVISIBLE stayed on the live path via parseAckTrailers but lost every test when the old suite was pruned. Back under test against the current surface -- all six invisible codepoints individually, whitespace collapse, trim, CRLF, and two seeded fast-check properties. Dropping any single codepoint now fails. MAX_ACK_TRAILERS counted raw trailers before de-duplication, so one trailer carried forward across rebased commits counted once per commit and could throw on a legitimate branch. Now counts distinct entries; 100 identical repeats dedupe to one. diffEmitted validated baseline, current and changedPaths but not the new ackHash/ackGrowth, so a bad shape raised an unhandled TypeError instead of an error verdict -- the same defect shape this file documents for #2778. Docs: CONTRIBUTING and TESTING-SUITES were rewritten only in their first sections; the later passages still taught fragments, git rm and the deleted guard, contradicting the new text directly above them. Finished. Also extends lint-removed-but-needed to exempt docs/adr and docs/research. That gate fails on any docs mention of a file deleted in the same diff, which makes it impossible to document a deletion in the PR performing it -- an ADR's whole job is naming what it retired. Exemption is narrow and comes with a test proving the gate still fires for a live consumer elsewhere under docs/. A guard that cannot fail is worse than no guard. Maintainer-approved. CONTEXT.md names the retired machinery by role rather than by filename: its generated projection lands in docs/, which that gate does scan. Adds docs/how-to/acknowledge-emitted-drift.md. The required docs set is Reference and Explanation, so the task quadrant can be empty with every gate green -- and this change has a real multi-step journey, including the fragment migration ten open PRs now need. lint:ci exit 0. Refs #3942 * docs(#3942): correct the duplicate-trailer rule in CONTRIBUTING Both axes of the code review independently flagged the same passage, without seeing each other's output. It claimed two declarations of the same key are always "a hard, loudly-reported error, not a silent last-wins". That is only half true, and the missing half is the one contributors hit: identical declarations -- same key, same reason -- dedupe silently, because a trailer legitimately survives a rebase and reappears on every rebased commit. Failing there would red a branch for doing nothing wrong, which is exactly why the dedup exists. Only a same-key/different-reason pair errors, and that one is a genuine ambiguity about which explanation holds. As written, the paragraph told a contributor that a rebase-carried trailer breaks the gate -- the opposite of the behavior. CONTEXT.md's parallel entry already stated it correctly; this brings CONTRIBUTING into line. Doc-only, root-level markdown. Refs #3942 * chore(#3942): backfill changeset PR number to 3954 --------- Co-authored-by: sim <sim@local> |
||
|
|
c5f2b94b27 |
enhance(#3907): gates report no-input instead of a verdict they never reached (#3932)
* feat(#3907): gates report no-input instead of asserting a verdict they never reached The three stdin-reading gates bound 2 to a stdin read error only, with no arm for stdin closed at zero bytes - so empty input flowed into the detector, found nothing, and exited 1, which each module's own comment defines as a negative verdict. An unset PHASE_SECTION made the UI gate assert the phase has no UI. Empty and whitespace-only input now exit NO_INPUT, and a read error exits UNAVAILABLE rather than a locally-invented 2, both resolved through the registry and delivered by terminateNow. The exit code was only half of it: under --json the same input emitted {detected:false}, byte-identical to the fabricated payload #3909 exists to fix, and the blocking coverage gate reads that payload. Empty input now emits the in-tree {skipped:true,reason} form with no detected key at all. teams-status is excluded: it never reads stdin and has no invented 2, so the four-module framing in the issue and ADR is wrong. The dead root bin/lib/ui-safety-gate.cjs is deleted - no installer reference, no workflow invocation, and the live fallback chains are for other modules. Its removal restores the unit tests to the module that actually ships; they had been asserting the stale copy's two-field shape, which is why it drifted unnoticed. * fix(#3907): drive gate tests through the process seam, and make removed-but-needed basename-precise CONTRIBUTING requires every subprocess go through tests/helpers/process-seam.cjs; two of the three gate suites hand-rolled spawnSync while the third, added in the same change, used runNode correctly for the identical injection case. Converted the blocks this change added, leaving pre-existing ones alone. Deleting one of two files sharing a basename made lint-removed-but-needed report 14 references that were all to the surviving canonical module - the false-positive class its own docstring names. It now matches on the deleted file's full path when a surviving file shares its basename, which is more precise rather than weaker: a genuine full-path reference still fails, and behaviour is unchanged when no basename collides. It immediately caught a docstring on this branch that spelled the deleted path. * test(#3907): update the one existing assertion that pinned the old empty-stdin verdict A pre-existing test asserted exit 1 on empty stdin - the defect this phase removes - and was missed because the change added new blocks without auditing existing ones pinning the old contract. Audited the rest: the other three status-1 assertions in that file all feed real input and are the genuine-negative controls that must keep returning 1, so exactly one was stale. The retired 2 is gone from the describe's contract comment too. * chore(#3907): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
39673ae9ff |
fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config Regression tests (RED first): --skills-root and gsd-tools query surfaces, install-plan dest dirs, and converter skills-path rewrite. * fix(#3738): antigravity global skills/agents install to ~/.gemini/config Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents}; the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the ADR-1239 skills/agents 'home' override on the antigravity global layout — the same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references in converted global content to ~/.gemini/config/skills/. configHome, settings, probe/migration semantics, and the local .agents layout are unchanged. * fix(#3738): retire deprecated configHome artifacts via installer migration 010 Next install converges an existing antigravity install: manifest-managed skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does not scan) are removed — modified files backed up first, unmanifested and non-gsd entries preserved — and now-empty containers retired. Global scope only; the local .agents surface is live. Docs + inventory updated. * fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline - bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/ rewrite as src (ADR-1508 dual copy must stay in sync). - Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so antigravity's emitted skills/agents stay differential-visible at their new install root; install-tree fixture regen confirms an unchanged key set. - skills-from-commands rule declares the antigravity converter as a runtime-scoped transform; one ack fragment covers the identity-classed workflow whose antigravity copy embeds the old skills path. - Migration 010 checksum baseline + home-override set doc updated; existing tests updated to the #3738 contract (global dest, golden parity via layout dest, integration expectations). * fix(#3738): tolerate an absent extra emit root on baseline-side measurement The base tree's installer predates the home override, so <HOME>/.gemini/config does not exist there; walk() threw ENOENT and the in-job baseline build failed. An absent extra root is the legitimate pre-override shape — skip it. * fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment - writeManifest resolves the agents-kind home override (_kindDestDirSafe), so the manifest records agents at their actual install root and drift detection keeps working (isolated review finding 1, major). - Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing slash) divert to ~/.gemini/config/skills instead of falling through to the retired configHome path (finding 2). - real-home-guard comment updated: antigravity's global agents kind is the first agents-kind home override (finding 3, doc-only). - Regression tests for both behavioral findings. * chore(#3738): changeset fragment (pr number backfilled after PR creation) * chore(#3738): backfill changeset PR number (3921) * fix(#3738): sandbox HOME in tests that install antigravity global artifacts antigravity is the first home-override runtime in the golden-parity and skills-wrapper suites (codex is not in their runtime lists), so those tests never needed HOME sandboxing — the real-home guard now (correctly) refuses their un-sandboxed global installs on CI, where HOME is the passwd home. * fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans K3's two back-to-back sandboxHome calls leave HOME pointing at the first sandbox once the after-hooks restore (each call saves the env as it found it, so the second saves the first's sandbox as 'original'). On the windows matrix that leaked gsd-k3-qwen-* home into the L2 property, whose antigravity/global run then (correctly) refused via the #3712 real-home guard — antigravity is the runtime that made L2's plan escape into os.homedir(). K3 now manages the env with a single restore; L2 sandboxes HOME per run, mirroring L1. * fix(#3738): L2 property's HOME sandbox must exist on disk The #3712 guard's sandbox exemption fails closed when identify(effectiveHome) is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir under the real home) the antigravity/global run refused even with HOME sandboxed. Create the per-run sandbox dir and clean it up. --------- Co-authored-by: sim <sim@local> |
||
|
|
6b7df61938 |
enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises ADR-3473 §8.1 carries a blocking open question with a forcing function: it must be answered before any implementation PR for the rule opens. Answered here as (a), a string-coercing adapter, with the measurement that settles it. The sequencing note bet that §8.8's schema would make (b) tractable. Measured against merged reality it does not: only 33 of extractFrontmatter's 78 non-test call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists survive real types, leaving ~31 lines across 3 call sites as the actual prize. Also corrects three claims verified false while answering it. §8.1's justifying sentence names #3349 and #3360 as defects a real parser would fix; both are already fixed on next, confirmed by executing the compiled parser rather than reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an expected casualty of this rule, but it guards shell grep idioms in workflow bash fences and never touches our parser. The same roster calls lint-vendored-deps.cjs reusable as-is; it is hardcoded to re2js throughout. The last two were caught by applying the rule this amendment records -- a factual claim in this ADR is a hypothesis until the implementing phase executes it -- on its first use. Refs #3881 * docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable An adversarial pass on the Phase 4 design established by execution that extractFrontmatter is not a YAML parser but a line-oriented scanner whose output is a function of raw source text. Four spellings of the same value collapse to one js-yaml tree but produce four distinct legacy strings, one of them mangled. No adapter over a tree can choose among outputs the tree does not distinguish, so fork (a) -- keep a string-coercing adapter so the existing contract holds -- cannot be built. For any document with a non-scalar value, (a) collapses into (b); about 26 percent of frontmatter-carrying documents have one. Also records three design defects and one new attack surface, all confirmed by execution: catching a parse failure and returning {} would delete the frontmatter block on the next write at eight call sites that conflate empty with unparseable; an empty value yields null where legacy yields {}, and reconstructFrontmatter omits null-valued keys, so the shipped state template's empty progress key would vanish; the #1882 truncation probe is parseYamlRegion itself rather than a pre-parse heuristic, so it cannot both stay unchanged and survive that deletion; and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB. The rule is not deferred. The measurement is the deliverable and the re-scoping is recorded as an open question with a forcing function, per section 8's own rule. Refs #3881 * test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed. Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states. B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today. B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today. B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today. Refs #3881 * chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496). gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own. src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare. scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored. docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds. Refs #3881 * feat(#3881): parse .planning frontmatter with the vendored js-yaml ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and parseQuotedScalar are deleted (not patched); parsing now goes through the vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA, json: true }. Everything js-yaml does not do is layered on top, in one place, carrying the seven design-doc consequences: 1. Empty value: a null js-yaml value is coerced to {} (matching legacy's own empty-value contract) so reconstructFrontmatter — which omits null-valued keys — still round-trips a bare `key:` line instead of deleting it. Verified live: progress: with no value survives parse -> reconstruct -> re-parse. 2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS channel, is carried on the {} returned for malformed/refused YAML. Invisible to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that never inspect it are unaffected; wiring the 8 hasFrontmatter sites to consult it is a separate change, not done here. 3. Non-scalar object-list items (the four spellings of `- test: a b` that js-yaml collapses into one tree shape) are rendered as a canonical `key: value[, key2: value2]` string per item, keeping the existing array-of-strings value SHAPE. A full corpus differential over all 1702 tracked markdown files found 11 residual divergences from the legacy parser (enumerated in the PR/report), most of them the parser now being MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key defect, a dropped Unicode key). 4. The #1882 truncation probe still runs the one real parser, but derives its key count from js-yaml's own thrown error and mark.line when the whole region doesn't parse cleanly (the dominant real truncation shape: fence opened, well-formed keys, no closing fence). Verified against both the clean-parse and the exception-fallback path. 5. The #3257 comment channel now attributes each pending column-0 comment against js-yaml's own parsed top-level key list (matched by literal key text, in document order) instead of the legacy ASCII-only key regex, so a comment above a Unicode key attaches correctly. 6. Anchors, aliases and merge keys are refused outright (a raw-text pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences today: zero. A 7-line billion-laughs fixture is verified refused rather than expanded. 7. A literal U+0000 is swapped for a private-use sentinel before the parse and restored in every resulting string afterward, since js-yaml rejects NUL unconditionally under every schema. escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump() (forced double-quoted style), with control-char hex escapes lowercased to keep serialized output byte-stable (#1779 emitted lowercase); it keeps its exported name and signature for its two other call sites (commands.cts, runtime-artifact-conversion.cts), which need no change. frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments, regenerateFrontmatterKey's guard, noOpObjectListSetError and parseMustHavesBlock are all unchanged — retiring them is fork (b) and is not this phase. Refs #3881 * fix(#3881): quote template placeholders and preserve unparseable frontmatter SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax under the vendored js-yaml parser, not the literal placeholder text intended. Quote them so they parse as strings. Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the 8 call sites in state.cts/state-transition.cts that compute hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and reassemble the document without a frontmatter block when false. That check conflated 'no frontmatter' with 'unparseable frontmatter' (both parse to {}), so a document with a merge-conflict marker or refused alias in its frontmatter had that block silently dropped on write. Each site now preserves the exact raw bytes stripFrontmatter removed when the marker is set, leaving the genuinely-empty case unchanged. Refs #3881 * test(#3881): consequence and boundary coverage for the js-yaml migration Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix. Refs #3881 * docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to Refs #3881 * docs(#3881): correct the frontmatter glossary entry Two errors in the entry as first written: it named parseYamlRegion as part of the read path when that function is deleted, and it recorded the eight hasFrontmatter call sites as unwired follow-on work when they were wired in e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the frontmatter block independently, so the marker binds at the transform layer. Refs #3881 * docs(#3881): record the semantic-migration decision and the counted guard ledger The maintainer chose the full semantic migration over splitting the rule into its own epic or patching the scanner, so section 8.1 is answered as "the fork was ill-posed and the migration is semantic" rather than as (a) or (b). Also replaces the pre-implementation guess that this phase would shrink the guard surface with the counted result: excluding vendored third-party lines the hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines despite four functions being deleted, because the compatibility layer over js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is therefore not delivered as written; what improved is the kind of code maintained, not the amount. Decision 6 requires recording that rather than netting it away. Refs #3881 * chore(#3881): changeset for the vendored YAML parser migration Refs #3881 * test(#3881): golden parity, round-trip property and packaging coverage Refs #3881 * fix(#3881): refuse anchors structurally and fold in review findings ADR-3473 §8.1 review findings, addressed inline: Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping ({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor mechanics while never matching that line shape, so the exact expansion the guard exists to stop went straight through unrefused (a 303-byte quoted-key bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener` callback, which reports `state.anchor` for every event belonging to an anchored node in every spelling, and throws from inside the callback to abort before any expansion (~1-2ms vs full expand-then-discard). A merge key with an alias is still refused (merge always requires a previously anchored node, so the alias itself trips the listener); a bare merge key with NO alias is no longer separately refused, documented as intentional: FAILSAFE_SCHEMA never resolves `!!merge`, so it carries no expansion risk. Table-driven tests added for all four bypass spellings + merge key, plus a quoted-key-spelled billion-laughs fixture registered in the adversarial matrix and README. Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/ aliases were "simply UNREACHABLE from typed code" through the twin. Corrected to state the truth: anchor/alias resolution is document-level `load` mechanics reachable through exactly the declared surface, and refusal is enforced at RUNTIME (Finding 1's listener), not by the type surface. Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree back to NUL, including one the document author legitimately wrote, silently corrupting it. Now refuses outright whenever the raw region already contains U+E000 (consistent with the existing anchor/merge-key refusal path), making the substitution provably injective. Tests added for a real NUL alone (preserved), a pre-existing U+E000 alone (refused, not corrupted), and both together (refused, not merged into one byte). Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead for a hand-authored row (only read inside the upstream-verbatim branch) — exactly how Finding 2's stale docblock drifted unnoticed. Added checkHandAuthoredTwin: every value-level export the twin DECLARES must be an actual own property of the vendored runtime module at require-time. Tests added, including a sensor that a declared-but-nonexistent export IS caught. Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/ unparseableFm/reassemble preamble, copy-pasted at 7 sites in state-transition.cts plus a sixth hand-inlined copy in state.cts's cmdStateCompletePhase, is now one exported helper (beginFrontmatterReassembly) every site routes through, including the hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore) keep a literal `body = stripFrontmatter(content)` assignment alongside the helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward scan (which does not chase aliases) still sees the strip; stripFrontmatter is pure/idempotent so the extra call changes nothing observable. Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a separate change" claim (the 8 call sites are wired on this branch) and the changeset's backlink from (#3473) to (#3881). Finding 7: fixed the lint:ci failures blocking the gate — an @typescript-eslint/only-throw-error violation from throwing a bare Symbol as the anchor-detected signal (now a real Error subclass), unused-var warnings left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by two migration-specific test files (allowlisted with justification), and the lint-state-write-path-drift false positive from Finding 5's helper (fixed above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already carried an explicit timeout; no change was needed there. Golden fixture: added a golden entry for the new anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line scanner would also produce, since it independently dropped every quoted top-level key). No other corpus document diverges: real .planning/ documents carry zero anchors/aliases/merge keys/U+E000 today. Refs #3881 * fix(#3881): fold in second-round review findings Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration ASCII-only key regex for the Unicode fixture; updated to require the 相 key's value now that js-yaml has no such restriction. Audited the rest of the file for other pre-migration pins (block scalars, quoted keys, flattened values, empty values, duplicate keys, unclosed blocks, null bytes) by execution against real fixtures; found none regressed. Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts so no function still answers to the deleted hand-rolled scanner's name (ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three external call sites (src/commands.cts, src/runtime-artifact-conversion.cts) updated in the same change — a mechanical rename, not an ADR-amendment matter. Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found by execution. A top-level key named constructor/__proto__/toString/ valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read resolving an inherited Object.prototype member); a key literally named __proto__ was silently DROPPED entirely (bracket assignment on an ordinary {} invoked the inherited __proto__ setter instead of creating a data property). Fixed by building every parsed Frontmatter object with Object.create(null), and replacing an `in` check with hasOwnProperty.call in propagateCommentChannel. Added round-trip tests for all five hostile keys, each with its own leading comment. Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed full byte-stability across the migration. Verified by execution: BEL/NUL/ NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw- literal forms. Proved round-trip equivalence (each escape re-parses to the exact source codepoint) and corrected the docstring. Found and fixed a related real defect while verifying: a lone UTF-16 surrogate was emitted BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely unparseable YAML that silently collapsed to {} on re-read — extended scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped path. Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real truncation shapes (unquoted colon, open flow collection, mis-indented sibling key, refused anchor). Root cause: the mark-based prefix recovery excluded the very line whose key needed counting, and a mark-less refusal never entered the recovery branch at all. Fixed by taking the max of two lower bounds: the longest parser-verified line-prefix, and a raw-text count of key-shaped lines (reusing the same key-shape pattern this file already uses for isFrontmatterShaped). Extended test-matrix row A5 table-driven over all 4 regressed shapes. Finding 6: the design doc's claim that no test owned the #3594 adversarial fixture corpus was false — consolidation epic #1969 had already folded it into tests/frontmatter.test.cjs. An earlier commit on this branch re-created a standalone duplicate under that false premise; folded its genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures, block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the ADR's §8.1 note, including the roadmap-sibling claim (no such file exists). Finding 7: the golden serializer sorted object keys, making it structurally blind to the key-order-parity invariant ADR-3473 §8.1 actually claims. Made it order-preserving and regenerated the golden fixture from a standalone compile of the legacy (pre-#3881) parser at ddde001af; the current parser matches it with zero undocumented divergences, confirming key-order parity genuinely holds. Extended row A2 table-driven across 6 of the remaining 7 transitionCore kinds (all pass) plus documented, by execution, a newly-discovered 8th-site regression: state.cts's cmdStateCompletePhase calls the same preservation helper but its result is clobbered by a later unconditional resync — filed as a distinct finding rather than fixed here (touches syncAndPreserveStateMd, outside this change's verified scope). Refs #3881 * fix(#3881): preserve unparseable frontmatter through the CLI write path Characterization (executed, before/after shown): case (b), not (a). The frontmatter FENCE survives — `state complete-phase` on a conflict-marked STATE.md returns success and a well-formed, freshly-derived frontmatter block, not a document with no frontmatter at all. But the block's actual content (the merge-conflict markers, and with them any signal to a human that the document was in conflict) is silently discarded and replaced. Root cause was two clobber sites, not one: 1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved `transformedContent` from readModifyWriteStateMd, finds {} + the FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh frontmatter block from the body anyway. 2. Even after (1) is fixed, applyPostSyncPreservation's own postFm/applyStatePreservation/authoritativeFm-reassertion machinery re-extracts frontmatter from syncedContent, restores curated fields from the pre-write snapshot, and reconstructs a NEW block again — confirmed live via `state begin-phase`, which still lost the markers after fixing (1) alone. Both are now guarded by the same predicate (isUnparseableFrontmatter, checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not parse and the caller is not on ADR-3408 §8.3's closed "body wins" list, both functions return their input content unchanged rather than re-deriving over it. The closed list (cmdStateSync #905, /gsd-health --repair's REGENERATE_STATE, both routed only through writeStateMd, which never reaches applyPostSyncPreservation and passes sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is untouched — neither widened nor narrowed; verified by execution that `state sync` still overwrites the conflict-marked block exactly as before. Other verbs sharing the same readModifyWriteStateMd path were checked and were equally affected before this fix: state update, query state.patch, and state begin-phase all lost the conflict markers (RED, shown by execution), and all three now preserve them (GREEN). Covered table-driven in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe block, which drives the real CLI verbs via runGsdTools — not just the pure transitionCore layer the earlier A2 rows exercised — plus a control asserting state sync's body-wins contract is unchanged. Refs #3881 * fix(#3881): restore the parse surface's prototype and fix remote-runner failures Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched. Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write. Refs #3881 * chore(#3881): backfill changeset PR number Refs #3881 * test(#3881): make golden parity resistant to unrelated tree churn A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus. Refs #3881 * test(#3881): make the parser golden hermetic instead of tree-keyed This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md, docs/*.md) the prior golden pinned by tracked path. Any PR editing one of those files' frontmatter for reasons unrelated to the parser (an argument-hint addition, an allowed-tools tweak) turned the suite red, and the reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to catch a real regression. Excluding .changeset/** was not enough; the design itself was wrong: a regression fixture must not be keyed to mutable repo paths, and a single 376-entry JSON every such PR touches is also a guaranteed merge-conflict surface. Rebuilt the fixture to carry its own documents: each of 51 entries stores a stable id, literal documentText (shrunk from a real ddde001af-era corpus document), and an expectedParse captured independently from the pre-migration legacy parser (git show ddde001af:src/frontmatter.cts, compiled standalone against its byte-identical sibling modules). The test reads no tracked path, shells out to no git command, and enumerates no tree -- a PR editing commands/gsd/help.md cannot affect it. Every entry's reconstruction was verified at capture time to reproduce both the current and legacy parser's output on the original document; 0 of 51 candidates were dropped by that check (1, the deliberately-unterminated unclosed-block.md adversarial fixture, has no closing fence to truncate at and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now diverges:true entries) and the D2 order-preserving structural serializer that keeps the comparison from passing vacuously; dropped the tree-enumeration helpers, the coverage floor, the post-capture-skip logic, and the vanished-file check -- all artifacts of the path-keyed design. Refs #3881 * fix(#3881): resolve vendored-deps paths independently of cwd shape Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on windows-latest CI: the test passed absolute scratch-file paths into compareFiles()/checkRow(), whose helpers joined every input onto ROOT via path.join(ROOT, rel), producing garbage when the input was already absolute. It surfaced on windows-latest specifically because GitHub's Windows runners checkout the repo on a different drive than TEMP, so path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged (no relative traversal is representable across drives) rather than the relative form the test assumed. The remote gsd-test runner this repo gates pushes on is Linux-only and could never have caught this; GitHub CI's windows-latest job is the only signal that does, and it did. Fixed the helper itself (scripts/lint-vendored-deps.cjs's new resolvePath()) to treat an already-absolute input as absolute-in, absolute-out instead of silently mis-joining it, and updated the test to pass the scratch file's absolute path directly rather than relying on a relative conversion that is not always representable. Kept every mutation-sensor assertion intact and added coverage proving resolvePath is a no-op for relative inputs and correctly passes absolute ones through unchanged. Refs #3881 * fix(#3881): warn when state sync regenerates over unparseable frontmatter state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly overwrites an unparseable frontmatter block per its 'body wins' contract — that overwrite behavior is unchanged here. The defect was the silence: synced:true/exit 0 gave no signal that the existing block (including git merge-conflict markers) could not be parsed and was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be reported as authoritative when the derivation dropped input it could not resolve') and §8.4 ('failure is a value'). Adds a gsd: warning — ... (#3881) line on stderr, matching the existing #3573 precedent, and surfaces the same disclosure in the JSON result's existing changes[] array so a machine consumer sees it too. Exit code and synced:true are left unchanged — sync did what its contract says. REGENERATE_STATE (/gsd-health --repair's sibling on the same sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally refused by applyRepairs's dispatcher before runRepairAction ever runs (src/health-diagnostic.cts), so it is not a live path today and is not in scope for this fix. Refs #3881 * fix(#3881): exit non-zero when a state command returns an error Refs #3881 * chore(#3881): changeset for the state exit-code fix Refs #3881 * fix(#3881): honor the documented --project-dir flag Refs #3881 * revert(#3881): restore exit-0 result envelopes for state errors Reverts 9638f2936 and its changeset. The change was wrong and the revert is the correction. This repo distinguishes two error mechanisms deliberately. error() in src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure path. output({error: ...}) writes a JSON result envelope to stdout and returns normally with exit 0. The reverted commit converted 23 result-envelope sites into hard failures, which is a different contract, not a bug fix. tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope contract directly -- a failing command exits 0 with a JSON error envelope and must not publish state.json -- and the remote matrix run caught it along with four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that the original commit rewrote were encoding that real contract, not the bug it claimed; they are restored. Whether an error envelope on stdout with exit 0 is the right CLI design is a genuine question, and it is section 8.4's rule ('failure is a value') with its own phase. It is not something to flip inside this PR. Refs #3881 * chore(#3881): backfill changeset PR number for the project-dir fix Refs #3881 * test(#3881): keep the frontmatter mutation shard inside its time budget The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard cap. Root cause is NOT row-level spawn overhead (contrast the #2790/ core-utils precedent): the three shard test files' own logic runs in ~413ms total (356+30+27ms) with all 392 assertions passing. Instead, src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to the vendored YAML parser, proportionally growing the mutant count Stryker generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner bills the full 'node --test <3 files>' invocation once per mutant, and node:test's default per-file process isolation forks a child process for each of the three files on every one of those invocations — pure fork overhead multiplied by a much larger mutant population. Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares isolation: 'none', and .github/workflows/mutation.yml passes --test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e. unchanged behavior — for the other 8 shards, which were not individually audited for cross-file state leakage under shared-process execution). Measured locally via node:test's run() API on the exact 3-file set: isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same 392 passing assertions. The true CI-shard number can only be confirmed on the GitHub Actions run (Stryker cannot run locally, and 'node --test' is hard-blocked in this environment). Refs #3881 * test(#3881): register the vendored-parser tests in the frontmatter mutation shard stryker.config.mjs's own rule ("Keep this list in sync with the tests arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew src/frontmatter.cts from ~825 to 1496 lines but its new tests (tests/feat-3881-yaml-parser-consequences.test.cjs, tests/frontmatter-golden-parity.test.cjs, tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in tests/frontmatter.test.cjs) were never added to the frontmatter shard's tests array, so Stryker's mutants in the new vendored-js-yaml adapter had nothing constraining them. PR #3888 measured 55.8% against the 65 floor (748 killed / 593 survived / 17 timeout) and the shard was separately cancelled at 15m04s against the 15-minute per-shard cap. Registers all four files (each earns its slot on evidence of a unique constraining assertion, documented inline), gives the shard a measured/projected 180-minute budget via a new per-module timeoutMinutes field threaded through mutation.yml's job-level timeout-minutes the same way isolation is threaded, and removes the prior isolation:'none' override (re-measured at this file-set size, its savings are within run-to-run noise, not worth the unaudited cross-file-state-leakage risk). Refs #3881 * feat(#3881): derive the mutation test list and ratchet the score floor Refs #3881 * test(#3881): ratchet five stale mutation floors and close the frontmatter gap Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1): config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78, context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs RATCHET_BASELINE in the same diff per the ratchet's own contract. Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter), repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via extractFrontmatter), and the null-byte sentinel round-trip surviving at region offset 1. Did not lower minScore. Refs #3881 * test(#3881): decouple the ratchet test from real module floors The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour. Refs #3881 * refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency Refs #3881 * fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives. Closes #2571 Refs #3881 --------- Co-authored-by: sim <sim@local> |
||
|
|
a84f756303 |
fix(#3078): sweep all-spent ack fragments on next, name the collision remedy (#3823)
* fix(#3078): sweep all-spent ack fragments on next, name the collision remedy
`guard-no-ack-on-next` only ever watched the legacy tests/emitted-drift-ack.json.
#2914 exempted the fragment directory on the premise that a persisting fragment
"cannot conflict with any other PR". Fragments do not share a FILE, but they do
share a PATH KEY SPACE, and a path claimed by two sources is a hard failure in
the same script -- so a fully-spent fragment on next owns keys it can no longer
gate, and the next PR to grow one of those paths can declare it neither there
(spent) nor in its own fragment (duplicate). Measured at the sweep: 45 fragments
owning 403 paths, up from 13/272 at triage 19 days earlier.
- `assertNoAllSpentFragments` fails a fragment only when EVERY surviving entry is
spent against the copy at HEAD^, so a partially spent fragment -- and the
re-arm-by-appending route #2639/#2993 ship on -- keeps working.
- `ackProse` duplicates the gate's zero-width/whitespace stripping across the
scripts-ship/tests-do-not line, bounded by a prose-parity test.
- The guard job's checkout takes fetch-depth: 2; at depth 1 HEAD^ is absent and
every fragment reads as brand-new, i.e. the guard passes vacuously.
- The duplicate-ack error now names both resolutions, since the guard is
post-merge by design and cannot stop the colliding PR.
- All 45 spent fragments deleted, 0000-legacy-migration.json included, and the
three tests that pinned its permanence corrected.
Verification is the remote runner (gsd-test), not a local suite.
Closes #3078
* fix(#3078): make the prose-parity test two-sided, cover the git seam, base on the pre-push tip
Three review findings, all fixed:
- The parity test was a tautology: it checked ACK_INVISIBLE against a
hardcoded list matching its own definition, never against the gate. The
gate's INVISIBLE and its reason normalizer (hoisted out of diffEmitted as
normalizeAckReason) are now exported for that sole purpose, and the test
sweeps 0x00-0xFFFF against both surfaces. Mutation-checked: adding a
codepoint to one side and not the other now fails.
- resolveBaseRef, readFragmentAtRef and assertUsableBaseRef had zero direct
coverage -- the tests reimplemented the git reads in a local helper, so the
ls-tree-vs-show discrimination, the root-commit fallback and the
option-injection guard were never executed. All are exported and tested
against real temp repositories now, plus an end-to-end --base-ref subprocess.
- HEAD^ is not 'the state of next before this push'. The default branch allows
REBASE merges, so one push can carry N commits, and a 2-commit rebase-merge
whose first commit adds a fragment would be told to git rm it on the very
push that introduced it. CI now passes github.event.before via --base-ref and
fetches it explicitly; HEAD^ remains only the local fallback.
Also adds the safe.directory guard every other git call in this repo carries
(#2767), and stops naming the deleted migration fragment by filename in
CONTEXT.md, which tripped lint-removed-but-needed.
Refs #3078
* fix(#3078): keep the fragment directory alive after the sweep empties it
Sweeping every fragment leaves the directory untracked, and check-glossary-refs
then fails: CONTEXT.md references tests/emitted-drift-acks, which no longer
exists. The empty directory IS the intended steady state, so it has to survive
its own remedy.
Adds tests/emitted-drift-acks/README.md documenting the create/use/delete
lifecycle where a contributor actually meets it, matching the existing
tests/qa/smell-acks/README.md precedent. Every reader filters on .json, so the
README is invisible to the gate.
Also sweeps #3809's ack fragment, which the rebase onto origin/next brought in
and the new guard immediately reported as all-spent -- its own remedy applied.
Refs #3078
* fix(#3078): guard the added tests' git calls, drop a second fragment-existence pin
Both defects surfaced by the remote runner (linux-node24, 4/37445 failed).
- The new --base-ref E2E test ran `git rev-parse HEAD` against the checkout
without the #2767 safe.directory guard. The runner mounts the repo at a path
owned by another uid, so git refused every operation there with 'detected
dubious ownership'. Every git call the new tests make now names its own
specific directory as safe, via one local helper, mirroring safeDirArgs in
helpers/emitted-runtime.cjs.
- tests/agent-tracked-source-rule.test.cjs pinned the existence and contents of
the 3645 and 3409 ack fragments. That is a merged PR's paperwork, not live
behavior: once the growth is in next's baseline the acks are spent and this
PR's guard sweeps them. The third assertion pinned the hand-appended
workaround for the exact collision #3078 removes. Deleted; #3645's real
protection is the two behavioral tests above it, untouched.
Also restores #3809's ack fragment, which merged one commit before this branch.
Deleting an ack in the same window as its introducing PR races any consumer
whose baseline predates it -- the runner's container proved it, resolving
origin/next to
|
||
|
|
1178c5f995 |
test(#3108): make the overlay ENOENT tolerance and hooks/dist readiness check honest (#3772)
* test(#3108): failing-first suite for the overlay vanish-retry and hooks/dist staleness RED by construction, and deliberately narrower than the issue. #3108 reports a bare `ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'` and attributes it to hooks/dist never having been built. That mechanism cannot produce that error: buildOverlayRepo enumerates with readdirSync and links what it enumerated, so a directory that never existed yields no names and no link is ever attempted. The error requires the file to have existed at readdir and vanished before the link -- which is the atomic-replace race placeVanishableLeaf was written for in #3285, three days AFTER this issue was filed. What is still genuinely broken, and what these tests bind to: placeVanishableLeaf's retry is unguarded. On ENOENT it re-checks existsSync and calls attempt() once more, bare. A second atomic replace inside that window throws an unhandled ENOENT of exactly the reported shape. Rows 1/2/3/5/6/7 fence the surrounding contract -- most already pass, which is the point: they are what stops the fix from widening into "swallow every error". Row 6 in particular covers a non-ENOENT on the RETRY, the exact path the new code will live on. ensureHooksDist's staleness predicate is extension-blind. It rebuilds only when hooks/dist is absent or holds zero .js files, while build-hooks.js also ships .sh -- including gsd-session-state.sh, the very file in the report. A dist with .js present and every .sh missing reads as populated and the rebuild is skipped. Rows 10 and 11 are mirrored on purpose: asserting only the .sh direction would permit swapping one extension heuristic for another, so both directions force the predicate to be about the expected set (HOOKS_TO_COPY, which build-hooks.js already exports) rather than about counting an extension. Rows 12 and 13 are cost guards. ensureHooksDist runs per suite and its rebuild is a real subprocess, so a predicate that over-triggers turns a correctness fix into a throughput regression nobody attributes to it; and hooks/dist legitimately carries files the expected list does not name, so an exact-set match would rebuild forever. The warning-text row is t.skip()'d rather than faked: the only ways to assert it were a source-grep (banned by local/no-source-grep) or a full overlay build, and a test that cannot be written honestly is better skipped visibly than written vacuously. Filename note: this started as fix-3108-*.test.cjs and tripped lint-regression-test-names, which bans new fix/bug/issue-NNNN files, then as install-overlay-helpers.test.cjs and tripped lint-test-file-count, whose `install` bucket is already at its limit. Module-named under the overlay bucket satisfies both. The allowlists were left untouched -- both are empty, so nothing here is grandfathered and adding an entry would have been the wrong instinct. * fix(#3108): guard the overlay retry and make the hooks/dist check see .sh Two holes, both reachable from the failure #3108 reports, neither of them the cause it names. placeVanishableLeaf's retry was bare. On ENOENT it re-checked existsSync and called attempt() once more with no catch, so a second atomic replace landing inside that window threw an unhandled ENOENT -- exactly the reported `ENOENT ... link '/work/hooks/dist/gsd-session-state.sh'`. Two vanishes inside the window means the same thing one does: the path is going away and is not part of the snapshot. It now returns false and skips the leaf, reaching the conclusion the single-vanish case already reached. Still ONE retry. No loop, no backoff, no sleep -- the existing comment argues that a timing-based wait here would be the flake rather than the fix, and that reasoning did not change. Non-ENOENT still propagates from either attempt, which is the invariant a careless widening would eat; the suite pins it on the retry path specifically, because a fix that guarded only the first attempt would look right and be wrong. ensureHooksDist could not see the file class that caused the report. It rebuilt only when hooks/dist was absent or held zero .js files, while build-hooks.js also ships .sh -- including gsd-session-state.sh itself. A dist with .js present and every .sh missing read as populated and the rebuild was skipped. The predicate is now membership against build-hooks.js's own exported HOOKS_TO_COPY, so it asks "is everything expected present" instead of counting an extension, and it cannot be blind to a file class again. It is extracted as isHooksDistStale(dir) with ensureHooksDist calling it, so there is one predicate rather than two that can drift. Extra unexpected entries are explicitly not stale -- hooks/dist legitimately accumulates subdirectory and hooks/lib output, and an exact-set match would rebuild forever. One readdirSync into a Set, no per-entry existsSync, no stat: it runs per suite and its rebuild is a real subprocess, so an over-triggering predicate would turn this into a throughput regression nobody would attribute to it. The skipped-leaf warning now names `npm run build:hooks`. It already named the cause; a reader still had to know what produces that directory. Deliberately NOT done: nothing here makes an absent hooks/dist fail. Absence is a legitimate package shape that bin/install.js:11191 treats as "nothing to verify", and six of the seven install suites never read the directory at all. * fix(#3108): count hooks/dist subdirectories, and never throw out of the predicate Two gaps found reviewing the predicate I had just written. It ignored HOOKS_SUBDIRS_TO_COPY. That is ["lib"], and hooks/dist/lib carries gsd-graphify-rebuild.sh, so a dist with all 27 top-level files but no lib/ read as populated. That is precisely the blindness the .js-count heuristic had, one level down: a whole file class invisible to the check. Fixing the extension case and leaving the subdirectory case would have been half a fix, and the half left behind is the one nobody would look at again. Subdir names are bare (no slashes), so they slot into the same top-level readdir Set — no second readdirSync, no stat. Whether lib is really a directory is not checked; that would cost a stat per entry and buys nothing, since the build owns that. It could also throw. existsSync passing does not make readdirSync safe: the path may be a regular file (ENOTDIR), unreadable (EACCES), or retired in the race between the two calls. This helper runs in every install suite's before(), so an unhandled throw there fails a suite on a condition it cannot act on. Unreadable is indistinguishable from unusable for this question, and rebuilding is idempotent, so it now reports stale instead. Both are the same shape as the original bug and the same shape as each other: a guard that answers "is this ready" must not have a blind spot or a hard edge, because every caller treats a false negative as "carry on". Two existing tests asserted a "complete" dist without lib and had to be corrected to keep meaning what their names claim, rather than being left passing against a definition of complete that no longer holds. * test(#3108): close the review findings, including a half-closed subdir check An isolated correctness reviewer found no blockers and four real gaps. The wiring was untested. Every Group-2 test exercised the pure predicate; none called ensureHooksDist. So restoring the old inline .js-count check INSIDE ensureHooksDist -- keeping isHooksDistStale exported and correct -- left the whole suite green, and that wiring is the actual #3108 defect. Two tests now drive ensureHooksDist itself through the process seam: build invoked exactly once when stale, never when fresh. The second is the one a permissive revert fails. Reaching that seam meant requiring process-seam as a module object rather than destructuring runNode, so a test can replace it in place. That is a testability affordance in a test helper, not a production change, and it is commented as such so it does not read as an accident later. The subdir check was only half closed, and the half left open was the important one. It required `lib` to be PRESENT in the top-level readdir, never looked inside -- so an EMPTY dist/lib, missing gsd-graphify-rebuild.sh, still read as populated. That is precisely the missing-file-class case the subdir check was added to catch, which made the fix a gesture at the problem rather than a fix. Each subdir entry must now be a readable, NON-EMPTY directory. A stray regular file named `lib` throws ENOTDIR into the same try/catch and reads stale too. Cost stayed honest: one extra readdirSync total (there is exactly one subdir entry), no stat, no per-expected-file syscall. Probed against the real hooks/dist -- still reports fresh, so no suite gains a rebuild. Two nits, both real: the error-code sweep re-tested EACCES already covered standalone, and the property ignored presentAtFinalAttempt whenever vanishCount was not 1, making roughly half the 200 runs duplicates. The flag now varies meaningfully across the whole range and the assertions depend on it. One reviewer finding was already stale: the subdir and ENOTDIR work was uncommitted when the reviewer snapshotted the tree, and had landed in a67aefb9c before the report arrived. Verified rather than assumed. * fix(#3108): stop the vanish tolerance from swallowing a dest-side ENOENT A defect this PR introduced, caught by an isolated security reviewer. linkSync(src, dest) throws ENOENT for the DESTINATION path too, not only for a vanished source. The widened retry caught that, saw the source still present, retried, got the same dest-side ENOENT, and returned false -- recording the leaf as "vanished mid-walk" and printing a warning that tells the reader to run `npm run build:hooks`. A remedy with nothing to do with the actual cause, an overlay quietly short a file, and the install under test proceeding against an incomplete tree. It also falsified the function's own documented invariant, which says in as many words: "Returns false only when the path left the source tree entirely." Widening the tolerance without re-reading the sentence above it is how that happens. The retry now re-checks existsSync(srcPath) before tolerating: source still present means the ENOENT was about something else and it propagates untouched. Chose the existsSync re-check over comparing retryErr.path to srcPath -- err.path normalization is not guaranteed across platforms, and a path-equality test is a subtler thing to get wrong later. The FIRST catch was probed and is already correct: for a dest-side ENOENT the source is present, so it falls through to the retry rather than returning false. Left unchanged rather than "fixed" symmetrically. Two regression pins, deliberately opposed: a dest-side ENOENT with the source present must THROW, and a genuinely absent source must still return false. The second exists because the obvious over-correction -- always rethrow on the retry -- passes the first and silently undoes what this PR set out to fix. Also closed the skipped placeholder. It claimed no non-flaky seam existed for asserting the warning text; the reviewer pointed out an injectable `warn` param is trivial, and they were right. buildOverlayRepo now takes opts.warn defaulting to console.warn (byte-identical for every existing caller) and the skip is replaced by real tests: fires with the remedy named on a skipped leaf, silent on a clean walk. "No seam exists" was a design choice presented as a constraint. Recorded the sequential-only constraint at the two sites that monkeypatch fs process-wide: adding { concurrency: true } to this file would cross-contaminate every other suite in the process. Better written down than rediscovered. Known limit, disclosed rather than fixed here: a legitimately dropped leaf can still pass vacuously downstream -- agent-fragments-emission asserts a negative over filesContaining, and mcp-catalog-parity has only an anti-vacuity floor of one. That is a pre-existing property of those suites and the tolerance #3285 already chose; this change narrows which drops are possible rather than adding the completeness assertion those suites lack. * fix(#3108): discriminate ENOENT by the dest parent, not by re-checking the source The previous commit's dest-side guard was wrong, and the remote run said so: "a leaf that vanishes again during the retry is skipped, not a bare ENOENT" Got unwanted exception. Actual message: "ENOENT: no such file or directory" That test was right and the guard was wrong. It rethrew when existsSync(srcPath) was still true, on the theory that a present source means the ENOENT was about the destination. But in the genuine race the source is being atomically REPLACED, so it is legitimately present again at the re-check while the ENOENT was entirely source-side. The gate therefore threw on precisely the race #3285 exists to tolerate -- trading one misclassification for a worse one, since the old bug was a bare crash and the new one broke the working tolerance. The security reviewer's alternative discriminator does not work either, and a probe settles it. Node populates BOTH `path` and `dest` on a link ENOENT, and `err.path` is the SOURCE in both directions: linkSync(existingSrc, missingDir/a.txt) -> ENOENT path=<source> dest=<dest> linkSync(missingSrc, validDest) -> ENOENT path=<source> dest=<dest> So the error object cannot tell you which side failed. What CAN: the dest parent. buildOverlayRepo builds its own dest tree -- place() mkdirSync's recursively into a private mkdtempSync root no other process touches -- so a missing dest parent is always a bug (Windows MAX_PATH, a concurrent cleanup, a bad dest), never the replace race. A present dest parent means the ENOENT was about the source, which is the case we tolerate. placeVanishableLeaf therefore takes an optional destPath and uses the dest parent as the sole discriminator when it has one; with no destPath it behaves exactly as before. linkOrCopyFile and the copy-mode call site both pass it, because those are the two places that actually know the destination. The doc comment now records BOTH failed discriminators and why each fails -- existsSync because the source is legitimately replaced mid-race, err.path because it names the source either way. Those are the two things a future reader reaches for first, and both look correct until they are not. The dest-side regression pin was rewritten to drive the real mechanism: a real temp source and a dest whose parent does not exist, through linkOrCopyFile. Previously it forced a throw through a present source, which is what encoded the wrong theory into a test and made it look verified. --------- Co-authored-by: sim <sim@local> |
||
|
|
b6977d9d11 |
fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps (#3766)
* test(#3762): failing-first roster gate for docs/INVENTORY.md rows
Anchors the human half of the inventory-drift rule: every entry in
docs/INVENTORY-MANIFEST.json must have a hand-written row in
docs/INVENTORY.md. Expected RED on this commit -- next carries 32
unrostered surfaces, which is the defect the gate exists to catch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#3762): enforce docs/INVENTORY.md roster rows, backfill 32 gaps
docs/INVENTORY.md calls itself the authoritative roster of every shipped
GSD surface, and CLAUDE.md's inventory-drift rule requires both a roster
row and a manifest regen. Only the manifest half was anchored, so a PR
could ship a surface, regenerate the manifest, omit the row, and stay
green -- as PR #3758 did with gsd-core/references/planner-coupling.md.
Adds the roster half to tests/inventory-manifest-sync.test.cjs, backed by
a pure matcher in tests/helpers/inventory-roster.cjs, and backfills the 32
surfaces already missing rows on next.
Fixes #3762
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(#3762): harden roster matcher against fenced blocks and trim exports
Skips fenced code regions when splitting level-2 sections so a documented
'## ' example inside a fence cannot truncate a family section (false red)
or contribute a phantom row (false pass); makes the heading pattern linear
rather than a backtracking lazy match; narrows the module surface to the
three names the gate consumes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#3762): close two false-RED gaps found in orthogonal review
Indented headings: CommonMark permits an ATX heading to carry 1-3 leading
spaces, and a ^##-anchored pattern read such a document as having no family
sections at all -- reporting all six missing, a structural red for zero real
drift. Reproduced, then fixed and pinned.
Link-wrapped cells: a row written as [`x.md`](../x.md) was not recognized,
though docs/INVENTORY.md already uses that form elsewhere. Unwrapping now
peels a whole-cell link and a whole-cell code span, and only layers that
wrap the cell entirely -- a file mentioned mid-prose is still not a row, so
the false red is not traded for a false pass.
Adds a second fast-check property over the commands source-link rule, the
one family whose matching rule differs from the other five.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(#3762): boundary trio for the fence-marker length limit
RULESET.TESTS.boundary-coverage — the matcher's only numeric limit is the
fence marker's {3,}. Exercises 2 (inline markup, swallows nothing), 3, and
4 characters, for both backtick and tilde delimiters.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: remove stray pwned_cmdsub injection-test canary from the repo root
A zero-byte file committed by
|
||
|
|
72819a4616 |
fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi (#3731)
* test(#3031): failing-first coverage for opt-in ~/.kimi legacy reclaim Drives the user-reachable installer surface against a sandbox HOME seeded with the pre-#2755 wreckage: a GSD [[hooks]] block, hooks bundle and CommonJS marker orphaned in ~/.kimi by a --kimi-code install. Covers the reclaim itself plus the four guards the diagnosis identified as negative space: opt-in only (no flag, no deletion), user-authored TOML and hook files preserved, a --kimi install never reclaiming its own root, and the KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one directory. Adds a fast-check property that stripping the block never destroys user content. Red until --reclaim-kimi-legacy exists. Refs #3031 * fix(#3031): opt-in reclaim of GSD hooks orphaned in ~/.kimi A --kimi-code install older than 1.10.0 wrote its GSD [[hooks]] block, hook bundle and CommonJS marker into Kimi CLI's ~/.kimi. #2755 fixed the destination but could not reclaim what the old bug already wrote: the stale block is byte-identical to a legitimate Kimi CLI one — both runtimes render the same bytes for the same root, since the command paths derive from the hooks root, not the runtime — so no inspection can tell litter from a working install. Cleanup is therefore opt-in. `--reclaim-kimi-legacy` on a --kimi-code install removes GSD's own artifacts from the legacy root; without it nothing is touched, so a dual-product machine keeps Kimi CLI's hooks and #2755's acceptance criterion holds. Extracts the uninstall path's removal sequence into reclaimKimiHooksRoot() and drives both callers through it, so the reclaim removes precisely what a real uninstall removes rather than a hand-copied second implementation. Guards the wrong-runtime case (a --kimi install would delete its own hooks) and the KIMI_SHARE_DIR/KIMI_CODE_HOME collision where both roots resolve to one directory. Also corrects two pre-#2755 leftovers in the same surface that told users to run `--kimi --config-dir ~/.kimi-code` — the form that produces this very defect, since --config-dir moves only the skills root — and adds the missing --kimi-code entry to the installer's own help. Regression coverage folded into tests/kimi-upgrades.test.cjs beside the #2755 cases, per the regression-test-naming lint. Fixes #3031 * fix(#3031): never reclaim ~/.kimi when this run also installs kimi Found by the isolated adversarial review pass and independently while tracing --all ordering, then reproduced. selectRuntimesFromArgs orders 'kimi' before 'kimi-code' in both --all and an explicit --kimi --kimi-code, and installAllRuntimes installs in that order. So --all --reclaim-kimi-legacy installed a fresh, legitimate Kimi CLI hooks block into ~/.kimi and then deleted it moments later from the kimi-code leg — exiting 0 and reporting success while leaving the user with no Kimi CLI hooks at all. The collision guard could not catch it: kimi-code's own root is ~/.kimi-code, a genuinely different directory. The flag asserts "I only use Kimi Code"; installing kimi in the same invocation falsifies that, so the reclaim is skipped with a notice. Also hardens the collision guard itself. It compared path.resolve strings, which returns false for two spellings of ONE directory — measured, not assumed: a symlinked alias and a case variant on a case-insensitive filesystem both compared unequal, so the guard would not have fired and the install would have deleted its own freshly-written hooks. isSameDirectory now compares directories via resolve, then dev+ino identity, then realpath. Regression tests for all three cases; the two alias tests probe the real filesystem and t.skip() where the alias cannot exist. Refs #3031 * docs(#3031): reattach reclaimKimiHooksRoot's JSDoc to its own function Inserting isSameDirectory anchored on the function name, which placed the helper between reclaimKimiHooksRoot's doc block and the function it documents. isSameDirectory ended up with two stacked doc blocks above it and reclaimKimiHooksRoot with none. Refs #3031 * fix(#3031): warn when --reclaim-kimi-legacy cannot apply The flag only acts inside the kimi-code GLOBAL install branch. Passed with any other runtime, or with --local, it was consumed in silence: exit 0, no cleanup, no message. For a cleanup the user explicitly asked for, silence is indistinguishable from "it ran and found nothing". The scope warning is raised at argument-resolution time rather than inside install(). kimi-code declares hostBehaviors.localInstallDeferred, so install() returns early at the deferral check long before the kimi-hooks-toml branch — a guard placed there is unreachable, which is both dead code and a linted drift shape in this repo. Verified reachable by spawning the real installer. Neither case is a hard error: the flag stays composable with --all, where it is legitimately inert for the other seventeen runtimes. Refs #3031 * docs(#3031): document every case where --reclaim-kimi-legacy skips Refs #3031 * fix(#3031): resolve local config dirs from RUNTIME_META alone in the install harness The remote runner surfaced this: the #3031 warning test drives a local kimi-code install and died with "The path argument must be of type string. Received undefined". runMinimalInstall carried a SECOND, hand-maintained local-dir map beside RUNTIME_META, and it had drifted — four runtimes present in RUNTIME_META (hermes, kimi, kimi-code, zcode) were missing from it, so scope:'local' for any of them resolved path.join(root, undefined) and threw a bare TypeError naming neither the runtime nor the map at fault. #3023 had already hit exactly this for pi and fixed it by adding one more entry, which left the divergence itself in place for the next runtime to rediscover. Local scope now reads RUNTIME_META.localDir, the same table the global branch already reads, with the same loud named error the global branch raises. Parity verified for all 14 previously-supported runtimes: every one resolves to a byte-identical configDir. cline keeps its ternary — its local artifacts land at the project root itself, which is a real exception, not a directory name. Guarded in golden-parity-single-source.test.cjs beside the buildParityManifest anti-divergence test, and both arms of that guard were proven able to fail. Refs #3031 * chore(#3031): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9a69a86f42 |
enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe for /gsd-pr-branch across six layers: pure classification and forbidden-path predicates, real-git fixtures that run the cherry-pick filter loop end to end, config-key registration through the real CLI and both manifests, the executed worktree-materialization claim the issue's triage asked to establish, fast-check properties over arbitrary path sets, and a drift guard over the shipped workflow. Two live defects in today's shipped recipe are pinned as regressions, both reproduced empirically first: `git rm -r --cached` stages a deletion of any .planning/ path the target branch already tracks, so the generated PR removes the base branch's planning files; and the same command leaves the cherry-picked file untracked on disk, so a second commit touching that path aborts the pick with "untracked working tree files would be overwritten" and every remaining commit is silently dropped. The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md rather than restating them, so the workflow stays the single source of truth and the suite cannot drift from what ships. Refs #2971 * feat(#2971): strict planning filter mode for /gsd-pr-branch Adds planning.pr_strict — a boolean, default false, that selects what /gsd-pr-branch means by "filtered". Default mode is unchanged: structural planning state survives into the PR branch and the nine transient subdirectories do not. Strict mode drops every .planning/ path, structural files included, and carries a commit over only when it touches at least one file outside .planning/. Strict mode is what makes planning.commit_docs: true safe for a project that versions its planning tree locally but publishes none of it. The alternative posture, commit_docs: false, silently costs parallel executor isolation — a worktree is checked out from a commit, so an untracked or ignored .planning/ is simply absent inside it and the executor has no PLAN.md to read. That claim is now established by an executed fixture rather than inherited. The two path lists are declared once and both projections derived from them, so create_pr_branch and verify can no longer disagree about what the filter promised. verify previously counted every .planning/ path against a documented success criterion of zero while create_pr_branch was specified to preserve five structural files, so a correct run reported itself as failed on every phase that touched STATE.md — which is every phase. It now asserts against the active mode, and names the .planning/ paths default mode deliberately keeps rather than trading a wrong signal for silence. Two verified defects in the same recipe are fixed alongside, because strict mode would have amplified both. `git rm -r --cached` staged a deletion for any .planning/ path the target branch already tracked, so the generated PR removed the base branch's planning files — under strict mode that would have been the entire tree. The same command left the picked file untracked on disk, so a second commit touching that path aborted the cherry-pick with "untracked working tree files would be overwritten" and every remaining commit was silently dropped. Both were reproduced against real git before being fixed. The filter now forces excluded paths back to what the PR branch's HEAD carries, in the index and the working tree; a conflict outside the filter halts instead of being improvised past; a commit left empty by filtering is skipped rather than failing. A clean-working-tree precondition makes the worktree half safe. Closes #2971 * fix(#2971): unwind the checkout on a conflict halt, and test the real recipe Two review findings, both fixed in place. The isolated adversarial pass found that the conflict-outside-the-filter branch exited while leaving the user checked out on the half-built PR branch with cherry-pick state still live — this loop runs in the user's own working directory, so stranding them there is a real cost even though it is not a vulnerability. The branch now aborts the pick, returns to the original branch, removes the partial PR branch, and says so before exiting. The standards pass found the L2 fixtures executed a hand-written mirror of the cherry-pick filter recipe rather than the recipe itself, so a reordering in the workflow would not have been caught — and the order is load-bearing, since restoring a path from HEAD before removing it inverts the filter. The helper now extracts the canonical loop from the shipped workflow and the fixtures execute that verbatim, which also gives the conflict-halt unwind above real coverage. The drift guard additionally pins the two commands' relative order and asserts the workflow carries exactly one canonical loop. Also records the publication gate in the CONTEXT.md glossary next to the commit gate it is distinct from. Refs #2971 * fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form The remote matrix caught two defects in the previous commit. The halt path claimed to restore the original branch but did not. `git cherry-pick --abort` does not apply to a single `--no-commit` pick with no sequencer file, and the fallback left the unmerged index in place, which makes `git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so the user was told they had been restored while still sitting on the half-built PR branch. The unwind now drops sequencer state, hard-resets the disposable PR branch to clear the unmerged index, and only claims a restore when the checkout actually succeeded; when it does not, it says where the user is and gives them the two commands to finish it by hand. Verified against real git: exit 1, the conflict named, HEAD back on the original branch, the partial branch gone, a clean tree and no CHERRY_PICK_HEAD. Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form, which names a command no runtime registers. The canonical authoring token for workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the documentation added in this branch is unaffected. The comment in src/config.cts moves to the colon form too, since it propagates into the generated lib. Refs #2971 * docs(#2971): backfill PR number into the changeset fragments (#3720) --------- Co-authored-by: sim <sim@local> |
||
|
|
4e70b245e8 |
fix(#3582): stop the cold-tree fixture racing concurrent hook builds (#3656)
* fix(3582): stop the cold-tree fixture racing concurrent hook builds Two tests in tests/gsd-check-update-worker-platform-gate.test.cjs failed a verification run with `ENOENT: no such file or directory, lstat '/work/hooks/.dist-staging-20836'`. This is a race I introduced in #3582, not a flake, and it passed when #3582 merged because it only fires when the timing lines up. buildColdInstallTree() copied the LIVE repo hooks/ directory with a filter that excluded only the basename 'dist'. scripts/build-hooks.js writes atomically through a per-PID staging dir, hooks/.dist-staging-<pid>, and removes it when finished — and the archived build-hooks-atomic-write changeset records that NINE test files invoke build-hooks.js from their before() hooks. So several test processes create and delete staging directories inside hooks/ while other tests are reading it. cpSync enumerated one, and the owning process removed it before cpSync got to it. The helper's own header already states the rule it needed: hooks/dist is excluded because it "is not present in a raw marketplace checkout either". hooks/.dist-staging-* is gitignored (.gitignore:21) and equally absent from a raw checkout — it was simply missed. Fixed by enumerating hooks/ explicitly and skipping 'dist' and any '.dist-staging' prefix BY NAME, before anything stats or copies the entry, then copying each surviving entry individually. A name-first skip means a vanishing staging dir is never touched at all. Worth recording because it corrects the assumption this fix was written under: cpSync's filter IS invoked before the entry is lstat'd, and returning false leaves it untouched (verified by deleting inside the callback and returning false — no throw). So merely adding '.dist-staging' to the old filter would also have closed the race. The explicit enumeration was kept anyway so correctness does not depend on that Node implementation detail. Proven by execution both ways: with a staging dir planted in hooks/, the OLD cpSync-with-filter form copied it straight through into the fixture, while the new form succeeds and produces no .dist-staging entry with the real hook set intact. Regression test added beside the existing cold-tree tests: it plants a real hooks/.dist-staging-test-<random>, asserts the fixture builds clean without it, and removes only the directory it created. Repo swept for the same exposure: this helper is the only place doing a bulk enumeration of the whole live hooks/ tree. The other hooks/-touching tests reference specific named files or hooks/dist/ and are not exposed. scripts/build-hooks.js is deliberately untouched — its per-PID staging is what makes its own writes atomic and is correct. Refs #3582 * fix(3582): make the race regression test hermetic instead of mutating the live tree The regression test added in the previous commit failed the runner with "failed running after hook", and it was wrong in two ways — the second one worse than the first. cleanup() (tests/helpers.cjs:452-487) deliberately THROWS for any path outside the known temp roots. The test planted hooks/.dist-staging-test-<random> inside the repo and then asked cleanup() to remove it, so the after-hook threw. That guard is correct and is left alone. The real problem is that the test mutated the LIVE hooks/ directory while other test files concurrently read it — the exact shared-state hazard this change exists to remove. A regression test for a race must not introduce one. buildColdInstallTree now takes an optional opts.repoRoot (defaulting to the real REPO_ROOT and used for both copies it performs), so the test builds a fake repo root under the temp dir, plants representative hooks plus dist/ and .dist-staging-99999/ THERE, and asserts the fixture excludes both. All six pre-existing callers pass no arguments and are unaffected. The test also asserts the real hooks/ listing is identical before and after, so a future edit that reintroduces live-tree mutation fails loudly. The name rule is now pinned directly rather than only through the copy. shouldCopyHookEntry is exported and asserted, including the two cases a sloppier implementation would get wrong: 'dist-staging-no-dot' and 'distant.js' must both be KEPT. Anything matching on a loose 'dist' substring or startsWith passes every other case and fails those two. Also corrected the issue number on the tests introduced here: they were labelled #3631, which is the unrelated capability-consent bytecode work. This is #3582. Verified by execution: the predicate rule holds on all nine cases; a fake-root fixture yields exactly the representative hooks with dist and .dist-staging excluded; the no-arg default still copies the real tree (29 entries); and the real hooks/ listing is byte-identical before and after. Refs #3582 * chore(3582): re-trigger CI after an orphaned Validate Branch Name run The Validate Branch Name run for this branch (32211622051) sat queued from 03:17 and was never picked up — updatedAt never advanced past createdAt while the same workflow completed normally for other branches. `gh run rerun` refused it ("already running") and `gh run cancel` returned HTTP 500, so the run is orphaned on the GitHub side. Closing and reopening the PR re-fired the other pull_request workflows but not that one, whose triggers evidently do not include reopened. An empty commit is the remaining way to get a fresh run. No file changes: the tree is identical to dc71534b6, whose remote-runner pass carries forward unchanged. Recording this rather than admin-merging past the pending check. Everything else was green (24 pass, 0 fail), but admin merge is sanctioned only for the missing-secondary-reviewer case, never to skip a gate that has not actually run. Refs #3582 --------- Co-authored-by: sim <sim@local> |
||
|
|
bf2332e67c |
fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in its fail-closed catch and is misreported as an unreadable dispatch-isolation configuration, blocking every executor dispatch. These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor guard and update worker, plus the seam's actionable build error surfacing instead of the generic misreport. Also adds the drift lint the acceptance criteria require, with a fixture proving it CAN fail — a guard never shown to fail is worthless. It is red here by design: it flags today's unfixed hooks, which is exactly the defect. * fix(3582): route every hook's compiled-module require through the self-heal seam RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard, cursor guard and update worker, the fail-closed-with-actionable-message assertion, and the lint's own real-tree check. The compiled runtime library is produced by build:lib and gitignored (ADR-457, build-at-publish). The npm package builds before publishing; a plugin-marketplace or git-clone install materializes the raw tree and never does, so on that channel those modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly this, and the CLI entrypoint already calls it — no hook did. The isolation guard's Cannot-find-module therefore landed in its fail-closed catch and was reported as 'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was misdiagnosed as an unreadable project config and every executor dispatch was blocked. All SEVEN affected files now call the seam before their first compiled require. The issue named four; a scan found six; implementing it surfaced a seventh — the shared isolation sentinel helper, used by BOTH guards, which requires two compiled modules itself and would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so fixed here rather than left as a known-broken remainder. Failure posture is deliberately split by hook kind: - Gates (agent isolation guard, cursor subagent start) surface the seam's actionable build error distinctly instead of swallowing it into the generic text, and stay fail-closed — a genuinely unreadable project config still DENIES exactly as before. - Cosmetic and detached hooks (statusline, update worker, update check, update banner) DEGRADE rather than crash: the statusline draws on every render and the worker is a detached process, so a build failure there must not take down the prompt. The npm path is untouched: the seam's already-built fast path returns immediately, so prebuilt installs pay nothing and behave bit-for-bit as before. Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than remembered — without it the next hook to add a compiled require reintroduces the class silently. It is proven able to fail: a fixture hook requiring a compiled module without the seam is flagged, and one that uses the seam is not. Verified directly — on the unfixed tree it named all seven offenders; with the fix it passes. While writing the lint's comment stripper, a naive whole-text block-comment regex ate its own fixture, because this repo's comments legitimately spell the compiled-lib glob whose star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression test pinning that case. * fix(3582): test the three untested seam call sites and assert typed reason codes Two independent reviews converged on the same major gap: the fix wired the seam into seven files but only four had cold-tree tests. The adversarial pass put it plainly — deleting the shared isolation-sentinel helper's seam call would not have failed any test in the diff. That file was my own addition beyond the issue's four, so it shipped untested; that is now closed. - Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT directly under cwd, and every existing cold-tree fixture puts it there, so the early return always fired first. Now covered, and proven load-bearing by mutation: with the call removed the spy records zero seam invocations and the test fails. - update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED VERDICT — the fallback cache filename, and silent suppression when the package name degrades to null — rather than merely 'did not throw'. The banner hook previously had no test file at all. Standards violation fixed: two tests asserted on free-form prose via assert.match against a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch it. Both isolation guards now emit a machine-readable reason_code from a frozen enum, following the repo's existing REASON convention, and the tests assert that instead. The human-readable message is unchanged for operators; only the assertion target moved. The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT extracted, and the reason is recorded at each site: both viable shapes — a path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file literal co-occurrence check, so extracting would require the lint to special-case its own helper. Triplication is the lesser evil while the lint stays a co-occurrence scan. The lint's header now states what it does and does not catch (literal quoted requires only; hooks/ scan root), so a future reader does not over-trust a guard that a concatenated path or a require inside a non-hooks helper would evade. * chore(3582): regenerate the committed install-tree fixtures Adding a new shipped hook helper changed the install tree, and those fixtures are committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>' tests failed on 541a1913. Regenerated rather than hand-edited. The delta across all 15 runtime fixtures is exactly two lines — the new helper under both its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in no unrelated drift. This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from lint:ci, which passed both before and after. * chore(3582): backfill changeset PR number (#3629) --------- Co-authored-by: sim <sim@local> |
||
|
|
7c649a9970 |
fix(#3585): close raw-git bypasses of the commit_docs gate (#3590)
* test(#3585): repo-wide guard for unguarded .planning/ git add Replaces the two-file #1783 scan, which required .planning/ on the git add line and so was structurally blind to fast.md's `git add -A` and to new-milestone.md (never scanned). Extracts the shell tokenizer, comment-position rule and gsd-scan-ignore marker from the #2269 guard into tests/helpers/shipped-command-scan.cjs so both guards consume one implementation. Commit-specific logic stays in commit-files-pathspec.test.cjs; every pre-existing test there passes unedited. Fails RED on five sites: fast.md:58, new-milestone.md:262, spec-phase.md:480, eval-review.md:148, ai-integration-phase.md:263. The last three carry a markdown prose conditional outside the bash block it claims to guard. * fix(#3585): close raw-git bypasses of the commit_docs gate Five shipped workflow steps staged .planning/ with raw git. Two had no check at all; three had a markdown prose conditional sitting outside the bash block it claimed to guard, so the block ran unconditionally. spec-phase, eval-review and ai-integration-phase now route through the gsd_run query commit seam, which performs the commit_docs and gitignore checks internally and returns a skipped envelope -- this deletes the raw git pair rather than wrapping it. new-milestone stages directories for a later commit and cannot use the seam, so it takes the executable guard form, fail-open on a tooling error. fast writes no planning artifacts and has no gsd_run in scope at that point, so it excludes .planning via pathspec instead of reading config. Guard now reports 0 offenders. * test(#3585): pin skipped_gitignored to production behavior COMMIT_REASON was a test-local frozen enum joined to production only by a hand-maintained keep-in-sync comment -- the Generative Fix Divergence class, whose required remedy is a parity assertion. B1-B3 already pinned SKIPPED_COMMIT_DOCS_FALSE. SKIPPED_GITIGNORED was pinned by nothing: production could rename it and every test still passed. G1-G3 drive the gitignore auto-detect path and assert the canonical reason. The fixture must OMIT .planning/config.json entirely -- with config.json present the loader resolves commit_docs to false first and cmdCommit returns skipped_commit_docs_false, never reaching its own isGitIgnored branch. * docs(#3585): document the planning commit gate and its guard CONTEXT.md had zero commit_docs entries. Adds a Planning Commit Gate glossary entry covering the resolution chain, the typed skip envelope, the measured ordering of the two reason codes, and why the gate is enforceable only as a text guard. CONTRIBUTING.md gains the contributor rule for the new guard, with the prose-is-not-a-guard example that caused three of the five defects. * fix(#3585): address review findings in the planning-add guard Spec review (blocker): fast.md excluded .planning unconditionally, changing behavior for commit_docs=true users and violating epic AC4. Now gated -- the launcher preamble was MOVED from log_to_state into the commit block rather than copied, so gsd_run is in scope for +4 lines instead of +4KB, and the else branch is byte-identical to the previous git add -A. Security review (major): git -C <dir> add was a false negative because the flag-skip loop never modelled flags that consume a separate value. Fixed for -C/-c/--git-dir/--work-tree/--namespace. The fail-closed rule now also covers $(...) substitution args and --pathspec-from-file, which were opaque in the same way $VAR is. git commit -a/-am is now classified as reaching, since it stages every tracked modification. Self-review: isSkippable treated any NAME= token as a skippable prefix, so V=$(git add -A) escaped -- the exact divergence the shared-helper extraction existed to prevent. Adopted the sibling predicate verbatim. eval, xargs, one-line function bodies and line-continuation remain blind and are now enumerated as declared limits in the guard docblock and CONTRIBUTING. The ifDepth clamp is defensive only: a 200k-case differential fuzz found no reproducing input, so its test is labeled a pin, not a failing-first test. * test(#3585): acknowledge emitted growth in three workflow files emitted-attribution has two arms: hash attribution AND per-file growth. The growth arm needs an acknowledgment even when every moved byte is attributable to the diff, which is why the first remote run went red on it. fast.md +417: the launcher preamble moved into the commit block so gsd_run is in scope for the commit_docs guard, plus the guard itself. new-milestone.md +281: the executable guard plus one line recording that the unstaged archive move is deliberate. spec-phase.md +21: reworded prose describing the skipped envelope. eval-review.md and ai-integration-phase.md shrank; no entry needed. * test(#3585): drop duplicate spec-phase ack, shrink its prose instead The base already acknowledges spec-phase.md (from #2733), and two ack sources may never name the same path. But a base-side ack is SPENT -- it cannot clear new growth -- so the two gates were in direct conflict: attribution wanted an ack, the ack lint forbade one. Resolved by removing the growth rather than the conflict. spec-phase.md's +21 was purely a prose reword; rewritten shorter, the file now shrinks 36 bytes against base and needs no acknowledgment at all. fast.md and new-milestone.md have no base ack and keep theirs. * chore(#3585): backfill changeset pr number to 3590 --------- Co-authored-by: sim <sim@local> |
||
|
|
9448736872 |
fix(#3547): exercise the real global config-home shape in the install harness (#3567)
* test(#3547): failing-first regression for collapsed global install shape * fix(#3547): exercise the real global config-home shape in the install harness * test(#3547): align ripple suites with the real global install shape * test(#3547): update stale collapsed-shape pins in provenance and migration suites * fix(#3547): bump emitted-baseline schema version for the real install shape --------- Co-authored-by: sim <sim@local> |
||
|
|
147856040b |
fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker toggled on any delimiter, so a backtick fence could be closed by a tilde one and an include in the gap was rewritten inside a code block. Fixed by reusing scanFencedBlocks - the canonical engine already behind stripFencedCode and extractFencedBlock - rather than carrying a fourth copy of fence detection, which also closes the duplication the standards review flagged. sanitizeForRender now strips combining marks and zero-width characters alongside the ANSI, control and bidi classes it already handled. Adds the C, E and F matrix rows the spec review found missing, including installer-level coverage that spawns the real install rather than calling the report builder. Ships the how-to, the reference and command docs in five locales, the changeset, the inventory and glossary entries, and regenerates health.md for the new W028 rule. Refs #2873 |
||
|
|
dc3c81e93d |
chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both suites fail with MODULE_NOT_FOUND, which is the intended RED. Locks the measured behavior rather than the assumed behavior: RegExp.escape hex-escapes the leading character of nearly every string ("abc" -> "\x61bc"), so the suite asserts match-equivalence against an inlined historical oracle (the implementation being deleted) rather than byte-equivalence of pattern text — 200 seeded fast-check runs plus a fixed corpus, 0 mismatches. Also locks the latent character-class range bug this phase fixes as a side effect: a hyphen-bearing value interpolated into [...] currently forms a real range and matches an unintended character; post-migration it must not. * chore(#3412): src/pattern.cts owns runtime-value regex construction Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam delegating to the built-in RegExp.escape, deletes every hand-rolled copy, and raises the Node floor to the Active LTS line. The census was low, three times over. ADR-3212 counted 10 copies; a graph query found 12; the new lint rule — once live — found 27 more. The difference is that the census counted named helper FUNCTIONS while the rule counts the escape SHAPE, so inline .replace(<class>, '\$&') copies were never in scope. ADR §1's actual requirement is that no module outside the seam escapes a value for regex use, so all of them are, and CLAUDE.md's no-defer rule makes them this change's work. Fourth consecutive epic here whose copy count was low — the argument for ADR-3180 Amendment 3's "state N found by the guard" rule. Also corrected mid-implementation: the survey reported phase-id.cts's escapeRegex had 0 external importers. It had 8 production importers, making its removal a public-surface change to an ADR-2121-owned module and requiring an update to that ADR's locked-surface test. Blast radius revised Medium-High -> High. RegExp.escape is match-equivalent but NOT text-equivalent: it hex-escapes the leading char of nearly every string ("abc" -> "\x61bc"). Equivalence is proven by a seeded fast-check property test against the deleted implementation as oracle. It also fixes a latent bug: a hyphen-bearing value interpolated into a character class previously formed a real range and matched an unintended character. Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines, .nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate `required-tests` context is unchanged and no job was added or removed, so branch protection cannot be orphaned by the dropped lanes. Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with structural provenance for reviewed pattern-fragment constants rather than a name heuristic) plus a whole-tree companion guard covering the directories ESLint's globs miss. * fix(#3412): close the _SOURCE guard evasion, correct two false claims Three findings from the orthogonal review pass, all fixed. 1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier- name matching with no binding check, so `new RegExp(userInput_SOURCE)` — a function parameter — sailed past the guard. That is the same rename-evasion class issue #3410 documents, reopened by the very fallback meant to complement the structural check. Now bound to the identifier's actual binding kind: import, require-derived const, or module-scope const; parameters, `let`/`var`, and unresolvable bindings fail closed. Four RuleTester cases cover the evasion and prove the legitimate cross-module case still passes. 2. src/pattern.cts's own header carried the stale pre-correction counts (12 copies / 17 call sites) while CONTEXT.md and the design doc carried the corrected ones (~39 / ~44) — a self-contradiction inside the PR whose entire purpose is deleting divergent copies. Rewritten, preserving the durable lesson: a named-function census cannot see inline copies; only a shape-matching guard can. 3. The claim that all deleted copies threw TypeError on non-string was false. phase-id.cts's copy — the one with 8 external importers — did String(value).replace(...) and never threw. The seam's locked signature does not coerce, so this is a real, now-disclosed behavior change rather than the pure preservation the tests asserted. Audited all 32 invocations across the 8 importers and 6 in-file callers: every one is safe by construction (upstream truthy guard or a string-producing derivation), verified by runtime probe against the compiled modules rather than by TS compilation, which cannot see a runtime undefined. Corrected the false claim in both the test comment and the design doc, and added it to Known limits. * docs(#3412): add Changed changeset for the Node 24 floor The only user-visible break in this phase. The escape-behavior change is internal and match-equivalent, so it carries no user-facing note. * fix(#3412): resolve the seam's require graph in script fixtures and packaging Checkpoint 2 came back red with 90 failures on the node24 lane. Three distinct defects, all introduced by routing scripts/ through the new pattern seam, none reproducible by any local gate: 1. ~82 failures — tests/adr-index-gate.test.cjs and tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an mkdtemp fixture and spawn it there (necessary: those scripts resolve their scan root from __dirname/.., so running the real script would scan the real repo). Each harness hand-listed the dependencies to copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to gen-adr-index.cjs made both lists silently incomplete -> MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON' failures from the same crash. Fixed as a class, not an instance: new tests/helpers/copy-script- fixture.cjs walks a script's transitive static relative-require graph and copies it, so dependencies are derived and never re-declared. It throws (naming the unbuilt artifact) instead of letting the child die with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host- contract, sync-runtime-launcher. 2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so the new scripts/lint-no-adhoc-regex-escape.cjs would be MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from the tarball, matching the existing precedent for gen-emitted- baseline.cjs, which is excluded for the identical reason, and locked with a test modeled on that one. Confirmed against a real npm pack: 890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs present (so the other four scripts' requires are legitimate). 3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to the retired hand-rolled escaper but NOT text-equivalent: it hex- escapes the leading character and all hyphens ('0*\x329', '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match decisions across all three real interpolation prefixes, zero divergence. Those tests now compile each source into the same heading regex src/roadmap.cts's searchPhaseInContent builds and assert what matches and what does not, including the 'i'-flag canonicalization the hex escape has to preserve. Re-pinning the new literals would have rebuilt the same brittleness one layer down. Adds a test for the property the escape exists for: a dot in '1.2' must not act as a wildcard. Also shares one definition of 'a require' between the packaging guard and the fixture copier, so the two cannot disagree about what they scan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): refuse to copy a fixture dependency outside the fixture root copyScriptWithDeps resolved each relative require and joined the repo-relative result onto fixtureRoot. A require resolving OUTSIDE the repo yields a '../'-prefixed relative path, so path.join climbed out of the fixture and wrote into the surrounding temp dir (verified: repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd). No script in the tree does this today, so this closes an available escape rather than an active one. Refuses via the existing unresolved- require path so the failure names the offending specifier. Covered by a negative proof that the guard fires and that nothing lands outside the fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract Applies all findings from the second orthogonal review round, re-run because real code changed after round 1. HIGH (security) — extractRequires stripped BLOCK comments before LINE comments, so a '//' comment containing '/*' opened a phantom block comment, and a '//' inside a string literal truncated the line. Both hid real requires: 'const u="http://x"; require("./real.cjs")' returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were invisible. Replaced with a real AST parse via espree. This is ADR-3212's own Decision 4 — tokenizer-first for stateful grammars — applied to the case it describes; comment/string/regex nesting is exactly such a grammar, which is why the regex version was wrong. The function was moved byte-identical out of the #2858 packaging guard, so the bug PRE-DATES this branch and has been a live blind spot there: a shipped script could have required an unshipped path undetected. Fixing it makes that guard strictly stronger than on next. espree is promoted from a transitive eslint dependency to an explicit devDependency rather than relying on hoisting. The script parse attempt sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a function, making a top-level return legal — scripts/check-coverage-gate .cjs relies on it, and without the flag the guard throws on a file it is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a real npm pack, so the exact extractor does not newly fail the guard. MEDIUM (security) — the repo-containment check guarded dependencies but not the entry path. One escapesContainment predicate now guards both. LOW (security) — containment was lexical while fs follows symlinks, and a directory symlink could mint a fresh dedupe key per level. realpath now resolves both repoRoot and each dependency before the decision, and the realpath-derived path is the dedupe key. Destination layout still uses the original repo-relative path, so copied trees are unchanged. MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests lost the foreign-prefix contract: every assertion was satisfied by an impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599 bug class the exact-source prevents. The literal assertions it replaced were catching this. Now asserts the compiled regex REJECTS a different prefix with the same number. MAJOR (standards) — the test hand-duplicated production's heading regex with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed the parallel surface instead of policing it: src/roadmap.cts exports buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports it. Byte-identical .source and .flags verified for both escaped forms. MINOR — '..foo' no longer false-flagged as an escape; the inverted spurious-vs-missing doc claim corrected; the dead allow-test-rule header removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3412): backfill changeset pr number to 3416 * fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision Two CI failures on PR #3416, both in code this branch added. CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a bracket run be consumed EITHER by the character-class branch OR one character at a time by the trailing catch-all, so a failing match explored both parses of every pair. Measured on the real regex: n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script scans repo source, so a file with a long bracket run after '.replace(/' would hang CI outright — a guard against undisciplined pattern construction was itself the worst pattern in the diff. Fixed the way ADR-3212 already prescribes: the catch-all branch now excludes '[' and ']' so a bracket can only be consumed by the class branch (this is what makes it linear), and every quantifier is bounded (the locked bounded-quantifiers decision) as a second line of defense. Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the constant: a regex literal with a BARE unescaped ']' outside a class is no longer matched by this backstop. No census shape has that form, and the AST rule remains the primary detector. Verified the guard did not go blind doing it: a real census-shape violation is still reported, and an allow-adhoc-regex-escape suppression comment is still honored. Regression test drives the exported findViolations on a 2000-repetition adversarial input and asserts the RESULT. It makes no wall-clock assertion — elapsed-time tests are forbidden — so a regression surfaces as a harness timeout, which is the correct signal. Prompt injection scan — 'must not act as a regex wildcard' in a test comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if| my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a whole test file over one phrase would blunt the scanner permanently, and the comment has nothing to do with injection. Neither failure was reachable from the remote runner — CodeQL and the injection scan are not in that matrix, so the sha it passed was green and still wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7a7bf19fc1 |
enhance(#2872): record scope and runtime in the install manifest (#3323)
* enhance(#2872): record scope and runtime in the install manifest gsd-file-manifest.json gains manifestVersion, runtime and scope, and a new read-only Installed Surface Resolver Module reads both install scopes for a runtime in one call -- the first code path in the repo that does. Phase 3 of epic #2866 (ADR-2866). Blocks Phase 4 (#2873), which resolves #2218: the resolver's shadowedBy field is that defect expressed as a value for the first time. It ships computed-and-unread here. Installed-ness is decided by manifest PRESENCE, never by the new fields, so a manifest written by an older GSD stays fully functional and no user needs to reinstall. Recorded runtime/scope are corroboration: a disagreement with the probed config dir is reported as declaredScopeMatchesProbe: false, never silently corrected. readInstallManifest is widened additively -- version/timestamp/mode/files keep their exact names, types and meanings for all four existing callers. manifestVersion is a new field rather than a reinterpretation of version, which holds the package version and is read by the golden-parity fixtures. Stems are derived from the installed manifest's own file keys, the inverse of Phase 2's filename composition, guarded by a fast-check round-trip property plus a kebab-case charset check so a crafted manifest key cannot put a traversal segment, control character or ANSI escape into a trigger that Phase 4 renders back to the user. Also fixes two defects found while working: - bin/install.js hardcoded manifestVersion: 2 while the reader owned MANIFEST_SCHEMA_VERSION = 2. Now single-sourced, with a parity test. - docs/installer-migrations.md documented an install-state schema of five snake_case fields that have never been written; InstallState has only ever been { schemaVersion, appliedMigrations }. Corrected with a dated note. Verification runs on the remote runner. * fix(#2872): fold review findings from three independent engines Standards axis: - convert the manifest-schema suite from a hybrid setup(t) closure to beforeEach/afterEach (CONTRIBUTING.md:319-354 Pattern 1). The hybrid was neither approved pattern and a new test forgetting the call got no warning. - SCOPE_ORDER was declared twice with no parity test -- this repo's recorded generative-fix-divergence class. Give the ordering one owner: install-scope exports it frozen, the layout module and the resolver both import it, and a test locks it against scopeRank so the constant and the ranks cannot drift. - drop the defaultReadManifest passthrough (Middle Man). Spec axis: - add the VOLATILE_FILES exclusion test and source comment the acceptance table promised and did not deliver. gsd-file-manifest.json stays excluded: the new fields are deterministic, but timestamp -- the original reason -- is unchanged. Security axis: - bound the reported manifest runtime at 64 chars, matching the truncatePostureValue convention already used in this subsystem. It reached declaredRuntime unbounded while the adjacent stems were gated by SAFE_STEM; an inconsistent posture on the same attacker-influenceable document. The charset stays ungated on purpose -- declaredRuntimeMatchesProbe needs to see the real value -- so Phase 4 must sanitize before rendering, recorded in the design's Known limits. Both new parity tests were verified to FAIL when the two sides are made to disagree, then pass again on revert. Verification runs on the remote runner. * chore(#2872): backfill changeset pr number to 3323 * fix(#2872): give git fixture construction its own timeout class PR #3323's full test (windows-latest, 22, shard 2/3) failed with gitOrThrow: 'git init' failed -- outcome=timed_out exitCode=null gitOrThrow: 'git commit --allow-empty' failed -- outcome=timed_out from drift-detection.test.cjs's beforeEach, a file this branch never touched. Every other lane passed the same commit, including windows-latest node 24 on all three shards, and next is green. Root cause is a bound sized for the wrong class. DEFAULT_GIT_TIMEOUT_MS is 15000 and its own comment scopes it to plumbing READS -- rev-parse, branch, log -- against an existing repo. createFixture uses it for six sequential repo-CONSTRUCTION spawns: init, three config writes, add -A, commit. init and commit each write dozens of files, and on Windows every spawn is Defender-scanned. Sibling tests in the failing block took 15.6-22.0s against a 15000ms bound. This repo already diagnosed this exact shape once: timeouts.cjs's HOOK_FANOUT_TIMEOUT_MS records PR #3285 failing in the SAME job with the SAME outcome=timed_out exitCode=null signature at the SAME bound while every other lane passed, and concludes 'a bound sized for the wrong class, not a slow machine'. It was fixed by splitting out a heavier class-norm at 60000. Same remedy here: GIT_FIXTURE_TIMEOUT_MS = 60000, 4x the bound that failed and half INSTALL_TIMEOUT_MS. DEFAULT_GIT_TIMEOUT_MS deliberately stays at 15000 -- a blanket raise would stop a genuinely hung plumbing read from surfacing fast. Verified the value reaches the spawn rather than being an ignored option: spawnSync was monkeypatched before requiring the fixture module, and all six git construction calls were captured carrying timeout: 60000. This branch's two new test files shift shard composition, which is how a pre-existing fragility landed in the heaviest shard on the slowest lane. Fixed here rather than deferred, per the no-defer rule. Verification runs on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
c28134ab39 |
fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by |
||
|
|
2e2b8ba4a7 |
enhance(#2704): resolve documentation links and compare H1 status brackets in the ADR gate (#3266)
* test(#2704): failing-first coverage for ADR link resolution and H1 status brackets Binds the gate to two assertions it does not yet make: every relative markdown link under docs/adr/ must resolve, and an H1 trailing status bracket must agree with the Status: field instead of being silently stripped. Covers all 51 rows of the phase test matrix across two altitudes - the pure extractLinks/maskCode IR for fence and inline-code-span boundaries, hostile input and the fast-check totality properties, and the real CLI verdict for the end-to-end classes. Includes the DEFECT.GENERATIVE-FIX parity test that iterates the exported STATUSES array so a sixth status is covered the day it is added. * feat(#2704): resolve ADR documentation links and compare H1 status brackets The ADR gate validated naming, relation symmetry and index freshness but never resolved a link target, and it stripped an ADR's trailing H1 status bracket for display rather than comparing it against that ADR's own Status: field. Both classes were structurally invisible: #2691 found five dangling references by manual audit roughly a year after they were introduced, one of which reached the published npm payload, while CI reported green throughout. Both are now assertions on the same --check path, using only node:fs and node:path - no dependency and no subprocess. Fenced blocks and inline code spans are masked before scanning, because markdown does not render a link inside code. That is not a policy choice: the corpus contains exactly two such sequences today and both are ordinary JavaScript. Masking preserves length and column positions so findings still name a real line. Resolution is case-exact on every platform - a link that resolves only through macOS or Windows case-folding still 404s on github.com and still fails the Linux lane - and a destination resolving outside the repository is reported before any filesystem call is made. Also single-sources two duplicated surfaces this change would otherwise have extended: the H1 bracket vocabulary (a second hand-written copy of STATUSES with nothing asserting agreement, a DEFECT.GENERATIVE-FIX instance) and the docs/adr directory traversal. Two tests added by #2691 that reimplemented link resolution and bracket comparison inside the test file are removed for the same reason; the corpus assertion is now made by running the real gate against the real corpus. * fix(#2704): reject symlinks that leave the repository and linearize code masking Four defects from the isolated adversarial security review, plus one it noted. BLOCKER - a symlink defeated path containment. path.relative(ROOT, abs) is purely lexical, but the case-exact walk then calls readdirSync, which follows symlinks at the OS level: a contributor-committed docs/adr/x -> /etc together with a link through it passed containment and listed the real external directory, and a wrong-case probe echoed a real external filename through the "Did you mean" hint into publicly-readable fork-PR logs. Every segment is now lstat'd before descent; a symlink is realpathed and re-checked against realpath(ROOT) - realpath on both sides, so a root under /var does not produce false escapes - and an escape emits no hint and reads nothing further. The same rule now governs which FILES are read: an ADR entry that is a symlink out of the repository is excluded and reported rather than parsed, closing the vector this change had widened by newly reading README.md, naming-violation files, and full bodies rather than only header fields. MAJOR - inline-span masking rescanned the line remainder per backtick run, roughly O(n^1.6) on adversarial input: 1.76s for an 800KB line. Rewritten as a single linear pass pairing runs through forward-only per-length cursors. Same input now takes 3.31ms, with behavior unchanged. MINOR - an unreadable or broken entry threw, and the generic handler wrote a raw stack trace carrying absolute CI paths to stderr. The scan is now fault-tolerant and reports excluded entries as ordinary violations. The status vocabulary is escaped before being interpolated into a dynamic RegExp - defence-in-depth, not a live bug. The containment predicate had reached three hand-written copies while fixing this; it is now the single escapesRoot() helper used by all four call sites. * feat(#2704): add a --json report so the gate's tests assert on typed values Maintainer-directed addition. CONTRIBUTING.md's "Prohibited: Raw Text Matching on Test Outputs" requires that a system under test producing text also expose a structured intermediate representation, and that tests assert on that IR rather than on rendered prose. This gate had no such surface, so its verdict tests matched on stderr. --json runs exactly the same validation as --check and writes a report to stdout with the same exit code, following the frozen-REASON-enum pattern already used by verify-reapply-patches.cjs. Every violation carries a stable reason code plus the fields a consumer needs, so nothing has to pattern-match an error message. Adding a reason stays three coordinated changes - the enum, the emitting site, and the test locking Object.keys(REASON).sort(). The human output is unchanged, deliberately: a large pre-existing suite asserts on it and migrating that is not this PR's concern. Verified by running the pre-change and post-change scripts against an identical violating corpus and diffing their stderr - character-for-character identical. This PR's own verdict tests now assert on parsed --json. Absence checks improve the most: "no bracket violation" is now a reason-code predicate rather than a negative regex over prose, which could pass for the wrong reason. The security assertions were strengthened rather than translated - no leaked filename may appear in ANY field of the serialized report. Unknown flags are now rejected instead of silently falling through to printing the index. * test(#2704): fix the status-parity fixture and guard hooks/dist before overlay builds Two failures from the matrix run of 79b29909. The status-parity fixture was mine. It built, per status token, an ADR whose H1 bracket and Status field both carried that token - but Superseded carries an obligation beyond the bracket: it must name its successor as a file link and be symmetric with it. The fixture declared a bare Superseded, tripped that unrelated invariant, and the test reported a bracket-parity failure for a reason that had nothing to do with bracket parity. The fixture now satisfies each token's own obligations in both the agreeing and contradicting corpora, derived from the status actually declared rather than special-cased on one name, so a future token carrying obligations is handled rather than silently skipped. The second failure was not mine but is fixed here rather than deferred. mcp-catalog-parity.install.test.cjs hardlinks hooks/dist/* while building its overlay, but hooks/dist is a gitignored build artifact produced only by build:hooks. The suite had no guard, so it passed only when some other suite happened to build it first - an execution-order dependency, which is why it failed on node22 and passed on node24 for identical code. install.test.cjs already documents this exact hazard and guards it. Six behaviorally identical copies of that guard existed across three files. Rather than add a seventh, they are now one canonical tests/helpers/hooks-dist.cjs - idempotent and bounded by the shared BUILD_TIMEOUT_MS class norm - which is the same single-sourcing this PR applies to the ADR gate itself. * docs(#2704): add a how-to for contributors the ADR gate rejects The reference and explanation quadrants were covered by Lifecycle rules 5 and 6, but the task-oriented one was thin: a contributor meets this gate because it failed on their PR, under pressure, and the rules told them what is checked without telling them what to do about it. Adds the command to reproduce the CI failure locally and a message-to-remedy table covering every reason code that can be hit - unresolved target, wrong case with the did-you-mean hint, repository escape, symlinked ADR file, bracket contradiction - plus the backtick escape hatch for illustrative links and the caveat that indented code blocks are not skipped. The table is itself written in backticked inline code, so the gate skips it: the escape hatch demonstrated on the page that documents it. * chore(#2704): backfill changeset PR number pr:0 placeholder replaced with the real PR number now that #3266 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
35df0891af |
fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates $HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard is right — this is #2665's class arriving through the one door the scrub set is structurally unable to close. bin/install.js writeNonClaudeDefaults() (#2834) writes path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime. os.homedir() consults NO GSD variable, so: - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is derived; and - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back to exactly that homedir(). Only a sandboxed HOME contains it, and HOME is deliberately excluded from TEST_ENV_BASE because blanking it would break far more than it fixed. So the containment belongs per-spawn, which is the discipline the suite already applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment naming this exact hazard for the --codex spawn, while the parameterized --${runtime} spawn 130 lines below it does not. Instance fixed, class open. Rather than add a fifth hand-synced env shape to a PR whose subject is that hand-synced copies drift, this adds ONE export -- installSpawnEnv() in tests/helpers.cjs -- and routes every raw installer spawn through it, including the shared tests/helpers/install-shared.cjs installerEnv(), which every install suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are unaffected: overrides spread last. Census (measured, not reasoned): of the 119 test files that spawn bin/install.js, exactly four wrote into a sandboxed $HOME before this commit -- install, copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu went red on one file while the macOS full lanes went red on four. Attribution: the leak reproduces unchanged at upstream/next itself, so the defect is base-owned and pre-existing; only the detector is new. The guard found a real leak on next within one run. No new failures: the surviving names under a sandboxed HOME (folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too, and base additionally fails folded:bug-3288-model-catalog-install-path, which this tree does not. |
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |
||
|
|
cbd180c5cd |
test(#3147): bound the lint/changeset/docs cluster onto the process seam (#3181)
* test(#3147): bound the lint/changeset/docs cluster onto the process seam Migrates 69 unbounded sync spawn sites across 24 files. Allowlist 73 to 49. Two shared helpers move: tests/helpers/graphify.cjs (6 importing suites) and tests/fixtures/index.cjs, whose three quoted-argument shell strings became single argv elements rather than whitespace splits. changeset-lint's throw-native git() helper routes to gitOrThrow; migrating it to bare runGit would have silently swallowed a failure that is loud today. ingest-docs goes the other way -- its catch never rethrew, it degraded failure into data every call site asserts on, so throwIfFailed would have thrown where the original returned. The design doc said otherwise and was corrected. tsconfig-noemit runs a real tsc --noEmit and takes a bespoke 180000ms per the ensure-runtime-build precedent, not the 30000ms build-hooks norm -- that norm is for a file copy, and sizing against a label rather than the work is the same error in the opposite direction. * test(#3147): add toLegacyResult and settle review findings The seam exposed a throwing adapter (throwIfFailed) but no non-throwing one, so eight files independently re-derived the same unwrap back to the legacy {status, stdout, stderr} shape. That is the third time this epic produced N copies of one mechanism -- seven throw wrappers in Wave 1, fifty-two timeout constants in Wave 2, eight result adapters here. The pattern is that whenever the seam does not expose a mechanism, every suite re-derives it. toLegacyResult now sits beside throwIfFailed, with its own tests. Two sites are deliberately NOT converted: changeset-cli's runRender and runRenderIn return {status, report, stderr} from parsed JSON and never a raw stdout, so they are a different shape family. lint-legacy-dir-name keeps its local GUARD_TIMEOUT_MS: 30000 matches the build norm numerically but bounds a lint probe, not hooks bundling, and importing it would encode a coincidence as a relationship. --------- Co-authored-by: sim <sim@local> |
||
|
|
3fac6e629f |
test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73. Timeouts are sized from evidence already in the tree rather than a house default, because this wave spawns installers rather than git plumbing and an undersized bound does not catch a hang -- it manufactures CI flake, which is worse, since a flake gets re-run instead of investigated. install.test.cjs records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while another lane passed the same commit in 12.7s, so full installs are bound at 120000ms against that recorded incident. Also adds an auditable escape to the guard's timeout ceiling. The 600000ms cap was set in #3143 from partial evidence, but fragment-single-edit- propagation carries a documented, load-tested 900000ms bound on a run that chains a full build plus eight generators -- the guard would have rejected a correct timeout the moment that file left the allowlist. A value above the ceiling is now permitted only with an inline allow-spawn-timeout-ceiling marker carrying a non-empty reason. It raises the ceiling; it never waives the requirement for a bound, which is asserted directly. install-shared.cjs keeps its hand-rolled assert rather than routing through throwIfFailed: its message embeds both streams, and throwIfFailed carries only a trimmed stderr. The message now also names the outcome, so a bounded timeout reads as such across its 38 importers instead of as expected null to equal 0. * test(#3145): extract class-norm timeouts and correct the build-hooks sizing A pre-PR review found 52 copies of four class-norm timeout constants across this wave. These are not per-suite fixture bindings -- they are shared facts about how long a class of subprocess takes, derived from a recorded bench incident. That norm already moved once (60000 to 120000 after a real ETIMEDOUT), and 52 copies would have drifted the next time it moved. Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and converts the copies. A site that genuinely differs -- a real tsc compile, or regen:derived -- keeps its own local constant with its own justification. Also corrects a misclassification: scripts/build-hooks.js was sized as a build at 120000 in twelve places and 60000 in another, but it compiles and bundles nothing. Its own header says no bundling needed; it copies pre-built files and syntax-checks them with vm. Three different values bounded one script; now there is one. * test(#3145): fix red CI — lint self-match and a Windows chunk overrun Two failures on PR 3176. lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The fixture exists to prove an unrelated marker does NOT suppress the rule, so it carries that marker's literal text as test data. Split via concatenation, the same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own self-match problem. The explanatory comment needed the same treatment. The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped seven minutes before the kill, so this was an overrun rather than a slow chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs regen:derived bounded at 900000ms, which is larger than the whole chunk budget, so the chunk killer always fires first and it can never complete there. Both the test and that bound predate this change; modifying the file pulled it into the Windows targeted set and exposed it. Skipped on Windows with the reason recorded; the Linux lanes cover it. The 900000 bound and its ceiling marker are unchanged -- they are correct. * test(#3145): refresh the stale test-timings cost table The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs packs chunks by measured duration from tests/test-timings.json, and an unknown file falls back to the table's median weight -- advisory by design, but it silently underweights exactly the files that matter. Four of the failing chunk's 22 files were absent from the table, including the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s (it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s. Both were weighted as average, so the chunk's total weight read 53.68 against a budget of 60 and the packer produced a single chunk. Regenerated from a passing full-suite run, per the remedy the script itself documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since gen-test-timings.cjs replaces the table wholesale rather than merging. Proven against the real packer: the same 22 files now weigh 103.91 and split into two chunks. No logic, budget, or timeout was changed; raising a budget to make a red gate pass is not a fix. --------- Co-authored-by: sim <sim@local> |