c516c39b33fa3d298a9c9d73fd342cc4d0682a11
5165 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c516c39b33 |
fix(#3203): stop npm-global installs validating bundled agents against themselves (#3229)
* fix(#3203): stop npm-global installs validating bundled agents against themselves getAgentsDir's claude branch derived the agents directory from __dirname, which is correct for repo runs and runtime-config-dir installs (where <root>/../agents IS the user's agents dir) but on an npm-global install resolves to the package's own bundled agents/, so checkAgentsInstalled validated the package against itself and agents_installed could never be false. The new-project and new-milestone halt/warn gates were silently dead for npm-global users. Keep the install-relative path for the shapes where it is correct, but when it lies inside a node_modules tree — the provably self-validating case — resolve getGlobalConfigDir('claude')/agents like every other runtime, honouring CLAUDE_CONFIG_DIR. GSD_AGENTS_DIR stays priority 1. Repair the doc comment that asserted the __dirname form was correct for both install shapes. Regression test mirrors the published npm-global layout (package under node_modules with a complete bundled agents/) and pins the resolved directory plus the issue's negative control (one agent missing from the config dir → agents_installed:false). Verified red against pre-fix code, green post-fix; the repo-layout W010 health test stays green. * chore(#3203): set changeset fragment pr to 3229 * docs(#3203): describe the node_modules guard as lexical, in CONTEXT.md and at the call site The Agent Install Check Module glossary entry asserted that Claude resolves the agents directory `__dirname`-relative unconditionally. That is the premise this PR falsified: on an npm-global install the install-relative path resolves to the package's own bundled `agents/`, so the check validated the package against itself and `agents_installed` could never be false. The inline doc comment above `getAgentsDir` was repaired with the fix; this external predicate was left behind and has been false since. CONTRIBUTING.md's `Fixed`-fragment docs exemption names this case explicitly — "Edit the docs anyway if a fix corrects something the docs got wrong." Both surfaces now describe the guard as what it is: an exact, case-sensitive path-segment test that TARGETS those layouts rather than detecting them, so neither claims more certainty than the predicate has. The call-site comment carried the same conflation the glossary did. A path merely carrying a directory of that name resolves the same way — the edge already disclosed on this PR — and a non-empty GSD_AGENTS_DIR overrides it. Comment-only in `src/`; no behaviour change. `CONFIG.LOCATION.SEAM.two-families` needs no change: `GSD_AGENTS_DIR -> getAgentsDir priority 1` is still accurate. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3dff700aaa |
fix(#3102): render edge-probe coverage report so the resolution loop consumes it (#3391)
* fix(#3102): render edge-probe coverage report so the resolution loop consumes it Step 5.5 captured the edge-probe report into $COVERAGE, shape-checked it, and reduced it to coverage.applicable — the engine's per-requirement items[] never reached the model, so the resolution loop re-derived edge categories from prose (the data-flow twin of #2733's control-flow discard). The block's own comment claimed the opposite. Render $COVERAGE RAW into context after the well-formedness guard (schema-agnostic so an ADR-550 D7a-style re-cut cannot desync a bespoke renderer), and bind the rows in the resolution loop as a deterministic FLOOR the model unions with its own classification — floor, never ceiling, since the classifier has a measured recall gap (ADR-857 §98 / ADR-550 D7b). --auto consumes the same floor. Comment corrected to match. Regression test asserts a bare render of $COVERAGE, not a count-only cross. * chore(#3102): add changeset for the Step 5.5 edge-coverage render fix * chore(#3102): re-arm spec-phase.md emitted-drift ack for the Step 5.5 render growth The render + floor-binding prose grows spec-phase.md ~1818 bytes (32238 -> 34056, under the 40960 cap). Re-arms the existing spent spec-phase.md ack rather than adding a new fragment (a second key would collide with the base-relative duplicate check). |
||
|
|
bd31cf2bba |
fix(#3321): exclude .claude/.planning from no-phantom-issue-refs SKIP_DIRS (#3396)
* test(#3321): add failing-first regression proving walk() must skip .claude/.planning RED step: SKIP_DIRS does not yet exclude .claude/.planning, so this test is expected to fail until the next commit adds them. * fix(#3321): exclude .claude/.planning from no-phantom-issue-refs SKIP_DIRS walk() previously had no guard against descending into ambient, gitignored .claude/worktrees/** or .planning/** content, so this guard's pass/fail would have depended on the developer's local worktree layout the next time PHANTOM is repopulated. Currently dormant (PHANTOM=[] short-circuits via t.skip before walk() runs) but the gap was real, per #1885 F23. --------- Co-authored-by: sim <sim@local> |
||
|
|
86101ee612 |
fix(#3133): keep global Claude @-references on tilde, not $HOME (#3393)
* test(#3133): global Claude @-references must resolve on tilde, not $HOME Regression for #3133: on a global Claude install, _applyRuntimeRewrites rewrites @~/.claude/... -> @$HOME/.claude/..., but Claude Code does not expand $HOME in @-file references, so the include silently resolves to nothing and the skill loads an empty execution_context. Rows 1/2/6/9 fail RED on next; rows 3-5/7-8 guard #1284 (shell $HOME), local redirect, #3503 (no homedir leak), and non-Claude isolation. * fix(#3133): keep global Claude @-references on tilde, not $HOME On a global Claude install, _applyRuntimeRewrites rewrote every @~/.claude/ include to @$HOME/.claude/ because computePathPrefix returns the $HOME form for shell-context correctness (#1284: ~ does not expand inside double quotes). Claude Code does not expand $HOME in @-file references, so the rewritten include silently resolved to nothing and skills loaded an empty execution_context. Add a normalization pass in the claude case: when the prefix is the $HOME (global) form, restore @$HOME/.claude/ -> @~/.claude/ (the form Claude expands and the shipped tarball uses). Shell contexts keep $HOME; local installs (absolute prefix) are unaffected; computePathPrefix is untouched (all its tests stay green). * test(#3133): fix row 9 assertion — drop end-of-string anchor Row 9's regex used $ (end-of-string) but a realistic @-ref line ends with a newline, so the anchor could not match. The fix under test produces the correct @~/.claude/gsd-core/references/ui-brand.md tail; only the assertion was wrong. Drop the anchor and assert the tail + no @$HOME. * docs(#3133): add changeset * docs(#3133): backfill changeset PR number (3393) --------- Co-authored-by: sim <sim@local> |
||
|
|
bd2c9589cd |
Merge pull request #3392 from open-gsd/chore/3331-no-elapsed-assertion-error
chore(#3331): promote local/no-elapsed-assertion warn->error |
||
|
|
c6e49a5729 |
chore(#3331): fix stale CONTEXT.md predicates left by the no-elapsed-assertion promotion
Standards-axis code review caught 3 stale predicates (CONTEXT.md:470, 536, 541) still describing local/no-elapsed-assertion as warn and citing the superseded epic #1885 (subsumed into #3053 and closed stale). Regenerated the derived CONTEXT-INDEX.json snapshots. |
||
|
|
9bb4752f2a |
chore(#3331): promote local/no-elapsed-assertion warn->error
#3314 delivered the precondition (ADR-456 §(a) reachability-based clock-control rule + deterministic backfill for commands.cts/init.cts/ io.cts). Current tests/**/*.cjs corpus has zero violations, verified via `npx eslint 'tests/**/*.cjs' --rule '{"local/no-elapsed-assertion":"error"}'` before flipping the config, so no fix/delete-and-replace work was needed. |
||
|
|
b399b4a8c4 | Merge pull request #3386 from open-gsd/test/3339-fold-state-model-profile | ||
|
|
2b20b7e2cd |
fix(#3257): preserve full-line frontmatter comments through the parse→reconstruct pair + syncStateFrontmatter (#3387)
* test(#3257: full-line frontmatter comments survive the parse→reconstruct pair AND a mutating state verb parseYamlRegion dropped column-0 # comments and reconstructFrontmatter rebuilt from Object.entries alone, so full-line comments were silently destroyed on every mutating STATE verb. Add failing-first regressions: 3 unit tests for the public pair (comment between keys, leading+trailing, consecutive) and an e2e test running a state verb (state update) on a commented STATE.md — the e2e exercises syncStateFrontmatter's fresh-derivedFm rebuild path, which is the actual loss site the issue is filed against. RED — fails on next; fix follows. * fix(#3257: preserve full-line frontmatter comments through parse→reconstruct AND syncStateFrontmatter Carry column-0 # comments through the frontmatter pair via a Symbol-keyed channel (FULL_LINE_COMMENTS): parseYamlRegion captures ^# lines and attaches them to the next top-level key (leading) or a trailing slot; reconstructFrontmatter re-emits them in place. The Symbol is invisible to Object.entries/keys/JSON, so every existing reader is unchanged; the channel is created only when a comment is seen, so comment-less frontmatter is byte-identical. CRITICAL (isolated review): syncStateFrontmatter rebuilds its target via buildStateFrontmatter (fresh object) + an Object.keys carry-forward, both of which skip the Symbol — so the pair-preserving channel was lost on the very STATE verbs the issue names. Export propagateCommentChannel(source, target) from frontmatter.cts and call it in syncStateFrontmatter before reconstruct, copying the channel onto derivedFm (leading filtered to keys still present so a deleted key's annotation drops with it, trailing preserved). Decision A. * chore(#3257: add changeset fragment * chore(#3257: backfill changeset PR number (#3387) --------- Co-authored-by: sim <sim@local> |
||
|
|
23e26a9de8 |
chore(#3339): extend BUG_FILE_RE to fix-/issue- prefixes, close H3 (#3315)
Widens the identity ratchet in scripts/lint-regression-test-names.cjs from banning only new bug-NNNN-*.test.cjs files to also banning fix-NNNN-*.test.cjs and issue-NNNN-*.test.cjs, per epic #3053's own recorded decision: extend only after the 74-file fix-*/issue-* backlog fully drains, never before. That precondition is now real, not assumed: this branch is rebased onto origin/next with Waves 1-6 (#3341, #3342, #3373, #3376, #3378, #3383) all merged, and a repo-wide check confirms zero tests/fix-*.test.cjs or tests/issue-<N>-*.test.cjs files remain -- only the two known false matches (issue-dedupe.test.cjs, issue-version-gate.test.cjs, feature-named suites with no digits after the prefix) survive the glob, and the widened regex correctly excludes both. Allowlist snapshotted at zero (scripts/lint-regression-test-names.allowlist.json was already [] and stays [] -- confirmed via --update reporting "already in sync"). Updated docs/TESTING-SUITES.md's Regression tests section to name fix-/issue- alongside bug- as banned new-file prefixes. This is the last commit of H3 (#3315)'s 7-wave decomposition. |
||
|
|
c70e736c85 |
fix: eliminate resource-collision race in ship-notes-wedged-pr jq() mock
Unrelated defect surfaced by a real gsd-test failure while verifying #3339 (test/3339-fold-state-model-profile diffs zero files this fix touches -- confirmed via `git diff <this-branch> origin/next -- ...` returning empty for this file, ship.md, and gsd-core/bin/lib/*.cjs). Fixed inline per this repo's no-defer policy for defects found while building, rather than deferred. The test's jq() mock shelled out to a full new Node.js process for every mocked `jq -r .field` call inside the extracted track_shipping bash script. The "exhausts the polling bound" test hits up to 24 cold Node spawns in one spawnSync call (6 polling iterations x 4 jq calls), racing a hardcoded 10s timeout -- under CI contention this tipped over on the linux-node22 lane specifically while linux-node24 had headroom, producing spawnSync's signal-killed `status: null` instead of the process's real exit code, asserted against the expected 0. Replaced the Node-subprocess jq() mock with a pure-shell sed -nE field extractor that never forks a process, eliminating the variable-cost operation rather than just widening the timeout. Verified extraction correctness against sample JSON (head/status/checks/review, including empty-string and non-matching-field cases). Locally verified (scoped per session policy): all 9 tests in the file pass; the previously- marginal test dropped from ~617ms to ~198ms. The underlying track_shipping polling logic in ship.md was confirmed correct and untouched -- this was a test-fixture race, not a product bug. |
||
|
|
4036470ca0 |
fix(#3339): consolidated helper silently overwrote an unrelated module-scope function
The runVerifiedPhaseComplete(args, tmpDir) consolidation in the prior commit hoisted a `function` declaration into a bare (sloppy-mode) block. Annex B legacy hoisting semantics mean a block-scoped function declaration in sloppy mode also reassigns any enclosing var of the same name the instant the block executes -- and this file already had an unrelated module-scope runVerifiedPhaseComplete(args, tmpDir, env) at line 54, used by ~50 other call sites throughout the file. The block ran before any test() body did, so every one of those 50 call sites was silently pointed at the consolidated helper's different phase-matching logic (parseInt-based, loses dotted decimal sub-phase segments) instead of the real one (normalizePhaseToken/phaseTokenFromDirName-based), breaking multi-level decimal phases like 03.2.1. Fixed by wrapping the consolidated helper in a strict-mode IIFE, which disables Annex B hoisting and restores real block scoping. Caught by a genuine gsd-test failure (2 unique failures, both throw- class, both in phase-completion verification-gate behavior) -- root- caused via diff against the pre-consolidation commit and a scoped local node --test run confirming the fix, not asserted from a hunch. No test() count changed. No production code touched. |
||
|
|
c444051bef |
test(#3339): fix orthogonal-review findings — Wave 7 fold
Standards/Spec-axis review + Memtrace graph pass found real issues in the just-folded suites, all fixed here: - tests/phase.test.cjs: the issue explicitly asked to dedupe overlapping fixtures between issue-2945/issue-2949 — the fold preserved both test sets correctly (genuinely distinct code paths, 0 tests dropped, already verified correct) but left each fold block with its own near-identical copy of a runVerifiedPhaseComplete(args, tmpDir) helper, the fixture- level dedup the issue actually asked for. Consolidated into one shared definition without disturbing the file's other, unrelated same-named helper at module scope (would have collided if hoisted directly). Removed two now-unused local runGsdTools destructurings left behind by the consolidation (both blocks already close over the module-scope import at line 26). - tests/review-lane-descriptor.test.cjs: two .find() results dereferenced without a presence guard (same defect class Wave 3/#3335 already found and fixed once in this epic) — added assert.ok() guards matching the repo's established style. - tests/model-resolver.test.cjs: documented the 1-of-80 dropped duplicate test with an inline comment, matching this same wave's host-integration fold's convention of citing drops by exact reference instead of leaving a reviewer to reconstruct the justification via git archaeology. No test() count changed in any file. No production code touched. |
||
|
|
1bc7f7e6b0 |
test(#3339): fold the state/phase/dispatch & model-profile issue-* cluster — Wave 7
Folds 9 legacy issue-*.test.cjs regression files (140 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). LAST of 4 issue-* waves — closes out the 74-file fix-*/issue-* backlog (pending BUG_FILE_RE extension, held for a follow-up commit until Wave 6 is confirmed merged, per the epic's own zero-backlog precondition). - issue-2828-flat-roadmap-total-phases.test.cjs (1) + issue-3204-state- writer-phase-count.test.cjs (21): both target state-document.cjs buildStateFrontmatter via different CLI entrypoints — merged jointly into state-document.test.cjs, 0 dropped. - issue-2945-phase-complete-checkbox-rollback.test.cjs (4) + issue-2949- phase-complete-stage3-sentinel.test.cjs (4): both target phase.cts cmdPhaseComplete; issue explicitly warned of overlap — verified disjoint fixtures/assertions, 0 dropped, merged into phase.test.cjs. - issue-2927-reviewer-lane-overlay-invocation.test.cjs (10) merged into review-lane-descriptor.test.cjs. - issue-2939-dispatch-flatten-maxdepth.test.cjs (9) merged into host-integration.test.cjs, 2 dropped as verified exact duplicates. - issue-2977-frontmatter-bom.test.cjs (5) merged into frontmatter.test.cjs. - issue-2045-third-party-skills-surface.test.cjs (6) merged into capability-loader.test.cjs. - issue-2517-runtime-aware-profiles.test.cjs (80, the largest single fold in the epic) merged into model-resolver.test.cjs, 1 dropped as a verified true duplicate (checked against src/model-resolver.cts logic, not just title similarity). Fixed a genuine eslint irregular-whitespace finding: a literal BOM character embedded in a doc comment (pre-existing content from the original #2977 source, illustrating what a BOM looks like) — replaced with a readable U+FEFF notation. 3 stale doc references found and fixed (docs/adr/2313, 3180, 443). Zero net test-coverage loss. No production code changed. |
||
|
|
4a0b5e26b4 |
test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6 (#3383)
* test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6 Folds 10 legacy issue-*.test.cjs regression files (75 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). Third of 4 issue-* waves. - issue-2701-nul-corrupted-validators.test.cjs (9) + issue-429-comment- text-gate.test.cjs (31, incl. fast-check property tests): both target verify.cjs/validate.cjs via different calling styles (CLI vs direct- require) — merged jointly into verify.test.cjs, 0 dropped. - issue-2762-plan-reviews-chunked.test.cjs (3) merged into plan-phase-drift-guard.test.cjs. - issue-2771-advisor-subagent-type.test.cjs (1) + issue-2772-discuss- phase-text-inconsistencies.test.cjs (6) merged jointly into discuss-phase-power.test.cjs. - issue-498-update-backup-runtime-dir.test.cjs (3, rename basis) + issue-815-update-next-channel.test.cjs (7, merged in): both concern update.md workflow-text contracts, now update-workflow.test.cjs. - issue-498-update-context.test.cjs (13): pure rename to update-context.test.cjs, sole comprehensive suite for its module. - issue-2765-brace-expansion-lockfile.test.cjs (1, rename basis) + issue-3238-js-yaml-lockfile.test.cjs (1, merged in): two distinct CVE regression pins against package-lock.json, now lockfile-cve-audit.test.cjs, shared ROOT/npmLs helpers deduped instead of double-declared. Ratchet upkeep to keep this wave's own gates green: pruned 3 stale allow-test-rule allowlist entries, cited 2 previously-uncited comments that surfaced in the folded content (#3338), tightened the exemption-file ceiling 305 -> 297. Fixed one stale filename reference in production code (src/init.cts) plus one in docs/reference/workflow-fragments.md. Zero net test-coverage loss. No production code BEHAVIOR changed. * test(#3338): fix orthogonal-review findings — Wave 6 fold Standards-axis review found a real structural defect in two files, both the same root cause and both fixed here: - tests/update-workflow.test.cjs: the folded:issue-815-update-next-channel wrapper's closing brace was placed after the file's pre-existing tail instead of before it, making the already-established folded:bug-2470 and folded:bug-3130 wrappers CHILDREN of issue-815's block in the test hierarchy instead of independent siblings — confirmed via an actual node --test run showing the mislabeled TAP nesting. Moved the closing brace to the correct position; all three fold wrappers are now top-level siblings again (verified via node --test, TAP hierarchy correct, 10/10 tests, 5 suites, identical count before and after). - tests/plan-phase-drift-guard.test.cjs: same mistake in the other direction — folded:issue-2762-plan-reviews-chunked was spliced inside the pre-existing folded:bug-2492-context-coverage-gate wrapper instead of after it. Fixed the same way (227/227 tests, 38 suites, identical count before and after). - Reverted unnecessary 815-suffixed local renames (assert815/fs815/etc.) introduced by the fold — the wrapper is genuinely block-scoped once correctly closed, so no collision existed (same class as Wave 4's __foldSetNested finding). No test() count changed in either file. No production code touched. --------- Co-authored-by: sim <sim@local> |
||
|
|
aee83c8e9e |
test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 (#3378)
* test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 Folds 6 legacy issue-*.test.cjs regression files (120 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). Second of 4 issue-* waves. - issue-766-plugin-manifest.test.cjs (50 tests): pure rename (git mv) into plugin-manifest.test.cjs, sole comprehensive suite for its module. - issue-844-manifest-version-sync.test.cjs (14 tests, rename basis) + issue-1855-marketplace-manifest.test.cjs (17 tests, merged in): both target scripts/sync-manifest-versions.cjs via non-overlapping describe blocks (generic engine vs. marketplace.json-specific), now manifest-version-sync.test.cjs, 31 tests, 0 dropped. - issue-498-identity-drift-lint.test.cjs (8 tests, rename basis) + issue-498-package-identity.test.cjs (17 tests, merged in): both target package-identity.cjs's surface from different angles (lint-side drift detection vs. derive/slugify), now package-identity.test.cjs, 25 tests, 0 dropped. - issue-607-cache-lineage.test.cjs (14 tests) merged into the existing gsd-statusline.test.cjs, matching its already-established fold-wrapper convention from prior waves. Learned from Wave 4: fold agents checked src/** (not just docs/ and gsd-core/references/) for stale filename references — found and fixed 5 across 3 ADR docs (766, 2121, 457), zero in src/ this time. Zero net test-coverage loss. No production code changed. * test(#3337): fix orthogonal-review findings — Wave 5 fold Standards-axis review found real issues in the just-folded manifest-version-sync.test.cjs, both fixed here: - The folded:issue-1855-marketplace-manifest block omitted its own test/describe/assert/fs/path/os requires, silently closing over the #844 basis section's outer-scope bindings instead of declaring its own dependencies — inconsistent with every other fold block in this wave. Restored the local requires the original issue-1855 file declared. - Merging two independently-lettered legacy files (each A-F) left duplicate top-level describe() labels (two "A:", two "B:", two "C:"). Renamed the folded-in #1855 labels (A2/B4/C2) to make every top-level label in the file unique. No test() count changed (31). No production code touched. --------- Co-authored-by: sim <sim@local> |
||
|
|
864f76b46f |
fix(#3255): updateTableCell scans for the table carrying the requested column (#3377)
* test(#3255): updateTableCell reaches a later table when the first lacks the column A ## Traceability section holding a summary table (no Status) above the requirement table (with Status) made updateTableCell bind to the first table and return 'unknown column: Status', so requirements.mark-complete left the row at Pending. Add a failing-first regression at the primitive. RED — fails on next; fix follows. * fix(#3255): updateTableCell scans for the table carrying the requested column updateTableCell bound to the FIRST GFM table in the scoped text and returned 'unknown column' if that table lacked the column — so a ## Traceability section holding a phase-summary table above the requirement table never reached the requirement table, and requirements.mark-complete left the row at Pending while reporting table_unmatched (the #2140 silent-divergence class one level deeper, flagged but not closed by the #2245 cross-section scoping fix). Scan candidate table headers and pick the first VALID table (delimiter + matching column count) whose columns include the requested column. Single-table behaviour is byte-identical (the one table carries the column). Error semantics preserved: no table -> 'no table found'; valid table without the column -> 'unknown column'; lone malformed table keeps its specific reason. Also fixes the derived hasRow/doneTable reasoning, which probes via the same call. * chore(#3255): add changeset fragment * chore(#3255): backfill changeset PR number (#3377) --------- Co-authored-by: sim <sim@local> |
||
|
|
ad07f76a31 |
test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 (#3376)
* test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 Folds 10 legacy issue-*.test.cjs regression files (79 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). First of 4 issue-* waves (following the 3 fix-* waves, all merged). - 1 file with no prior target coverage: renamed (git mv) into legacy-cleanup.test.cjs (sole comprehensive suite for that module). - 9 files merged into 6 pre-existing suites: golden-parity-single-source, runtime-artifact-layout-surface, codex-config (4 sources merged jointly in one pass per the issue's own instruction, to catch overlap between the 4 sources themselves, not just against the pre-existing target — zero overlap found, all 20 blocks additive), runtime-config-adapter-registry (1 of 10 source blocks dropped as a proven subset of existing coverage), cline-install, install.test.cjs. Incidental fixes required to keep this wave's own ratchets green: - Fixed a stale ADR doc reference (docs/adr/1235) to a folded-away filename. - scripts/lint-allow-test-rule-refs: pruned 4 stale allowlist entries for renamed/merged-away files, cited 2 previously-uncited allow-test-rule comments that surfaced as "new" only because their file path changed, added 1 fresh allowlist entry for a pre-existing uncited comment that predates this PR, and tightened the exemption-file ceiling 309 -> 305 to match the real post-fold high-water mark. Zero net test-coverage loss. No production code changed. * test(#3336): fix orthogonal-review findings — Wave 4 fold Standards-axis review + Memtrace graph pass found real issues in the just-folded suites, all fixed here: - Standardized the fold-wrapper convention (block-scoped __foldDescribe) across golden-parity-single-source.test.cjs, runtime-artifact-layout- surface.test.cjs, runtime-config-adapter-registry.test.cjs, and cline-install.test.cjs to match the pattern already used by codex-config.test.cjs and install.test.cjs in this same wave (and by earlier folds elsewhere in the epic) — repeats the exact inconsistency Wave 3 (#3335) already fixed once in this epic. - Fixed a stale allowlist entry's alphabetical position (cosmetic, not tool-gated, caught by review anyway). - Fixed two stale test-filename references in PRODUCTION code comments (src/capability-writer.cts, src/runtime-config-adapter-registry.cts) caught by lint-removed-but-needed — a class of stale reference this wave's fold agents didn't check for, since they were scoped to docs/ and gsd-core/references/ only, not src/. First fix attempt wrongly edited the gitignored gsd-core/bin/lib/*.cjs BUILD OUTPUT instead of the tracked .cts source; caught and corrected before commit. - Fixed one remaining stale doc reference in docs/adr/1235 (a prior partial fix in this same wave missed it). No test() count changed in any file. No production code BEHAVIOR changed — comment-only fixes in src/. --------- Co-authored-by: sim <sim@local> |
||
|
|
23e6d49929 |
fix(#3233): no-op state update-progress when the milestone scan finds zero plans (#3375)
* test(#3233): zero plans (0/0) is a no-op; plans-but-none-done still writes 0% cmdStateUpdateProgress mapped 0/0 through clampPercent to 0% and rewrote the shipped Progress record after milestone close. Replace the stale 'handles zero plans gracefully' test (which asserted the buggy percent:0) with a #3233 no-op regression (100% record preserved, updated:false), and add a negative-space guard: plans exist but none done must still write a legitimate 0%. RED — fails on next; fix follows. * fix(#3233): no-op state update-progress when the milestone scan finds zero plans cmdStateUpdateProgress mapped 0/0 through clampPercent to 0% and unconditionally rewrote the body Progress line, so after /gsd-complete-milestone archived the phases (.planning/phases/ empty, scope COMPLETE) a routine update-progress run destroyed the shipped record ([██████████] 100% → [░░░░░░░░░░] 0%). Add an early-return no-op when totalPlans === 0 — mirroring the established scope-withholding no-op (stderr WARNING + {updated:false, reason}) and computeProgressPercent's null-for-empty contract ('nothing to measure' ≠ '0% done'). The legitimate 0% case (plans exist, none summarized) is unaffected: totalPlans > 0 reaches clampPercent(0, N>0) = 0 and writes 0% as before. * test(#3233): unshadow 'Progress field missing' — clear the zero-plans guard The new totalPlans===0 no-op guard fires before the 'Progress field not found' branch, so the existing 'returns error when Progress field missing' test (no phase dirs → 0 plans) was passing for the wrong reason and that branch lost coverage. Give that test a phase dir + PLAN so totalPlans > 0 clears the guard and it reaches the branch it is named for. (Isolated review finding.) * chore(#3233): add changeset fragment * chore(#3233): backfill changeset PR number (#3375) --------- Co-authored-by: sim <sim@local> |
||
|
|
a875372f18 |
test(#3335): fold the workflow-content & phase-lifecycle fix-* cluster — Wave 3 (#3373)
* test(#3335): fold the workflow-content & phase-lifecycle fix-* cluster — Wave 3 Folds 13 legacy fix-*.test.cjs regression files (131 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053): - 6 files with no prior target coverage: renamed (git mv) into new suites (spike-manifest-scoping, ship-note, add-todo, workflow-jq-dependency, resolve-execution-dynamic-routing, clock) - 7 files merged into 5 pre-existing suites (worktree-base-ref x2, model-resolver, phase-locator x2, frontmatter, verification-status), deduplicated against existing coverage Zero net test-coverage loss: every source assertion preserved or verified as a genuine pre-existing duplicate. No production code changed. Last fix-* wave (Wave 1 #3341, Wave 2 #3342 already merged); 4 issue-* waves remain in #3315. * test(#3335): fix orthogonal-review findings — Wave 3 fold Standards-axis review + Memtrace graph pass found real defects in the just-folded suites, all fixed here: - phase-locator.test.cjs: pinned an unseeded fast-check property test (CONTRIBUTING.md determinism requirement), matching the sibling test's seed:7 convention. - phase-locator.test.cjs: added assert.ok() presence guards after 9 data.plans.find() calls that were dereferenced unguarded, inconsistent with 5 sibling tests in the same file that already guard correctly. Latent robustness gap — an omitted plan would throw an opaque TypeError instead of a clear assertion failure. - Standardized the fold-wrapper convention (block-scoped __foldDescribe) across worktree-base-ref.test.cjs, verification-status.test.cjs, and phase-locator.test.cjs to match the pattern already established in frontmatter.test.cjs and model-resolver.test.cjs from earlier folds. - worktree-base-ref.test.cjs: moved a mid-file require to the top-of-file require block. - Eliminated duplicated env-isolation helpers: model-resolver.test.cjs and phase-locator.test.cjs each reimplemented GSD_WORKSTREAM/GSD_PROJECT save-restore independently; factored a shared isolateWorkstreamEnv()/ restoreWorkstreamEnv() into tests/helpers.cjs and pointed both call sites at it. No test() count changed in any file. No production code touched. --------- Co-authored-by: sim <sim@local> |
||
|
|
ae7dc52972 |
fix(#3225): guard W006/W007 + consistency loops with isSentinelPhaseId (#3371)
* test(#3225): sentinel phase dirs no longer trigger W007 / consistency warnings / gaps The W006/W007 (validate health) and the parallel consistency disk↔roadmap and gap-numbering loops never got the isSentinelPhaseId guard that phase.cts has (#2786/#2949), so every sentinel phase dir (999.x/0.x — never-on-roadmap by convention) produced a spurious W007 and a spurious 'Gap in phase numbering: N → 999'. Add failing-first regressions for both surfaces + the gap check, each with a non-sentinel orphan negative-space guard. RED — fails on next; fix follows. * fix(#3225): guard W006/W007 + consistency + gap loops with isSentinelPhaseId cmdValidateHealth's W006/W007 loops, cmdValidateConsistency's parallel disk↔ roadmap loops, AND its gap-numbering check never got the isSentinelPhaseId guard that phase.cts has at 10+ sites (#2786/#2949). So any repo using the sentinel-id convention (999.x backlog/interim, 0.x drafts) got a permanent spurious W007 and a spurious 'Gap in phase numbering: N → 999', with advice to add-to-roadmap (violates the convention) or delete (destroys archived work). Add isSentinelPhaseId to the phaseIdMod destructure and skip sentinel ids in: W006 + W007 (cmdValidateHealth); the two plain-warning disk↔roadmap loops and the gap-numbering integerPhases filter (cmdValidateConsistency — same bug family, folded in inline per no-silent-defer). Additive only: non-sentinel orphans and real numbering gaps still warn. The gap-numbering guard was surfaced by the isolated review (a 999-interim dir would otherwise create a false 'N → 999' gap). Same family as #3167 (since fixed). * chore(#3225): add changeset fragment * chore(#3225): backfill changeset PR number (#3371) --------- Co-authored-by: sim <sim@local> |
||
|
|
62f5f3b39e |
fix(#3224): register WINDOWS.md as a canonical .planning/ artifact (#3369)
* test(#3224): assert WINDOWS.md is a canonical .planning/ artifact The broken-windows ledger (.planning/WINDOWS.md, written by gsd-core's own windows command) was absent from CANONICAL_EXACT, so validate health flagged it W019 'Unrecognized' with advice to delete a file that can gate /gsd-ship. Add WINDOWS.md to the expected-canonical list + a dedicated predicate test. RED — fails on next; fix follows. * fix(#3224): register WINDOWS.md as a canonical .planning/ artifact CANONICAL_EXACT (src/artifacts.cts) was never updated when the broken-windows capability (#1950/#2441) started writing .planning/WINDOWS.md. validate health therefore flagged the ledger W019 'Unrecognized' with fix advice to archive or delete it — a file gsd-core itself produces (src/broken-windows.cts, LEDGER_FILE_NAME) and that can gate /gsd-ship under workflow.windows_enforce. Add WINDOWS.md to CANONICAL_EXACT, per the registry header's own maintenance mandate ('Add entries here whenever a new workflow produces a .planning/ root file'). isCanonicalPlanningFile now returns true for it, suppressing the false W019. Existing W019 behavior for genuinely-unrecognized files is unchanged. * chore(#3224): add changeset fragment * chore(#3224): backfill changeset PR number (#3369) --------- Co-authored-by: sim <sim@local> |
||
|
|
bfd749cb9a |
fix(#3213): segment-boundary membership for letter-named phase dirs (#3368)
* test(#3213): add letter-named phase regression for getMilestonePhaseFilter The custom-ID branch's greedy capture excluded every letter-named phase directory (Phase A:..Phase L:) from the milestone, fabricating counts. Add two failing-first regression tests: a single letter phase + numeric control, and the full A..L + 00 tree from the issue's reproduction. RED — fails on next; fix follows in a separate fix: commit. * fix(#3213): segment-boundary membership for letter-named phase dirs The custom-ID branch of isDirInMilestone (getMilestonePhaseFilter) used a greedy capture ^([A-Za-z][A-Za-z0-9]*(?:-[A-Za-z0-9]+)*) that swallowed the whole hyphenated directory name (A-tool-output-contract was captured as 'A-tool-output-contract', not 'A'). The set lookup then failed and every letter-named phase directory (Phase A:..Phase L: — GSD's own convention, ADR-612 first-class non-numeric IDs) fell out of the milestone, silently fabricating progress/plan counts over whatever numeric dir survived. Replace the capture-then-lookup with a segment-boundary membership test: a directory belongs if its lowercased name EQUALS a declared phase ID or BEGINS with that ID followed by '-' (so 'A-tool-output-contract' matches ID 'a'; 'PROJ-42-description' matches ID 'proj-42'; 'AB-combined' does NOT match 'a'). IDs are sorted longest-first so a hyphenated id (proj-42) is tested before a prefix of it (proj). Additive only — numericRe still handles every leading-digit dir first, and this can only ADMIT a dir the greedy capture wrongly excluded, never exclude one already matched. * chore(#3213): add changeset fragment * chore(#3213): backfill changeset PR number (#3368) --------- Co-authored-by: sim <sim@local> |
||
|
|
e7993d77bf |
fix(#3207): create-and-switch on the first phase/milestone commit (#3363)
* test(#3207): invert fresh-create branching tests to expect create+switch The four fresh-create branching tests (#3079) locked in create-without- switch behavior that #3207 identifies as the regression: a fresh phase/ milestone branch is created but HEAD never moves onto it, so the first phase/milestone-scoped commit lands on the base branch. Invert those assertions to expect create+switch, and add two new tests: - fresh-create is non-silent (logs the create+switch) (#3207 AC3) - a second phase commit does not re-warn once HEAD is on the phase branch (#3207 AC5) The existing-branch path (#2539) is deliberately byte-for-byte unchanged. RED — fails on next; the fix follows in a separate fix: commit. * fix(#3207): create-and-switch on the first phase/milestone commit The #3079 fix (PR #3141) replaced git checkout -b with git branch (create-only) unconditionally, including the case where the strategy branch does not yet exist. That regressed #1278: the first phase- or milestone-scoped commit no longer landed on the strategy branch — it stayed on the base branch, and the strategy branch was left as an empty marker pointing at the pre-phase tip. Every subsequent commit then warned 'already exists; committing on the current branch instead of switching', wording that misleads because the tool itself created the branch moments before. Re-separate the two cases #3079 collapsed: - branch does NOT exist -> create AND switch (git checkout -b). The #3079 resurrection hazard cannot apply: the branch was just verified absent, so there is no merged-and-deleted ref to resurrect and no silent move onto an existing unrelated branch (#2539 AC2 is honored by the existing-branch arm, which is unchanged). The create is logged so the first phase-scoped commit is not silent (#3207 AC3). - branch ALREADY exists -> unchanged: no switch, commit on current branch, non-silent warning (#2539/#3079). Once the first commit switches HEAD onto the strategy branch, the currentBranch !== branchName guard skips the block on subsequent commits, so the misleading 'already exists' warning no longer recurs. * chore(#3207): add changeset fragment * chore(#3207): backfill changeset PR number (#3363) --------- Co-authored-by: sim <sim@local> |
||
|
|
0396d9cab1 |
enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection
The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.
That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.
Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.
review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.
* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines
Two CI failures from the first push, both mine:
1. lint-tests: the new regression test split readFileSync content on a
literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
"\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).
2. golden-install-parity / workflow-size-budget / workflow-compat: editing
gsd-core/workflows/review.md changes its content hash and byte size, and
both are pinned in committed baselines. Regenerated via the repo's own
generators (npm run size:baseline, npm run gen:golden).
The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.
Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).
* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape
The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.
* enhance(#2483): also guard the claude leg against auto-memory injection
CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).
* enhance(#2483): match the claude binary in command position, not argument position
The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as
|
||
|
|
b65d04c044 |
docs(#3353): record #3346 and #3353 triage decisions as out-of-scope (#3362)
* docs(#3346): record new-host-as-in-tree-runtime as out-of-scope New host runtimes go out-of-tree as EoS host-plugins listed in the EoS Registry (ADR-1239), not as first-party in-tree registry entries. Sibling to omp-runtime-in-core.md; on-point precedent #2170 (Devin CLI). * docs(#3353): record commit_gates config as out-of-scope A parallel commit_gates config would duplicate the existing capability gate system (command-exit-zero, #2008/ADR-2008). The route is a commit:pre loop-point reusing the existing gate. Companion defect #3352 stays open. --------- Co-authored-by: sim <sim@local> |
||
|
|
68a199cf5a |
fix(#2783): address wedged PRs in ship note protocol (#2818)
* fix(#2783): address wedged PRs in ship note protocol * chore: acknowledge ship.md growth * fix(#2783): repair ship workflow structure * fix(#2783): gate ship-note recovery on current PR state * fix(#2783): avoid scanner collision in poll loop * fix(#2783): address reviewer feedback on ship-note wedge handling * test: add timeout to spawnSync in ship-notes-wedged-pr.test.cjs to satisfy lint * test: update ghCalls bound expectation in ship-notes-wedged-pr.test.cjs * test: restore ghCalls expected count in ship-notes-wedged-pr.test.cjs |
||
|
|
5e951540af |
fix(#3162): resolve active state phase before drift scan (#3208)
* test(02-01): reproduce template state validation drift - derive command fixtures from the shipped STATE template - pair passed-verification drift with a clean opposite-result control * test(02-01): cover state phase resolution boundaries - exercise precedence conflicts fallbacks and fail-closed directory handling - prove canonical equality and reject outside-root verification evidence * docs: add changeset for PR #3208 * Address review feedback * fix(#3162): preserve phase validation after state refactor * test(#3162): align validation scope cases |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
e87fb409ee |
enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age freshness proxy (state_commits_behind / state_commit_stale) through state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as health W024. The proxy is advisory: classify() deliberately does NOT consume it (ADR-1787 locks the classification/routing boundary — a signal, not a route). Composes with #3099 and #1882 (both merged to next after this branch): the commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE diagnostic reads `last_activity` — two different fields, not "two staleness signals on one field." A new regression test asserts a STATE.md carrying both an unparseable last_activity AND a valid state_head resolves each independently (diagnostic fires once; freshness reads state_head, commits_behind 0). Rebased onto next (flattened): resolved the add/add conflicts in src/smart-entry.cts (kept both the #2573 freshness import/derivation and the #3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B). Tests: smart-entry 62, state/state-transition/health/verify 639, all pass. * chore(#2573): allowlist health-validation test in the prompt-injection scan The scanner's `exec('` code-execution pattern matches the benign `re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests (pre-existing: 16 such calls on next, this PR adds none). The file entered the diff-mode scan's changed-file set only because #2573's W024 state_head assertions touch it. Allowlist it alongside the other test files that carry pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class). Scanner self-test 38/0; diff scan 14 files, 0 findings. |
||
|
|
9341d8b8d3 |
test(#3334): fold the workflow-dispatch & review-lane fix-* cluster — Wave 2 (#3342)
Folds 15 tests/fix-*.test.cjs regression files (191 test() blocks) into their module's main suite, per the wave decomposition of #3315 (H3 of epic #3053). 187 blocks land in 8 existing suites (4 exact-duplicate cases dropped, documented inline); 4 blocks move via git mv into 2 new suite files with no prior coverage to merge into. Zero production behavior change. Also tightens two H1 (#3313) ratchets that the fold's own file-count reduction moved past their grace window, per the ratchets' documented dual failure mode (a stale/too-loose baseline fails exactly like a novel violation): - lint-test-file-count.allowlist.json: removes the stale "audit" entry (folding fix-2766 into tests/uat.test.cjs drops that module back to its 2-file cap). - lint-allow-test-rule-refs.ceiling.json: lowers maxFiles 314 -> 309, the real post-fold high-water mark (gsd-test's own repo-baseline test caught this — CI, not a human, found it). Two orthogonal review passes (Standards+Spec code-review, isolated security-review) found and this commit fixes two issues before push: a genuinely-distinct #2287 test case (file-absent vs. file-present- resolved) that a prior fold pass had wrongly dropped as a duplicate — restored verbatim into tests/uat.test.cjs; and a missing same-line allow-test-rule citation on the #2196 block in tests/debug-session-management.test.cjs, added for consistency with its sibling #2257 block. lint-removed-but-needed also caught two stale doc references to the now-folded-away fix-2285-claude-orchestration-wiring.test.cjs filename (docs/adr/1143-claude-orchestration-capability.md, gsd-core/references/execute-phase-response-language.md) — updated both to point at tests/claude-orchestration.test.cjs, its new home. Co-authored-by: sim <sim@local> |
||
|
|
33fca50d8a |
test(#3333): fold the runtime & install surface fix-* cluster — Wave 1 (#3341)
* test(#3333): fold the runtime & install surface fix-* cluster — Wave 1 Folds 11 legacy tests/fix-*.test.cjs regression files into their module's main test suite: 6 folded into existing suites (host-integration-descriptors, effort-surface-axis, trae-imperative-reference, hermes-skills-migration, gsd-agent-isolation-guard), 5 renamed to become the module's sole suite (cursor-hook-workspace-roots, cursor-subagent-isolation, lint-compiled-artifact-sync, hooks-commonjs-marker, shared-hooks-dir-resolution). All 195 test() blocks preserved with zero drops; lint-test-file-count.cjs and eslint remain clean. No production code changed. Wave 1 of 7 in #3315 (H3 of epic #3053). * test(#3333): replace try/finally with t.after() in isolation-guard tests CONTRIBUTING.md bans try/finally inside test bodies (masks failures, not an approved pattern). The fold in the prior commit carried 27 instances forward verbatim from the deleted fix-3045-dispatch-isolation-resolver.test.cjs into an otherwise-clean file. Converts each to the approved per-test t.after() cleanup pattern — same cleanup call, registered instead of finally-wrapped. No assertion, fixture, or test-name change; test( count unchanged at 50. Found by the Standards review pass on Wave 1 (#3333, H3 of epic #3053). * fix(#3333): restore raw NUL byte mangled by the fold in hermes-skills-migration.test.cjs The prior fold commit copied fix-2284-hermes-agent-delegate-task-projection's "collision-robust" test via a text-based Read/Write pipeline, which silently turned a raw NUL byte (0x00) embedded in two string literals into a regular space character. That corrupted the test's actual purpose (proving a NUL byte survives a string-rewrite operation untouched) and produced a genuine gsd-test failure: `24 !== 1` for `out.split(' ').length`, because splitting on a space finds every space in the sentence instead of the single NUL byte the test meant to isolate. Root-caused by diffing the raw bytes (via `git cat-file blob` + `cat -v`) between the pre-fold source and the folded target — confirmed exactly two bytes differ. Restored via a byte-precise patch (latin1 round-trip) touching only those two lines; test( count and every other byte unchanged. * fix(#3333): use \x00 escape sequence instead of a raw NUL byte in test fixture The prior commit restored a byte-exact raw NUL byte matching the original fix-2284 source, and the production function (applyClaudeCodeBrandSwap) was confirmed correct in a standalone repro. But the same raw byte still failed through gsd-test's remote pipeline. Root cause is upstream of gsd-core: some step in that transfer path does not carry a raw 0x00 byte through untouched. A raw embedded NUL byte was never necessary here — `\x00` as a 4-character escape sequence in the source text produces the identical runtime character (U+0000) without ever putting a raw byte in the tracked file, sidestepping any byte-oriented transfer step. Applied at both call sites (the fixture string and the split() delimiter). No behavior change; test( count unchanged at 76. * fix(#3333): harden copyWithPathReplacement against a source file vanishing mid-copy (TOCTOU) Surfaced by this PR's own gsd-test run: tests/install-minimal-hooks.test.cjs and tests/opencode-command-dir-plural.test.cjs intermittently crashed with ENOENT reading gsd-core/workflows/zzz-e5-drift-fixture.md. Root cause is unrelated to test-file consolidation — tests/planning-prompt-drift.test.cjs writes that fixture directly into the real, shared gsd-core/workflows/ tree (main() hardcodes its scan root to the real repo) and deletes it in t.after(); copyWithPathReplacement's readdirSync-then-read loop has no protection against the listed file vanishing before it gets there, so a concurrently-running install path can crash entirely on what is otherwise a completely benign race. Fixed by skipping (not crashing on) a listed entry that no longer exists by the time the loop reaches it. Added a regression test that deterministically reproduces the race (readdirSync snapshot still lists the file; it is deleted immediately after) and proves both outcomes: no throw, and the vanished entry's destination is never partially written. Per CLAUDE.md's no-defer rule, a defect surfaced while verifying this PR is fixed inline rather than deferred — this overrides one-concern-per-PR. * fix(#3333): fix third NUL-byte-mangled occurrence missed by prior fix passes The fold originally mangled three raw-NUL-byte occurrences to spaces, not two — the earlier byte-restore and escape-sequence commits both only targeted the fixture string and the split() delimiter, missing out.includes('[ ]') a few lines below (should read out.includes('[\x00]')). A remote gsd-test run kept failing on this exact assertion even after both prior fixes, which is what surfaced the miss. Verified via a standalone repro using the file's real (not retyped) fixture content: all six assertions in the collision-robust test now pass. Zero raw NUL bytes remain in the file; test( count unchanged at 76. * chore(#3333): add changeset for the copyWithPathReplacement TOCTOU fix Fixed-type fragment for the production defect fixed inline in this PR (bin/install.js's copyWithPathReplacement). Exempt from docs/ requirements per CONTRIBUTING.md (only Added/Changed/Deprecated/Removed require it). * chore(#3333): backfill changeset PR number (pr:0 -> pr:3341) --------- Co-authored-by: sim <sim@local> |
||
|
|
7a7bf19fc1 |
enhance(#2872): record scope and runtime in the install manifest (#3323)
* enhance(#2872): record scope and runtime in the install manifest gsd-file-manifest.json gains manifestVersion, runtime and scope, and a new read-only Installed Surface Resolver Module reads both install scopes for a runtime in one call -- the first code path in the repo that does. Phase 3 of epic #2866 (ADR-2866). Blocks Phase 4 (#2873), which resolves #2218: the resolver's shadowedBy field is that defect expressed as a value for the first time. It ships computed-and-unread here. Installed-ness is decided by manifest PRESENCE, never by the new fields, so a manifest written by an older GSD stays fully functional and no user needs to reinstall. Recorded runtime/scope are corroboration: a disagreement with the probed config dir is reported as declaredScopeMatchesProbe: false, never silently corrected. readInstallManifest is widened additively -- version/timestamp/mode/files keep their exact names, types and meanings for all four existing callers. manifestVersion is a new field rather than a reinterpretation of version, which holds the package version and is read by the golden-parity fixtures. Stems are derived from the installed manifest's own file keys, the inverse of Phase 2's filename composition, guarded by a fast-check round-trip property plus a kebab-case charset check so a crafted manifest key cannot put a traversal segment, control character or ANSI escape into a trigger that Phase 4 renders back to the user. Also fixes two defects found while working: - bin/install.js hardcoded manifestVersion: 2 while the reader owned MANIFEST_SCHEMA_VERSION = 2. Now single-sourced, with a parity test. - docs/installer-migrations.md documented an install-state schema of five snake_case fields that have never been written; InstallState has only ever been { schemaVersion, appliedMigrations }. Corrected with a dated note. Verification runs on the remote runner. * fix(#2872): fold review findings from three independent engines Standards axis: - convert the manifest-schema suite from a hybrid setup(t) closure to beforeEach/afterEach (CONTRIBUTING.md:319-354 Pattern 1). The hybrid was neither approved pattern and a new test forgetting the call got no warning. - SCOPE_ORDER was declared twice with no parity test -- this repo's recorded generative-fix-divergence class. Give the ordering one owner: install-scope exports it frozen, the layout module and the resolver both import it, and a test locks it against scopeRank so the constant and the ranks cannot drift. - drop the defaultReadManifest passthrough (Middle Man). Spec axis: - add the VOLATILE_FILES exclusion test and source comment the acceptance table promised and did not deliver. gsd-file-manifest.json stays excluded: the new fields are deterministic, but timestamp -- the original reason -- is unchanged. Security axis: - bound the reported manifest runtime at 64 chars, matching the truncatePostureValue convention already used in this subsystem. It reached declaredRuntime unbounded while the adjacent stems were gated by SAFE_STEM; an inconsistent posture on the same attacker-influenceable document. The charset stays ungated on purpose -- declaredRuntimeMatchesProbe needs to see the real value -- so Phase 4 must sanitize before rendering, recorded in the design's Known limits. Both new parity tests were verified to FAIL when the two sides are made to disagree, then pass again on revert. Verification runs on the remote runner. * chore(#2872): backfill changeset pr number to 3323 * fix(#2872): give git fixture construction its own timeout class PR #3323's full test (windows-latest, 22, shard 2/3) failed with gitOrThrow: 'git init' failed -- outcome=timed_out exitCode=null gitOrThrow: 'git commit --allow-empty' failed -- outcome=timed_out from drift-detection.test.cjs's beforeEach, a file this branch never touched. Every other lane passed the same commit, including windows-latest node 24 on all three shards, and next is green. Root cause is a bound sized for the wrong class. DEFAULT_GIT_TIMEOUT_MS is 15000 and its own comment scopes it to plumbing READS -- rev-parse, branch, log -- against an existing repo. createFixture uses it for six sequential repo-CONSTRUCTION spawns: init, three config writes, add -A, commit. init and commit each write dozens of files, and on Windows every spawn is Defender-scanned. Sibling tests in the failing block took 15.6-22.0s against a 15000ms bound. This repo already diagnosed this exact shape once: timeouts.cjs's HOOK_FANOUT_TIMEOUT_MS records PR #3285 failing in the SAME job with the SAME outcome=timed_out exitCode=null signature at the SAME bound while every other lane passed, and concludes 'a bound sized for the wrong class, not a slow machine'. It was fixed by splitting out a heavier class-norm at 60000. Same remedy here: GIT_FIXTURE_TIMEOUT_MS = 60000, 4x the bound that failed and half INSTALL_TIMEOUT_MS. DEFAULT_GIT_TIMEOUT_MS deliberately stays at 15000 -- a blanket raise would stop a genuinely hung plumbing read from surfacing fast. Verified the value reaches the spawn rather than being an ignored option: spawnSync was monkeypatched before requiring the fixture module, and all six git construction calls were captured carrying timeout: 60000. This branch's two new test files shift shard composition, which is how a pre-existing fragility landed in the heaviest shard on the slowest lane. Fixed here rather than deferred, per the no-defer rule. Verification runs on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
80734a9694 |
chore(#3053): clock-seam ADR-456 amendment + backfill — H2 (#3332)
* test(#3314): backfill deterministic clock-seam coverage (failing-first) Replaces loose regex/range assertions with exact-value and boundary tests for the CLI-subprocess and in-process clock-touching call sites identified by H2's audit (epic #3053): cmdCurrentTimestamp, _wsParseRetryAfter (commands.cts), cmdInitManager's is_active gate and cmdInitQuick's quick_id generation (init.cts), and reapStaleTempFiles (io.cts). The CLI-subprocess-pinned tests are expected RED until a follow-up commit routes those call sites through realClock so GSD_TEST_MODE+GSD_NOW_MS can reach them. * refactor(#3314): route CLI-subprocess clock reads through realClock cmdCurrentTimestamp, cmdInitManager's is_active gate, and cmdInitQuick's quick_id generation read Date directly, which the GSD_TEST_MODE+GSD_NOW_MS subprocess pin cannot reach (it only fires inside realClock.now()). Behavior-preserving: realClock.now() falls through to Date.now() whenever GSD_TEST_MODE is unset, which is every real invocation. * docs(#3314): amend ADR-456 with reachability-based clock-control rule ADR-456 §(a) documented one mechanism (injected {clock=Date} + t.mock.timers). Adds the two this repo already relies on: t.mock.timers for in-process direct-Date reads, and the GSD_TEST_MODE+GSD_NOW_MS subprocess pin (routed through realClock) for CLI-spawned code. Updates TESTING-STANDARDS.md's matching passages, which already referenced this issue by number as the no-elapsed-assertion promotion precondition. * fix(#3314): address orthogonal review findings Spec-axis findings: file and link the no-elapsed-assertion promotion follow-up (#3331) instead of leaving TESTING-STANDARDS.md pointing at a dead #1885, and ship the module-by-module audit table in the ADR itself rather than only in a gitignored phase artifact. Standards-axis finding: pin the "hour-old file = not active" test via GSD_TEST_MODE+GSD_NOW_MS for consistency with the reachability rule this PR's own ADR amendment now documents. --------- Co-authored-by: sim <sim@local> |
||
|
|
fae2a0fa8e |
test(#3322): add dedicated secrets.cts test coverage (#3328)
Covers maskSecret's unset triad, the 8-char reveal boundary (length-1/length/length+1), falsy-but-valid inputs (0, false), non-string scalar coercion, and isSecretKey/maskIfSecret wiring. H8 of epic #3053. Closes #3322. Co-authored-by: sim <sim@local> |
||
|
|
8b4545f3c0 |
feat(#3218): the prompt layer asks the CLI for plan counts (#3327)
* feat(#3218): the prompt layer asks the CLI for plan counts Seven sites across four workflows counted plans with ls and wc -l instead of asking the CLI. A shell glob is not scanPhasePlans, so every fix that landed on the owner missed all seven: they counted superseded plans as live, reported zero for the nested plans layout, and missed loosely-named files. The 1762 figure of 30 plans and 24 summaries came from here. phase find is extended rather than a verb added - 3218 is an enhancement whose own checklist says it adds no new command, and CONTRIBUTING makes a new verb a feature needing approved-feature. It gains plan_count and summary_count for the live set and plan_count_all for the physical one, additively; the existing arrays are untouched. Both sets are exposed because the sites need different ones. Amendment 1 names two cases; three of these sites ask a third - did the planner write files to disk - and take the physical set, because a superseded plan is still a file the planner wrote. The progress.md dead route is fixed and was worse than the issue said. It read .plans and .summaries arrays that roadmap.analyze has never emitted, so the fallback always fired, both counts were always zero, and Route 0's resume-incomplete-phase check had never fired at all. The ratchet baseline is empty. Its own stale-entry check makes that self-enforcing. Verified on the remote runner. * test(#3218): acknowledge the workflow growth and update the stale guard The emitted-attribution gate named its own remedy, so it was followed rather than pre-guessed: four workflow files grew between 200 and 770 bytes because each replaced a shell glob with a find-phase call plus its jq extraction. plan-phase grew most - two sites, and it takes the physical count for its did-the-planner- write-files question. progress also carries the Route 0 dead-path fix. plan-phase-drift-guard asserted the literal old ls shape. Updated rather than deleted: what it protects is that a filesystem fallback exists and is reachable, and that is intact. It is not a regression - gsd_run is already load-bearing throughout plan-phase.md long before step 9, so the 9a and 11a fallback never existed to survive gsd_run being unavailable; it guards against the planner subagent's return hanging. Three ack sources collided with the new fragment, which the gate treats as a hard error rather than last-wins. Only the three colliding keys were removed, not the 421 spent entries, and two fragments left entryless were deleted per the convention that an empty fragment signals nothing. Verified on the remote runner. * docs(#3218): document the live and physical plan counts docs/CLI-TOOLS.md gains a find-phase counts section covering plan_count and summary_count for the live set against plan_count_all for the physical one, plus the null-not-zero not-found behavior. The live-versus-physical distinction is spelled out because a caller picking the wrong one gets a plausible number, which is the trap Amendment 1 records. Changeset leads with what a user sees: progress and execute-plan stop counting superseded plans as outstanding, a nested plans layout stops reporting zero, and Route 0 resume routing starts working after never having worked. No how-to. Nothing is enabled and nothing is sequenced - the user runs the same command and the number is simply correct. The one new distinction is field semantics, which is what a reference entry is for. * chore(#3218): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3327 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
bcf7b04864 |
chore(#2896): convert CONTEXT.md prose defect registry into enforced gates (#3325)
* chore(#2896): convert CONTEXT.md prose defect registry into enforced gates Squashes the prior 4-commit sequence and fixes defects found while resuming this branch: 5 orphaned/corrupted DEFECT fragment lines left by an earlier botched edit, 17 "Source of truth: Memtrace `find_symbol`" placeholders that had destroyed real file-path citations, and 3 DEFECT.GENERATIVE-* entries merged into one RULESET.GENERATIVE-FIX predicate (policy, not an unenforced defect) to satisfy the zero DEFECT.<NAME>.<field>= acceptance criterion. Six mechanizable defects get real gates: DEFECT.UNBOUNDED-SUBPROCESS (eslint-rules/require-subprocess-timeout.cjs), DEFECT.CANARY-VERSION-LEAK (scripts/lint-canary-version-leak.cjs + version-gate.yml), DEFECT.CHANGESET-PR-FIELD-DRIFT (findPrFieldDrift in changeset/lint.cjs), DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, DEFECT.REMOVED-BUT-NEEDED, and DEFECT.DEFAULT-FLIP-DOCUMENTATION (new lint scripts, wired into lint:ci). Already-enforced and unenforceable prose entries are deleted; the gate is the record. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#2896): route the new lint tests' subprocess calls through the bounded process-seam helper The 4 new test files for this PR's lint checks called cp.spawnSync/ execFileSync directly with no timeout, tripping this repo's own existing local/no-unbounded-spawn ESLint rule. Route every one through runNode/gitOrThrow (tests/helpers/process-seam.cjs, tests/helpers/git-fixture.cjs) instead, matching the pattern already used elsewhere in the suite (e.g. tests/changeset-lint.test.cjs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: register claude-orchestration.cjs and regenerate stale generated indexes Pre-existing drift on next, unrelated to this PR's own change, surfaced by running lint:ci as part of verifying #2896: two cli_modules (claude-orchestration.cjs, write-set.cjs) landed without a manifest regen, and CONTEXT.md's own edits in this PR staled its two generated indexes. Adds the missing docs/INVENTORY.md row for claude-orchestration.cjs (write-set.cjs already had one — only its manifest entry was stale) and regenerates docs/INVENTORY-MANIFEST.json, docs/CONTEXT-INDEX.json, and examples/dynamic-context-management/CONTEXT-INDEX.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): default-flip-documentation lint's local fallback base was main, not next Found in review: every other base-ref fallback in this repo (see scripts/changeset/lint.cjs's DEFAULT_BASE, #2988) defaults to `next`, the integration branch every PR actually targets — `main` is the release branch. This script's local fallback (used only when GITHUB_BASE_REF is unset, i.e. never in CI, but potentially on a local or direct invocation) diffed against the wrong ref. No test exercised the unset-env-var path, so it shipped unnoticed; every e2e test sets GITHUB_BASE_REF explicitly and is unaffected by this fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): stale eslint comment, overclaiming CONTEXT.md wording, and an incompletely-regenerated manifest Found by the isolated Standards code-review pass: - eslint.config.mjs's require-subprocess-timeout comment said "'warn' for now... flip to 'error' once migrated" while the rule already shipped as 'error' with all 8 sites migrated in the same commit — described a state that never existed. - The CONTEXT.md pointer block claimed the rule's bounded call sites "never throw", but roadmap-upgrade.cts's pre-mutation clean-tree check correctly still throws on failure (it gates a destructive real-run migration; degrading to "assume clean" would risk clobbering uncommitted work) — softened the claim to describe both shapes accurately instead of overclaiming one. - docs/INVENTORY-MANIFEST.json's claude-orchestration.cjs/write-set.cjs entries from the prior "fix: register claude-orchestration.cjs..." commit didn't actually land — re-running the generator now includes them; lint:generated-sync is green. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#2896): backfill changeset pr field with the real PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): normalize buildCorpus file paths to POSIX in lint-removed-but-needed Windows CI caught it: path.relative(root, abs) returns backslash- separated paths on Windows, but findSurvivingReferences's package-lock special case does file.startsWith('.github/workflows') — a forward- slash literal. On Windows the check silently never matched, so tests/removed-but-needed-lint.test.cjs's real-defect-shape fixture got exit 0 instead of the expected exit 1. Normalize at the production source (RULESET.CONTENT-PATH-NORMALIZATION) rather than the test side. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5339dd60e5 |
feat(#3313): allow-test-rule total-file-count ratchet, F17 promotions (#3326)
Extends lint-allow-test-rule-refs.cjs with a second, independent check alongside the existing uncited-citation identity ratchet: the total number of distinct test files carrying any allow-test-rule marker (cited or not) is now checked against a tight ceiling via the previously-unwired assertTightCeiling primitive (allowlist-ratchet.cjs, 0 prior callers). A cited exemption is legitimate under ADR-456 but nothing stopped the raw total from growing forever - this closes that gap without duplicating the file walk (both checks consume one shared walkTestFiles pass). Ceiling introduced at the exact measured high-water mark (314 files, grace 3) rather than a padded estimate, per "budgets may only decrease." Also lands the two F17 pieces (absorbed from the now-closed #1885) that had no precondition: - --max-warnings 0 added to lint/lint:ci - local/no-source-grep promoted warn->error in the scripts/bin/ eslint-rules glob block (already error in the tests/ glob) Both promotions were pre-verified against a zero-warning tree (fresh non-cached eslint run) before flipping, per the maintainer's clean- tree-first decision. Not included: local/no-elapsed-assertion promotion, which stays warn pending #3314 (H2) - 10 of 19 clock-touching src modules have no sanctioned time-control mechanism until ADR-456 is amended there. H1 of epic #3053, absorbing #1885 F17. Co-authored-by: sim <sim@local> |
||
|
|
aceea3ce4a |
refactor(#3217): withhold a percentage when its scope is not complete (#3318)
* wip(#3217): rule-4 scope withholding — parked, two open findings Implemented but NOT shippable. An isolated review found buildStateFrontmatter still hardcodes SCOPE.COMPLETE, so state json reports percent 0 where roadmap analyze, stats and query progress all correctly report null on the same disk state - rule 4 reintroduced at a site this phase claims to close. Also: roadmap analyze emits scope complete beside progress_percent null with nothing explaining it. Parked to build Phase 4 (#3186) first, which is unblocked. Findings recorded in .gsd/phase/refactor-3217-completion-ratio-scoping/60-review.json. * fix(#3217): withhold the sync percentage on a non-complete scope The parked blocker is fixed - buildStateFrontmatter no longer hardcodes SCOPE.COMPLETE, and the prose Progress fallback is gated too, which was a second leak found while tracing the first. roadmap analyze exposes progress_scope so a consumer can tell WHY a percentage is absent from the JSON alone. Then a residual gap was reproduced rather than assumed. cmdStateSync carried the same hardcode behind a written reason claiming it did not reproduce. It did: on a TRUNCATED window and on UNSCOPED row 4, state sync wrote Progress 0 percent to 100 percent while state json, roadmap analyze, stats and query progress all withheld - and it persisted a self-contradictory file, body claiming 100 percent while its own frontmatter correctly omitted percent. The excuse was also wrong. syncRoadmapRaw is already parsed in that function and is exactly what produces a real scope, so there was a scope to pass. Threaded through listMilestonePhaseDirs; a non-complete scope now skips the write with a reason in changes. milestoneBounded stays as the orthogonal 1761 guard for row 5. Second time this epic a does-not-reproduce claim was too generous. Recorded in ADR Amendment 8 as a correction rather than a quiet rewrite. Verified on the remote runner. * test(#3217): give the withholding fixtures a resolvable scope 40 matrix failures, all fixture drift - no code regression. My own hypothesis that this was over-withholding was wrong and is recorded as such: the worry case, a plain ROADMAP with Phase entries and no version heading, resolves to complete exactly as ADR 7.1 says it should. The real causes were two fixture shapes. Most had no ROADMAP.md at all, which is unreadable via a pre-existing graceful path, and asserted a numeric percent. The five vscode, pi-extension, mcp-server and shell-projection failures were that shape - bare temp dirs using progress json as a reachability proxy while asserting typeof percent is number, which under rule 4 is now null. The rest had a version token in a title or heading with no STATE.md milestone pointer to resolve it, which is classification row 4, versioned but unresolved, so withholding is correct per the contract. Verified on the remote runner. * test(#3217): make the LM-tools reachability tests dispatch against their fixture The gsd_progress reachability test was never testing its fixture. invoke() resolves cwd from vscode.workspace.workspaceFolders by design (the real LanguageModelToolInvocationOptions has no cwd field, per the 2103 fix in extension.js), the mock had no workspace at all, and the test passed a cwd option nothing reads - so it dispatched against the repo working directory. Writing a ROADMAP into the temp dir had no effect. Rule 4 only made it visible. Fixed by mocking workspaceFolders. The two siblings in the same file carried the identical dead cwd and were dispatching against the repo too; they were not failing only because their assertions did not touch scope-dependent output. Both now use their own fixture with assertions unchanged - the no-planning fallback paths already satisfy them honestly. Re-scanned the other five reachability files: no further instances. They thread cwd into parameters that genuinely read it, not through an options shape that ignores it. Verified on the remote runner. * chore(#3217): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3318 exists. * ci(#3217): give the coverage merge enough heap for the merged shards The coverage gate OOMed at exit 134. c8 report merges three shard artifacts, roughly 358MB of V8 dumps in coverage/tmp, and died holding their per-file position maps at the ~4GB default heap. Verified as this branch's delta rather than pre-existing: the same job succeeded on next at 14:18, after phases 4 and 5 merged. Both coverage-gate steps get the bump because both re-slice the same merged data. 8192 doubles what failed and leaves headroom on a 16GB ubuntu runner, matching the idiom the shard step already uses at 6144. This is a memory bound, not a change to what is measured. No threshold was touched. The test file was checked for gratuitous subprocess spawning and is already reasonable at 43 spawns, each a distinct fixture-by-surface pairing. Verified on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
3763d96a97 |
docs(#3316): standing counter-test rule for error/fallback branches (#3317)
Strengthens TESTING-STANDARDS.md contract 6 (counter-tests for negative space): a test on an error/fallback branch must assert the specific degraded verdict the branch produces, not merely that the call did not throw. Generalizes the liveness-vs-correctness finding from #3050 and epic #3051 (closed, all 14 enumerated modules drained) into a standing review expectation, now that the one-time enumeration is done. Worked examples are the real pre-fix and post-fix shapes of tests/worktree-safety.test.cjs's resolveWorktreeContext timeout counter-test, verified directly against source. Deliberately not lint-enforced: a pattern scan for fail-open shapes scored 1 true positive against 3 false positives during #3051's own measurement (source: epic #3051 body, Phase 3). Cross-linked from CONTRIBUTING.md's QA Matrix Requirements so reviewers see it where they already apply the matrix. H4 of epic #3053. Co-authored-by: sim <sim@local> |
||
|
|
e201cde73c |
refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4 The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a ROADMAP checkbox is a human annotation with no machine authority. Section 7.4 still carried the OPEN QUESTION and was marked blocked, so the contract said one thing and the tracker another. Recorded per section 7's own rule - a behavior not stated there is not decided, and amending a rule is an ADR amendment rather than a code change with a comment. The decision comment names Phase 4's PR as the carrier of this edit and makes it an acceptance criterion that the text be in the tree before implementation begins, so this lands first, alone, ahead of any code. Also clears the stale blocked-on-2957 row in the guard roster. * refactor(#3186): one shared phase-completion predicate, disk-strict isPhaseComplete in verification.cts becomes the single owner. It calls readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init gated the read on a plan count and synthesized a not_required sentinel, so phase.complete succeeded while init.manager reported incomplete for the same phase. The guard, built and run before scope was fixed per Amendment 3, found 9 re-derivations where the ADR named 3. Four were unnamed, including one in the prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under disk-strict is the divergence itself. Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no machine authority. The overrides in roadmap analyze and init manager are deleted rather than generalized; the user's checkbox stays in ROADMAP.md, only its authority goes. scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded - they answer 'are all plans summarized', which is a different question, and folding them would either over-report completion or invert the dependency direction between Phase 1's owner and this one. Verified on the remote runner. * fix(#3186): close seven review findings and record the missing-verdict rule The isolated review reproduced a write-path regression I introduced: migrating cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase with a fresh passing verification plus a newly-added unsummarized plan reported complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The owner stays right per 7.4 - plan count is not a completion precondition - so the gate is restored at the write site as an explicit composition, mirroring the separate 2648 unexecuted-plan gate cmdPhaseComplete already carries. The spec axis was right that my 0.x-split reasoning was too permissive. The 2957 decision names buildStateFrontmatter as one of the three that must converge, and buildWorkstreamInventory combined a summaries-met local with verification data to decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in a third surface. Both now route through the owner. The raw scanPhasePlans helper stays: it answers are-plans-summarized, which genuinely is a different question. Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so an absent VERIFICATION.md means not complete everywhere. That retires 2645's verifier-disabled tolerance and inverts its Goodhart incentive - deleting the evidence now lowers completion instead of raising it. Guard hardened: block-form count gates and algebraic restatements are caught, and the header now discloses its remaining limits instead of overclaiming. Verified on the remote runner. * fix(#3186): route state sync through the owner and catch bare completed reads The matrix found 52 failures. 51 were fixtures asserting the old semantics: a phase with plans and summaries but no VERIFICATION.md used to count complete and correctly no longer does. Each fixture now carries a passing verification where that is what the test was actually about, rather than having its assertion weakened. The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync destructured scanPhasePlans().completed directly - a bare field read, not a comparison - and used it as a completion verdict, so state sync and state json disagreed on completed_phases for identical disk state. Routed through the owner. Guard gains shape (d): any read of .completed off a scanPhasePlans() result outside plan-scan.cts, in chained, destructured and indirect forms, function scoped with no line window. It cannot tell a summaries-met read from a completion read - that is data flow - so it flags every one and requires a written-reason exemption, which is the same discipline shapes a-c already use. The blind spot is disclosed in the header rather than overclaimed. The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks. Verified on the remote runner. * test(#3186): give the nested-plans sync fixture a passing verification Last 3 matrix failures were one failure echoing up two describe levels. Phase 01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict completed stayed 0 and no Progress change was emitted - correct new behavior, not a regression. Added the passing verification rather than dropping the Progress expectation, so the test still covers what #3257 is about: that a nested plans/ layout is counted and not undercounted. Probe against the built lib confirms Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3. * chore(#3186): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3306 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
96a82bbffb |
enhance(#3245): report the detected host runtime in init (#3307)
* test(#3245): failing-first coverage for host runtime detection in init Locks the behavior epic #2313 Phase 5 must produce before any of it exists: init reports the detected host, explicit GSD_RUNTIME and config runtime still outrank detection, non-Codex sessions are untouched, and nothing is ever written to shared defaults (#2297). * enhance(#3245): report the detected host runtime in init init reported agent_runtime: claude inside a Codex session, and resolved agents_dir to the Claude agents root with agents_installed: true — a spuriously healthy triple. Runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json. Adds a detection rung beneath both explicit sources, in a new pure module. Codex is identified from its own documented session environment (CODEX_SANDBOX / CODEX_SANDBOX_NETWORK_DISABLED), else an explicitly exported CODEX_HOME whose config.toml exists. The default ~/.codex is never probed: that file exists on every machine that has run Codex, so probing it would misreport other runtimes' sessions. resolveRuntime keeps its exact contract and all 71 dependents, including formatGsdSlash command-style emission; only withProjectRoot consumes the new rung. Nothing is written on any path (#2297). Explicit config still wins (#2517). * fix(#3245): make the parity guard real and single-source the marker Four independent review passes found the generative-fix-divergence guard was vacuous: it asserted agreement at the one input where inferPreferredRuntime and detectHostRuntime do not differ, so it could not fail. It now pins the actual divergence point (CODEX_HOME set, config.toml absent) and records that the asymmetry is deliberate. The config.toml marker is now single-sourced from update-context.cts and imported, rather than carried independently by two surfaces. tests/helpers.cjs now scrubs CODEX_SANDBOX and CODEX_SANDBOX_NETWORK_DISABLED: GSD reads them, so an ambient Codex session would otherwise make the non-codex control test fail non-deterministically. Also: detection is throw-safe end to end rather than only around the fs probe; the Windows-join test is replaced with one that can actually fail (trailing-separator, catches hand-rolled concatenation); the #2297 no-write proof now wraps resolveReportedRuntime, the function that ships, across all three ladder outcomes. * chore(#3245): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
57437071e7 |
test(#3244): codex smoke test across the emit, validate and repair seam (#3298)
* test(#3244): end-to-end codex smoke test across the emit and validate seam Test-only. The only place in this epic where the emitter (Phase 1) and the validator (Phase 2) meet the same bytes. Every phase so far tested its own half against fixtures it authored, and per #2371 a fixture written by the gate's own author can only confirm what that author already believed. Rows 7 and 10 feed a REAL emitted tree to Phase 2's checkCodexModelPosture: if the emitter and the validator disagree about what "clean" means, nothing else in the suite can see it. Scope corrected from the epic, which scoped this to model_profile: inherit. That profile is the one this epic does NOT change — readGsdRuntimeProfileResolver returns null for it, so a Codex install under inherit omitted the model before Phase 1 too, and a test scoped only to it would pass identically before and after the change it exists to prove. Row 1 uses `balanced`, the default and the path that actually lost its pin; inherit is row 4, the unchanged control. Honest classification, because this is a smoke test and claiming otherwise would be false: NO row here is red-first. Phases 1 and 2 are merged and working, so every row passes today. Rows 1-6 and 9 are regression guards; rows 7 and 10 are cross-phase integration. Their value is that nothing else can catch the drift they cover, not that they are red now. The RED checkpoint is therefore skipped deliberately for this phase rather than spent proving a tautology. Rows 8 and 11 — the same cross-checks against Phase 3's repairer — are deliberately ABSENT, not forgotten. Phase 3 (#3243) is still in CI and its module does not exist on next yet. They land once it merges; writing them against a surface that does not exist would have meant inventing the assertion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3244): close the emit-validate-repair loop with rows 8 and 11 These were deliberately absent from the previous commit because Phase 3 had not merged; writing them against a module that did not exist would have meant inventing the assertion. It has merged, so they land now. Row 8: a freshly-installed tree reports ZERO changes from the sync's dry run. If a brand-new install needs repair, the emitter and the repairer disagree about what a correct file looks like. Row 11 is the sharpest row in the phase: after installing with an explicit real-Codex model_overrides pin, the sync reports that agent skipped rather than synced, and the file is byte-identical afterwards. An over-eager stripper that removes any `model` line passes every Phase 3 unit test and fails only here. Byte-identity is asserted by comparing file CONTENTS, not mtime — a rewrite with identical bytes still moves mtime, so an mtime check would pass exactly the implementation this row exists to catch. With these, all four cross-phase rows are in place, and this suite is the only place in the epic where the emitter, the validator and the repairer meet the same bytes. Every phase before this tested its own half against fixtures it authored, and per #2371 those can only confirm what their author already believed. Both rows are cross-phase integration and pass today, like the rest of this suite. No row here is red-first and none is described as such. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3244): extract withAgentsDir instead of try/finally in test bodies Review finding. Rows 7 and 10 put try/finally directly inside the test callback to save and restore GSD_AGENTS_DIR. CONTRIBUTING prohibits that in test bodies — it masks failures — and permits it only inside standalone helpers. The phase's own test matrix restated the rule and it was violated anyway. lint:ci does not catch this; it is a prose standard, which is precisely why it survived to review. It was also inconsistent with this file's own conventions: runCodexInstall, runEffortSyncDryRun and captureStderr already factor save/restore into helpers. withAgentsDir now follows the same shape and both sites use it. No assertion changed — both rows still drive the real checkCodexModelPosture against the real installed tree, which is the whole point of them. Swept the rest of the file: the only other try blocks are inside the two pre-existing helpers and were already compliant. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d28ab7c8f7 |
enhance(#3243): sync installed codex .toml model/effort to the passive posture (#3296)
* feat(#3243): sync installed codex .toml model/effort to the passive posture Implements ADR-2313 D7, and owns the Codex .toml typed IR that Phase 1's review assigned to this phase. The IR exists for a structural reason, not tidiness: this phase has to PARSE these files, and a parser kept bug-compatible with a separate renderer is the generative-fix-divergence shape this epic already dealt with once for the model predicate. So Phase 2's parsing MOVES here rather than being copied — agent-install-check now imports it, and its test file passing unchanged is the proof the extraction altered nothing. The load-bearing property is byte-identical round-trip: render(parse(x)) === x. Without it a sync silently reformats a user's file — line endings, key order, BOM, trailing newline — turning a two-line repair into a whole-file diff in their dotfile repo. The IR keeps original lines and removes targeted ones rather than reconstructing from parsed fields, which is what makes that property hold. It also reconciles a real contradiction between Phase 2 and ADR-2313. An unterminated developer_instructions block: the reader excludes the rest of the file, deliberately failing toward a false positive, because misreading prose as a pin only wastes a user's time. The writer must refuse, because proceeding on a malformed document rewrites it. A false positive is the safe direction for a reader and the dangerous one for a writer. So the parse reports the fact and the two consumers branch on it — one parse, one truth, two policies, instead of two parsers that agree today. The sync leaves a legal real-Codex pin and its coupled effort untouched, reported skipped rather than synced; strips a stale Anthropic or tier model and an orphaned effort; keeps dry-run as the default; refuses any file whose parse fails; and skips symlinks exactly as the Claude path already did. The Claude path itself is byte-identical. PARSE_REASON.NO_HEADER from the ADR's illustrative snippet is deliberately not implemented — a missing header is legal, not an error, so it would be a dead enum member that the enum-lock test then pins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve per-line endings and make the codex write atomic Two findings from an isolated review, both in the write path. BLOCKER: mixed line endings broke the byte-identical round-trip. `eol` was a single whole-file flag and split(/\r?\n/) discarded each line's own terminator, so render re-joined with ONE style and normalized every line — even with zero strips performed. A file with one CRLF line and the rest LF came back fully converted. That falsified the A14 guarantee, violated the design's "must not silently rewrite every line", and made the CONTEXT.md glossary claim wrong. It was untested because A12 and B15 only cover PURE CRLF; no mixed-ending fixture existed anywhere. Fixed by keeping each line's terminator alongside its content, so render is a plain concatenation and a strip removes only the target line and its own terminator. `eol` survives as informational metadata that render never reads. Seven fixtures added for the paths nothing exercised: mixed endings unmodified and with a strip, a lone \r, a file ending on the block's closing ''' with no newline, multiple trailing newlines, a BOM-only file, and an empty file. MINOR, but it contradicted this phase's own contract: the write was in-place open-truncate, so a failure between truncate and completion leaves a truncated .toml — exactly what ADR-2313 says must never happen. The Codex path now writes a sibling temp file and renames over the target, which is atomic on one filesystem, with cleanup on failure. It uses the repo's existing retryRenameSync rather than a hand-rolled rename, and deliberately NOT platformWriteSync, whose normalizeContent would mangle the very CRLF and trailing-newline bytes the round-trip property exists to preserve. The Claude path keeps its in-place write untouched. It has the same shape, but changing it is not this phase's business and its tests must stay byte-identical. B20 previously mocked writeFileSync to throw BEFORE touching anything, so it proved nothing about a mid-write failure — its passing comment was true only because of how the mock was built. It now performs a real truncated write wherever writeFileSync is called, catching both the naive direct-to-target path and the new temp path, and asserts the target is byte-identical afterwards with no stray temp file left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve the trailing-newline state when stripping a last line Caught by B17, one of this phase's own tests — the suite working, not a test problem. Content is reconstructed as the concatenation of lines[k] + terminators[k], so a file with no trailing newline has '' as its last terminator. removeLine spliced out both arrays at the same index, which is right for a middle line but wrong for the last one: it dropped the empty terminator and left the PREVIOUS line's newline in place. A file ending `...\nmodel = "sonnet"` with no trailing newline came back as `...\n`, gaining a newline the user never wrote. The new last line now inherits the removed line's terminator, so a removal leaves the file exactly as if that line had never been written. Removing the only line yields an empty file rather than a stray terminator. Both stripModel and stripReasoningEffort funnel through the one removeLine, confirmed rather than assumed, so a single fix covers both — including the row-B7 shape where a stale model and its orphaned effort are removed in sequence and the second removal targets the last line. Two of the four new cases are honestly not red-first and say so in their comments: removing a last line that HAS a trailing newline only exposes the bug under mixed EOL, since uniform files coincidentally have equal terminators on both sides; and removing the only line already degenerated correctly through Array.slice. They are kept as guards for the new branch rather than dressed up as catches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3243): document the codex repair path and close the loop How-to: the Codex-400 entry added in Phase 2 told users to re-run the installer, because that was the only repair available then. It now leads with `effort sync` and keeps the reinstall as the alternative, with the reason to prefer one — a reinstall regenerates the agent files wholesale, so anyone who hand-edited theirs loses those edits. Detect, preview, apply is now one continuous path in one place. Reference: docs/COMMANDS.md had no `effort sync` entry at all — the same gap `validate agents` had in Phase 2, found the same way. The entry documents BOTH runtimes, because the command genuinely forks on runtime and describing only the new half would misdescribe it. The write flag is `--apply`. The design doc and test matrix both said `--no-dry-run` throughout, which does not exist — verified against the actual arg parser in gsd-tools.cjs before writing. Documenting a flag that does not exist is worse than documenting nothing, because it fails at the moment someone needs it. Both surfaces state that only the targeted lines are removed and every other byte is preserved. That is a user-visible guarantee rather than an implementation note: it is the difference between a two-line diff and a reformatted file in someone's dotfile repo, it is what the IR's round-trip property exists to deliver, and writing it down makes it a contract a future change has to break knowingly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): inherit the trailing-newline state, not the line ending style My previous rule was subtly wrong and this phase's own test caught it. "The new last line inherits the removed line's terminator" copies the removed line's STYLE as well as its presence. A26 uses mixed endings on purpose — line one terminated \r\n, the model line terminated \n — so inheriting silently rewrote line one's ending to \n. That is precisely the defect class the mixed-EOL blocker fix existed to eliminate, reintroduced one layer down by the fix for it. The correct rule inherits the EMPTINESS only. If the removed line had no terminator, the new last line loses its own, preserving "this file has no trailing newline". Otherwise the new last line keeps its own terminator: it is already a newline, and already the right style for that line. A26's assertion moved too, and that deserves saying plainly rather than burying: it previously encoded my wrong rule. Changing a test to match the implementation is usually the mistake, so it was checked from first principles instead — a file whose first line ends \r\n and whose last line ends \n, with that last line removed entirely, must be the first line with its own \r\n intact. The new expectation is what the user's file should actually look like; the old one was wrong. A29 adds the interaction nothing covered: the compounding case (strip a stale model, then its orphaned effort, the second removal landing on the last line) with non-uniform endings either side. The two fixes meet there and nothing exercised the meeting point. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): drop the phantom trailing line from the IR representation Root cause, not another patch on the removal rule. Three consecutive fixes there each surfaced the next issue, which was the signal that the data model was wrong. splitPreservingTerminators left a phantom empty final entry for any file ending in a newline: "a\nb\n" became lines ['a','b','']. So for the common case the real last content line was NOT the last array element, removeLine's isLastLine check never matched it, and every rule I gave was reasoning about the wrong element. What hid it: render was already a plain concatenation, so a phantom empty line with an empty terminator contributes nothing to the output. A14's byte-identical round-trip could never have caught it — the defect is byte-neutral until a removal shifts the index arithmetic under it. That is worth recording, because "the round-trip test is green" was exactly the reassurance that kept the search pointed elsewhere. The representation is now 1:1 — terminators[i] follows lines[i] and may be '' — with no phantom, verified across empty, no-trailing-newline, trailing-newline, blank-line and mixed-CRLF inputs. render stays a plain concat and needs no special cases. With the phantom gone the removal rule is correct as stated and finally applies to the genuinely last element. Consumers checked rather than assumed: the block-range detector and header scanner are agnostic to array shape, and Phase 2's reader uses its own independent split, so tests/agent-install-check.test.cjs is untouched and still passes unchanged. One test expectation was wrong and is corrected rather than quietly adjusted: A18 asserted a 7-element terminators array whose trailing '' was the phantom itself. It now asserts the six real terminators, which is what the invariant lines.length === terminators.length requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3243): backfill changeset pr number (#3296) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c28134ab39 |
fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by |
||
|
|
95d0da9060 |
docs(#3287): add ADR-3180 decision 8, the diagnostic contract (#3292)
Co-authored-by: sim <sim@local> |
||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
693f12ad56 |
refactor(#3187): give state field extraction one canonical owner (#3283)
* refactor(#3187): give state field extraction one canonical owner stateFieldValue in state-document.cts becomes the single owner of the #1760 frontmatter-then-body fallback chain. The new whole-repo guard found 14 independent re-derivations where the epic scoped 5, all now routed through it: cmdStateSnapshot (11), cmdStatePrune (2) and smart-entry fmScalar (1). state validate was a gate that could not fail. Every warning it could emit sat behind a phase resolved without the frontmatter tier, so a STATE.md whose phase lives only in frontmatter skipped the drift scan entirely and returned valid:true. It also read unstripped content, letting a frontmatter status: key shadow the body field (#1255 class). Both fixed; output gains a scope field so could-not-look stops being output-identical to looked-and-clean. Verified on the remote runner. * docs(#3187): document the state validate scope field and its reason codes Adds docs/how-to/interpret-state-validate-results.md so a reader can tell nothing-to-report from could-not-look, updates the COMMANDS.md and USER-GUIDE.md entries, corrects the CONTEXT.md glossary overstatement about Current Position sole ownership, and drops the changeset fragment. * fix(#3187): close three drift-guard evasion shapes and test the refuse path The isolated adversarial review found the ladder detector was evadable by ordinary reformatting, not just deliberately: a member or computed operand (fm.key / fm[key]) missed the bare-identifier backreference, a swapped tier order missed a hardcoded number-then-boolean sequence, and a ladder wrapped across lines missed single-line detection. All three now caught, each with its own test plus a proven boundary control. The frontmatter-parse refuse path on the destructive complete-phase route was unreachable and therefore untested. It is now driven by an injected parse failure and asserts STATE.md is byte-identical after the refusal, rather than shipping untested defensive code on a path that rewrites user state. Verified on the remote runner. * fix(#3187): widen the drift guard to the prompt layer and disclose tier-2 changes The code-review spec axis found the guard's scan surface was src/ only, which is Decision 4(d)'s forbidden allowlist one directory wide - and it had a live miss: gsd-core/workflows/smart-entry.md tells an agent to read status from frontmatter or the body, a prose expression of this same chain. The surface now covers the prompt layer. That one site carries a permanent written exemption rather than a ratchet: it is the gsd-tools-is-down fallback, so it cannot call the owner by construction, and a ratchet would imply removable debt that does not exist. Two tier-2 output changes shipped undisclosed and are now named in the changeset and docs: complete-phase's idempotency guard consulting frontmatter, and the workstream inventory resolving frontmatter-only fields. docs/COMMANDS.md gains a state complete-phase entry, which it never had. Also records Amendment 5 on ADR-3180, extracts the duplicated frontmatter-parse block the epic's own thesis forbids, and re-points two assertions from free-form warning prose onto the structured drift object. Verified on the remote runner. * chore(#3187): backfill changeset PR number pr:0 placeholder replaced with the real PR number now that #3283 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
cf6de5e1c0 |
feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution 23 tests over the 50-test-matrix rows. RED by construction: resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet, and the validator silently ignores triggerPrecedence today. Written in the per-runtime describe idiom the other four runtime-artifact-layout suites use, not a table. The rows that carry the weight: windsurf must NOT report a shadow it does not have, since its global scope emits only agents and agents are not trigger-bearing; agents and kimi-agents must be absent from the output for every runtime; and reordering a runtime's triggerPrecedence must flip the winner, which is the only assertion that proves the axis is read rather than decorative. Stems are injected, never scanned, so the surface is assertable with no filesystem. * feat(#2871): resolve triggers and host precedence, not just placement resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger a runtime emits, with the scope and kind that produced it, whether the host registers it directly or only through a router, and which artifact shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers need placement only and the issue requires them unchanged. AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The host-integration matrix models command and dispatch as separate interface points: an agent is invoked through the Agent tool's subagent_type, not by typing a slash trigger, and _copyStaged never applies the kind prefix to an agents entry. So agents and kimi-agents are excluded from the surface entirely, and this commit amends ADR-2866 with a dated correction. #2218's conclusion is unchanged -- the collision is strictly commands-vs-skills, and claude's local /gsd-* trigger surface is still fully shadowed -- but the ADR implied the local agents surface was lost too, and it is not. That correction is what makes windsurf come out right. Its global scope emits only agents, so it has no global trigger and its local commands are unshadowed. Model agents as trigger-bearing and windsurf falsely reports a full shadow. The triggerPrecedence axis lands on all 19 descriptors as an ordered kind list, one value with one owner, rather than a numeric rank spread across N kind entries with nothing keeping them consistent. Validation uses a required-with-default shape that has no precedent in this validator -- every existing axis is hard-required -- so a third-party capability.json omitting the field still validates, which is what ADR-894's additive-only contract promises. Winner resolution reads Phase 1's scope rank first, then the kind ordering. A test reorders the axis and asserts the winner flips, since an axis that is added, validated and never consulted would pass every other assertion. shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this issue's out-of-scope note. Verified via the remote runner. * fix(#2871): single-source namespacedByDir and close two test gaps Four findings from the isolated adversarial review. The namespacedByDir rule had reached three copies -- install-engine, surface, and the new trigger resolver -- one of which carried a hand-written keep-in-sync comment and no assertion. That is this repo's generative-fix-divergence class. Extracted to one exported predicate all three now call. Verified by diverging one copy deliberately: the existing #816 parity test failed, and passes again on revert. The omission test was vacuous. Row 16 asserted that a descriptor without triggerPrecedence still validates, but built its fixture from claude's shipped descriptor -- which this PR had just added the axis to. It now clones and deletes the key, following the shippedDescriptorWithout pattern, and asserts both that validation passes and that the resolver still picks the right winner from the default. The second half is what makes it prove anything. resolveTriggerSurface silently dropped an unrecognized scope while every sibling in this epic throws. Two phases of one epic should not disagree about whether an invalid scope is an error, so it now rejects through the same shared validator; an empty scope list still returns empty rather than throwing. The ADR amendment had been spliced into the middle of the References list, orphaning its last bullet. Moved to the top, after the header block, which is where ADR-3660 and ADR-1016 both put dated amendments. No lint checks markdown structure, so this was green while malformed. * fix(#2871): single-source the command filename composition too The earlier fix shared the namespacedByDir boolean but left the filename composition around it written twice -- once in _copyStaged as what actually gets written, once in resolveTriggerSurface as what gets predicted. The predictor could go stale silently. One exported helper now composes it for both. The entry.name asymmetry that looked like it would block extraction does not: entry.name is filtered to end in .md and stem is entry.name minus those three characters, so the two branches are the same string by construction. Divergence proven to fail: injecting a marker into the helper broke the trigger-surface suite; reverting restored 25/25. The four sibling layout suites hold at 227 unchanged. * docs(#2871): correct the ADR timing notes that this phase makes stale The Amended by back-links on ADR-3660 and ADR-1016 were written in Phase 0, when the widenings they describe had not shipped. Each carried a forward-looking clause -- "the module changes at Phase 2, not before, until then this module resolves placement only" -- which becomes false the moment this PR merges. ADR-2866's own Amends header and its reciprocal-notes section carried the same tense. All four now describe what shipped. This is a tense and status correction on Accepted ADRs, not a change to any decision. Worth stating because it is the failure mode this epic keeps meeting: gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI would have caught either the missing back-link in Phase 0 or these stale clauses now. They stay correct only because someone checks. * chore(#2871): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
4a1ed2531f |
enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check Specifies ADR-2313 D6 before the implementation exists, so the tests bind to the contract rather than to whatever the code happens to do. RED is established by construction, not by a remote run: checkCodexModelPosture and POSTURE_REASON are absent from the compiled lib today, so every row fails on the missing export. A remote checkpoint here would prove only that the function is missing, which is already known — so the run is deliberately deferred to the combined green checkpoint rather than spent proving a tautology. That makes the NEGATIVE PROOFS the rows that carry real signal. Every positive row passes even for a naive implementation that greps /model\s*=/ over the whole file. Six rows fail it: light-tier service_tier/model_verbosity decoupling (#774), hand-added keys, a commented pin, the model_verbosity key-prefix collision, the runtime no-op ordering, and the headline case — a literal `model = "sonnet"` inside the developer_instructions ''' block, which the emitter fills with agent prompts that discuss models constantly. Row 14's fixture was verified to discriminate before being written: a whole-file scan matches it and a header-slice scan does not. Without that check the test would pass trivially and prove nothing, which is the vacuous-test failure this epic has already hit repeatedly. Adversarial TOML fixtures are hand-authored against the real Codex shape rather than generated by generateCodexAgentToml, per #2371 — a fixture from the writer can only confirm what the writer already believed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3242): validate codex .toml model posture, not just presence Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export, not a branch inside checkAgentsInstalled — that function carries 33 upstream dependents, cyclomatic 25, and sits in two traced process flows, so it is deliberately left untouched. It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf. That is what Phase 1's constant move bought: agent-install-check is documented as pure read/verify and imports only leaves, so reaching the rule through model-resolver would have dragged config-loader into it. Reads liberally, judges strictly, and never guesses. Tolerates comments, key order, whitespace, CRLF, and a BOM; anchors on full key names so model_verbosity does not satisfy a `model` probe; treats extra hand-added keys as none of its business, since the check is a predicate on the two fields the posture owns rather than a whitelist over the document. An unreadable file becomes a named violation and the loop keeps going. The scan covers only the header slice — the lines before the developer_instructions ''' marker. The emitter writes agent prompts into that block and GSD's prompts discuss models constantly, so a whole-file scan reports violations for prose. This is the highest-risk defect in the phase and the reason its fixture was verified to discriminate before being written. The non-codex short-circuit runs before any filesystem call, so a stray .toml under another runtime is never inspected. Wired through cmdValidateAgents as an additive codex_posture key, so a violating install is visible from a command a user actually runs rather than only from a library nothing calls. Also fixes a test defect found while implementing: .gitattributes forces `* text=auto eol=lf` repo-wide, so the committed CRLF fixture was normalized to LF in the index — `git ls-files --eol` reported `i/lf w/crlf`, the working copy being stale pre-normalization bytes. The CRLF row was asserting against a file that could not survive a fresh clone. CRLF is now derived at runtime, which puts it under the test's control rather than git's, instead of adding a .gitattributes exception that fights a deliberate repo-wide policy and that anyone could re-normalize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3242): document the codex posture check where a user will look Three quadrants, filed by where the reader actually arrives. How-to (recover-and-troubleshoot.md, under Install and update problems) is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400 about an unsupported model" — and opens with the verbatim error string. Someone hitting this does not know the words "posture" or "ADR-2313"; they have a 400 in their terminal and will search for that. Reference (COMMANDS.md) had no `validate agents` entry at all, though sibling gsd-tools subcommands are documented. Adding user-visible output to an undocumented command and then linking to it from the new how-to would have left a dangling reference. The entry carries the violation-reason table, since the frozen POSTURE_REASON enum is the machine-readable contract a reader needs rather than the prose. Both surfaces state that presence and posture are separate verdicts — a missing agent lands in `missing`, never as a posture violation. That is a deliberate design decision and would otherwise be invisible to someone watching one command emit both. Explanation stays in ADR-2313, which already covers D6 and the liberal-parse/strict-judge boundary. Pointing at it beats duplicating it into COMMANDS.md and creating two copies to drift. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): close two false negatives in the posture scan Both found by an isolated reviewer and reproduced before fixing. Both made the check report clean when it was not — the worst direction for this function, since the how-to tells users an empty violations list means the install is posture-clean. A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML, but the key pattern required a bare identifier, so the pin was silently invisible. Bare, "double" and 'single' quoted forms now normalize to the same key name. The block marker was found by unanchored whole-content search and used to truncate the header. A `description` value merely containing the literal text `developer_instructions = '''` truncated the scan before a real pin, and a user who hand-reordered `model` to sit after the block — still legal TOML — was never scanned at all. Fixed by changing the strategy rather than the regex: find the block's range, anchored at line start, and scan every line OUTSIDE it. That covers both failures and is strictly more correct than truncation, while still never reading prompt prose. An unterminated block excludes the rest of the file, which fails toward a false positive — the safe direction, since misreading prose as a pin wastes a user's time while the alternative hides a real one. Also corrects two overclaims of mine. The how-to named "v1.11", a version that does not exist — package.json is 1.10.0 and unreleased — so it now describes the boundary by behavior and links the ADR. And the test matrix asserted that a naive whole-file scan "fails exactly rows 12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16 produce the correct result against that baseline too. They guard real but *different* mistakes, and the matrix now says which one each catches instead of attributing them all to the header-slice defect. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): skip symlinked agent files instead of following them Security review finding. The scan listed entries with readdirSync and read them with readFileSync, which follows symlinks — so a symlink in the agents directory pointing anywhere would have its contents read, and any line matching the model pattern echoed into cmdValidateAgents' output through the `value` field. A read-and-echo primitive on an arbitrary path. It needs write access to the agents directory, so it crosses no new trust boundary today. Fixed anyway, for two reasons. This repo already does it correctly next door: cmdEffortSync filters with lstatSync().isFile() and the comment "Skip symlinks — only write regular files to avoid clobbering symlink targets." Being inconsistent with a sibling in the same subsystem IS the defect. And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these files. Establishing symlink-following as the house pattern for Codex .toml handling here would hand Phase 3 a worse starting point while it writes rather than reads. Skipped silently rather than reported, matching the sibling: a symlinked agent file is a structural install choice, which checkAgentsInstalled owns, not a model-content posture defect. An lstat that itself throws excludes the file rather than crashing the scan. That does narrow the guarantee slightly, so the how-to now says an empty list means every REGULAR .toml is clean, and tells anyone symlinking their configs to check the targets by hand. Claiming a clean bill of health over files the check declined to open would be the same kind of false confidence the two false negatives above produced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3242): backfill changeset pr number (#3290) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |