next
406 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6b0b92674a |
refactor: remove dead descriptor-driven mechanisms (no remaining consumer)
Drop code whose only consumer was a retired runtime: hostBehaviors readers for agentFileExtension, localTargetIsProjectRoot, sharedHooksDirName, skillsManifestPrefix and skipCodexSkillsManifest; the empty NON_REGISTRY_CONFIG_HOME_DESCRIPTORS array and live-config-guard plumbing; the empty RUNTIME_NOTE_AUDIENCE_BY_HEADING filter; the unused resolveVersionFrom export; and the WINDSURF_SESSION_ID workstream session key. Delete tests that only exercised those mechanisms. |
||
|
|
792139b5ed |
chore: sweep Kimi mentions from comments and notes
Some checks failed
Tests / PR mergeability (push) Successful in 19s
Tests / Base branch health (push) Successful in 10s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 1m43s
Tests / plugin-validate (push) Successful in 1m7s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 18s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 19s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 19s
Tests / test (ubuntu-latest, 24) (push) Failing after 17s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 18s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 8s
Tests / Required tests (push) Has been cancelled
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
|
||
|
|
6cfa0c55d2 |
refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes, cline, codebuddy and pi end to end: capability descriptors, installer branches and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters, hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi migrations, Kimi payload normalization in the hook guards, dead hostBehaviors vocabulary, launcher home probes, fixtures, runtime-specific tests and the prose that presented them as supported. Installer output for the six kept runtimes is byte-identical to before the prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and read-injection-scanner are left in place pending a decision. |
||
|
|
12ee75a509 |
chore: point MSD at git.golem15.com/golem15/msd-core
Some checks failed
Tests / PR mergeability (push) Successful in 1m37s
Tests / Base branch health (push) Successful in 11s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 2m16s
Tests / plugin-validate (push) Successful in 1m6s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 25s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24) (push) Failing after 19s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 19s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 9s
Close Draft PRs (sweep) / Sweep open draft PRs (push) Successful in 8s
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
Tests / Required tests (push) Has been cancelled
The fork lives on the golem15 Gitea forge, not GitHub. Package identity now parses either host and derives in-place raw URLs for Gitea; the identity-drift lint accepts the new host; README drops GitHub-only badges and the npm quickstart in favour of the checkout installer. |
||
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
1e3e1f7cd8 |
enhance(#4570): allow disabling planner stall detection (#4585)
* enhance(#4570): allow disabling planner stall detection * docs(#4570): add changelog fragment Emitted-Drift-Ack-Growth: plan-phase.md — the explicit opt-out gate covers all five planner and checker spawn classes Emitted-Drift-Ack-Growth: settings-advanced.md — the toggle prompt and bounded-recovery warning expose the new setting * chore(#4570): refresh compact-content baseline * fix(#4570): preserve default-on watchdog fallback * docs(#4570): qualify the chunked-mode orchestrator rules with the toggle The two chunked-planning-mode stall-watch imperatives read as unconditional, with the opt-out stated only in a following bullet. State the PLANNER_STALL_DETECTION_ENABLED condition inline, matching the three sites already qualified in plan-phase.md. * chore(#4570): refresh compact-content baselines against rebased next * fix(#4570): sync planner stall launcher Keep the stall-detection helper aligned with the canonical runtime launcher. * chore(#4570): refresh compact baseline after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
17bf25a3d2 |
docs(#4671): lock the config-resolution and CLI-encoding design — Phase 0 of #4633 (#4829)
* docs(#4671): lock the config-resolution and CLI-encoding design
Amend ADR-1411 in place with the design contract for epic #4633: one
per-key resolution owner that returns the producing layer, two declared
configuration families with no precedence invented between them, one
parser and one encoder at the config-get boundary, a round-trip property
scoped to what raw output can actually carry, and the migration census,
child boundaries and anti-divergence guard spec.
Every claim about current behaviour was reproduced on next at
|
||
|
|
76ef60ba25 |
enhance(#4836): prefer the graphify CLI for planner and researcher graph queries (#4874)
* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries The planner gets one knowledge-graph query per phase and the researcher two or three, and that single shot decides which modules the plan treats as related — and therefore how tasks are ordered into waves. It was spent on the built-in reader, which seeds by case-insensitive substring match over a node's label and description and then expands a hardcoded two hops. The phase "User Authentication" seeds on `author`, `authoring` and `unauthorized` with the same weight as `authenticate`, and when the inflated payload exceeds `--budget` the trimmer drops edges by confidence tier — so the highest-confidence tier can be discarded to fit a payload that bad seeding inflated in the first place. The graphify CLI is already a hard dependency of /gsd-graphify build, and it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context filters before traversal. Both prompts now prefer it and fall back to the built-in reader, branching on `command -v graphify` — the same degradation shape the repo already uses for Context7 to ctx7. Binary presence is a self-satisfying gate: a graph can only exist if the binary built it, so the fallback covers edge cases (a CI checkout with a committed graph, a binary since removed), not the common path. No new config key and no new tool grant — both agents already have Bash. The planner additionally runs `graphify affected`. The reference states its own goal as "which subsystems may be affected by changes in this phase", which is literally reverse traversal by relation; the built-in reader only approximates it with undirected two-hop expansion and has no equivalent verb, so `affected` is skipped on the fallback path. `graphify status` now reports `graph_path`, the resolved absolute graph location, on both the present and the missing branch. The CLI takes the graph location as `--graph`, and the prompts must not re-derive `.planning/graphs/graph.json` for it: that would point the CLI at a non-existent local mirror in exactly the umbrella multi-repo setup `graphify.graph_path` (#1825) exists to serve. For the same reason the presence gate in both prompts is now the `status` call itself rather than a bare `ls` of the default location, which was already blind to the override. Known limit, stated in both prompts rather than implied: the two paths return different shapes. `graphify query` emits prose and has no `--json` flag; the built-in emits JSON with per-edge confidence tiers and budget_met/budget_estimate. `--budget` also counts rendered output on one and estimated payload bytes on the other (#2738) — same flag name, different unit. Both are read by a model and nothing machine-parses the injected block. With graphify absent from PATH the injected context is byte-identical to before. Closes #4836 Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point. Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes. * docs(#4836): record the CLI-first graph query in the planner and researcher entries * chore(#4836): add changeset fragment * enhance(#4836): name the full domain word in the planner's query-term examples The reference's own example — phase "User Authentication" → term "auth" — is the exact collision the CLI-first path exists to avoid, and it stays a collision whenever the fallback path runs, since that path matches the term as a substring of label and description. * fix(#4836): surface graph_path on the unparseable-graph status branch graphifyStatus() returned graph_path on the exists:true and exists:false outcomes but not on the third, error, outcome (graph.json present but unparseable). The planner/researcher prompts gate CLI-first dispatch on exists, not on this outcome, so a corrupt graph file made them fall through to the CLI-first branch with the literal <graph> placeholder and no real path to substitute. * docs(#4836): note graph_path's trust boundary at the --graph interpolation graph_path is reflected verbatim into a double-quoted --graph argument the agent executes via Bash. It comes from graphify.graph_path, a config surface already trusted elsewhere, so this isn't a new trust boundary -- but it is a new injection site (no --graph flag existed on this call before). One-line caution for anyone hardening this later. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
55fba5f7ce |
feat(#4917): add the PlanningDoc parse → mutate → serialize seam — Phase 1 of #4906 (#4918)
* feat(#4917): add the PlanningDoc parse -> mutate -> serialize seam Phase 1 of epic #4906, implementing ADR-4910 and its 2026-09-21 amendment. Net-new leaf module; NO call site is migrated, so nothing in the twelve absorbed issues changes behavior yet. src/planning-document.cts composes the seams that already exist rather than reimplementing them: markdown-sectionizer for structure (fences and code spans come from stripFencedCode / scanInlineCodeSpans, never a second scanner), markdown-table for tables, frontmatter for frontmatter, write-set for Result<T>. What is structural rather than conventional: - A field node carries labelSpan, valueSpan and trailingSpan separately, and the only write entry point takes a node id and writes into valueSpan. trailingSpan is readable and has no exported writer, so the #2853/#3584/#4852 rule ("the verb owns the count token ONLY") stops depending on an author remembering a third capture group. - Mutation is node-addressed. A handle is minted by the parser, so a caller cannot name a node the parser did not find. No path strings — a path is a grammar, and a grammar needs a parser. - serialize splices staged spans into the ORIGINAL buffer. A no-edit serialize is byte-identical, which is what eliminates the #4499 defect without a targeted fix, and which also makes an already-escaped table cell impossible to double-escape on a round trip. - A node that fails to parse carries its own error and span; siblings stay readable. - Per the ADR amendment, serialize REFUSES whenever any node carries a parse error, even with zero staged edits, naming the offender. One defect was found while building and fixed in place rather than deferred: PLANNING_ARTIFACTS first derived from isCanonicalPlanningFile unfiltered, so config.json, state.json, milestone.lock and skill-manifest.json were accepted and returned {ok:true, nodes:[]} — "this document records nothing" when the truth was "I have no grammar for this file". That is the exact empty-vs-error confusion this epic exists to remove (#4899, #4900), reproduced inside the seam built to prevent it. The registry now filters to .md while still DERIVING from isCanonicalPlanningFile, because hand-writing a second list is the divergence this epic is about, and parsePlanningDoc refuses a non-markdown kind at the document level per ADR-4910 section 5. New bin/lib module bookkeeping: .gitignore, eslint.config.mjs ignore (ADR-457 — lint the .cts, not the emitted .cjs), docs/INVENTORY.md row, regenerated docs/INVENTORY-MANIFEST.json, and the CONTEXT.md glossary entry. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4917): cover the PlanningDoc seam across 28 input classes 30 cases in 14 describe blocks, one per row of the phase test matrix. The two load-bearing tests are the fast-check properties (seed 20260921, numRuns 200): - a single node mutation leaves every byte outside that node's valueSpan identical to the source - serialize with zero staged edits is the identity function Both are DOCUMENT-SHAPED per CONTRIBUTING.md fixture provenance (#2371): the generator assembles arbitrary frontmatter, heading, label and value text with join(), and never calls serialize or any other function from the module under test to build a fixture. Seeding the generator from the module's own writer would make the document shape a constant, and the property could then never explore a document the writer would not itself emit. Verified the byte-range assertion is not vacuous with a control run: it passes against the real writer and FAILS against a simulated #4852 writer (one capture group, replace-to-end-of-line), which visibly drops the trailing annotation. Boundary coverage is zero / one / two staged edits. Negative space carries its own rows — bold emphasis in prose, a field-shaped line inside a fenced block, the same inside an inline code span, and a horizontal rule mid-body all correctly mint no node. Row 24 is the Generative-Fix-Divergence parity assertion: PLANNING_ARTIFACTS must not diverge from isCanonicalPlanningFile. Row 28 covers the non-markdown canonical file found during the build. Assertions are structural throughout — a typed Result / NodeRead / SerializeOutcome shape, or a byte range computed from the node's own Span. Full-string equality appears only where the contract IS byte equality. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4917): apply review findings — refuse unrepresentable values, compose the layers the seam claimed Three review passes ran against this branch: /security-review, an isolated adversarial pass, and a standards+spec pass. Five findings, all fixed here. Every one is the epic's own failure class reproduced inside the seam built to end it, which is the thing worth noticing. 1. setFieldValue accepted a value containing a newline. It survived serialization and reparsed as a REAL sibling field — one write to Plans forged a second Owner into a document that already had one. Content became structure, which defeats ADR-4910 Decision 2's "cannot reach past its own token by construction": the token boundary is a LINE boundary. 2. setFieldValue accepted a value containing the trailing separator and SILENTLY TRUNCATED it. Staged "sneaky - annotation", read back "sneaky", with the remainder reclassified as trailing prose. No error, nothing unreadable, both resulting nodes parsing perfectly. Worse than (1) because it loses the caller's own value rather than adding something visible. Both are fixed by ONE general check, deliberately not a blacklist: setFieldValue rebuilds the candidate line, re-parses it through the same field grammar, and refuses unless the value reads back identical. Blacklisting the separator would close this instance and leave the class open for whatever separator the grammar grows next. The round-trip check is ADR-4910 Decision 4 stated executably. 3. The module reimplemented two layers it claims to compose. Checklist detection hand-rolled a checkbox regex that markdown-sectionizer's iterateBullets already owns. Frontmatter span detection re-derived fence handling because frontmatter.cts's frontmatterRegion was module-private — so ADR-4910 section 1's stated layering was UNREACHABLE as written, and the first implementation routed around it silently instead of surfacing the gap. frontmatterRegion is now exported (additive only; ADR-2143 section 2's extend-never-mutate lock is inherited) and both layers are consumed. 4. The CONTEXT.md glossary entry asserted "the Frontmatter Module supplies frontmatter" while zero frontmatter.cts code was invoked. That was a false claim in the repo's vocabulary of record, written by me, and it is now true rather than edited away. 5. Adopting iterateBullets narrowed GFM coverage: it classifies only dash-prefixed task items as checkboxes, so "* [ ] x" and "+ [x] y" stopped becoming checklist nodes. Widening the sectionizer is forbidden by the inherited lock, so the task-list MARKER is interpreted in this module while bullet STRUCTURE still comes from the sectionizer. The sharpest finding was not a defect. The fast-check generator constrained values to [A-Za-z0-9 .,!?], so it could not emit an em-dash, newline, backtick, pipe or asterisk — precisely where (1) and (2) lived. The property was real, seeded and non-vacuous, and structurally blind to the module's actual bug class. The generator now spans the grammar's own metacharacters, and a new property asserts that every value setFieldValue ACCEPTS round-trips identically. Proven able to fail: against a scratch copy with the guard stripped it fails after 7 cases on newValue "\n". Recorded as a measured boundary, not fixed: a bare CR inside a field line leaves that field unrecognised. Measured — sibling fields still parse, no error node, and serialize stays byte-identical, so the worst case is an unreadable field and never a damaged document. Flagged for Phase 4's empty-vs-error census. Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4917): backfill changeset pr number to 4918 Refs #4906 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b956bb7c67 |
fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs (#4902)
* test(#4834): failing-first launcher hijack regressions * fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs A gsd_run on PATH that cannot prove it is @opengsd/gsd-core (a foreign package, or a release older than the runtime-identity verb) is no longer accepted by the launcher snippet's PATH arm; resolution falls through to the hard error when no path-based candidate matches. The runtime-config-home arm now precedes the PATH arm, restoring the documented prefer-local-over-PATH order, so an installer-managed install wins even against a genuine global. The 16-home probe list is factored into _gsd_homes() and the identity gate into _gsd_id_ok(), keeping the per-copy delta at +141 bytes. The files whose frozen ceilings had no headroom (gsd-executor, gsd-plan-checker, gsd-verifier, gsd-planner, execute-phase, execute-plan) now load the resolver by @-include from gsd-core/references/gsd-run-resolver.md (the onboard.md pattern) instead of carrying an inline copy. Propagated to all other inlined workflow/agent copies via scripts/sync-runtime-launcher.cjs; the resolver reference re-copied byte-equal (parity B2); the hard-error text, docs/how-to/diagnose-a-foreign-gsd-tools.md, and the CONTEXT.md launcher predicate updated to match (#4834); the quick-batch row-48 guard gains the canonical-preamble sweep carve-out (#4834, per its own #3730/#2529 precedents); the compact-content benchmark baseline regenerated. Emitted-Drift-Ack-Growth: add-backlog.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-tests.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: add-todo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ai-integration-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: audit-uat.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: autonomous.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: check-todos.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: cleanup.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: code-review-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: code-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: complete-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: debug.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: diagnose-issues.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: discuss-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: do.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: docs-update.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: edit-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: eval-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: explore.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: extract-learnings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: fast.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: forensics.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: graduation.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-code-fixer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-debugger.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-eval-auditor.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-eval-auditor.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-intel-updater.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-intel-updater.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-project-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-project-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-research-synthesizer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-research-synthesizer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-ui-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: health.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: import.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: inbox.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ingest-docs.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: insert-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: list-seeds.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: list-workspaces.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: map-codebase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: milestone-summary.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: mvp-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: new-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: next.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: pause-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plan-review-convergence.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: plant-seed.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: pr-branch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: profile-user.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: progress.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: quick-batch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: quick.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: remove-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: remove-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: resume-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: scan.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: secure-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings-advanced.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings-integrations.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: settings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ship.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sketch-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sketch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: smart-entry.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spec-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spike-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: spike.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: stats.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: sync-skills.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: thread.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: transition.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ui-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ui-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: ultraplan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: undo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: validate-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) Emitted-Drift-Ack-Growth: verify-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834) * docs(#4834): backfill the changeset PR number * test(#4834): regenerate the compact-content benchmark baseline after the rebase --------- Co-authored-by: sim <sim@local> |
||
|
|
eea9247c93 |
enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select Auto-mode used to auto-select a checkpoint:decision's first <option> unconditionally, making a decision checkpoint's safety depend on option presentation order rather than an authored choice. Add an optional auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode now escalates to a human exactly like gate="blocking-human" does; present, it names the option auto-mode selects; naming an id with no matching <option id> is a hard structural-validation error at plan-parse time rather than a silent fallback to the first option. gate="blocking-human" continues to win over everything, unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys An isolated adversarial review of the auto_select work found that both new attribute regexes used \b as their left anchor, which is a word boundary, not a "start of attribute name" boundary. A decoy attribute ending in the same word (e.g. data-id="...") sitting before the real id="..." on the same <option> tag matched first, silently corrupting the extracted option id. Anchor on (?:^|\s) instead so only the real attribute name can match. Adds a regression test reproducing the exact decoy-attribute shape, plus a Unicode option-id test from the same review pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane lint-docs-guard-registration failed: the new test reads docs/reference/ plan-md.md but was not registered, so a future edit to that doc could silently desync from the test without the guard catching it on the PR that changed the doc. Registered alongside its direct precedents (precondition-element.test.cjs, reversibility-tagging.test.cjs), which read the same file for the same reason. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling The remote gsd-test run caught what local checks missed: execute-phase.md carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes of headroom before this change, and the original checkpoint:decision wording pushed it to 93772 (over the ceiling). Cascaded into failures in phase6-capstone-conformance, execute-phase-completion-reconciliation, claude-orchestration, and the compact-content drift-report test. Also caught: tests/package-legitimacy-gate.test.cjs anchors a "decision is conditional, not unconditional" safety check on the literal phrase "first option" in the decision bullet — which #4095 deliberately removes, since there is no more unconditional first-option pick. The test was asserting an assumption this change intentionally makes obsolete; re-anchored on tokens that still identify the bullet ('decision', 'auto-spawn') without weakening what the test actually verifies (the bullet must still carry a blocking-human carve-out). Also fixed a word-order mismatch between my own new test's regex and the actual doc text it was asserting against (tests/auto-select-attribute.test.cjs). Regenerated the compact-content benchmark baseline (tests/fixtures/compact-content-benchmark-baseline.json) to match the new byte counts. Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap. Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4095): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5e729445d3 |
fix(#4657): give the ui consideration probe a text_en language channel (#4804)
* test(#4657): add failing-first coverage for the ui probe's text_en channel Mirrors the #3717/#4156 test shape onto the UI adapter: a failing-first proposeConsiderations regression (Danish text + English text_en must classify as its English equivalent, not land in the #1110 unclassified sentinel), proposeElements/analyzeCoverage/CLI end-to-end pairs, fail-closed text_en validation cases (empty/whitespace/non-string, unconditional under an elements override), a ui-phase.md Step 9.5 workflow-prose contract test, a reference-doc Inputs parity test, and a fast-check property proving any cue-matching prose classifies identically under a cue-free Danish rendering plus text_en. All new assertions are RED until src/ui-consideration-probe.cts and the workflow/reference docs are updated. * fix(#4657): give the ui consideration probe a text_en language channel Element gains an optional text_en; classifyElement's own signature stays untouched (a locked, directly-tested export) and the text_en ?? text selection is pushed to the two classification call sites (proposeConsiderations, proposeElements) instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero kinds. Mirrors #3717/#4156 onto the UI adapter: ui-phase.md Step 9.5 gains the Non-English projects section (mirroring spec-phase Step 5.5) and the ELEMENTS_JSON shape comment documents the field with both zero-applicable guard arms named; the reference doc's Inputs section, the PROBE.ui CONTEXT predicate (with both derived indexes regenerated), and the nav-override test expectation stay in sync. The ui-phase contract test carries the site-scoped allow-test-rule marker and its cluster is registered in the test-file-count allowlist ratchet. Emitted-Drift-Ack-Growth: ui-phase.md — Non-English text_en section, ELEMENTS_JSON shape comment, and two-arm guard wording (#4657) * docs(#4657): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
003d982c83 |
fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint Two defects in the covered-input fingerprint (#4155), one issue. 1. `computeCoveredDigest` hashed the whole bytes of every declared path uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` — which every phase rewrites as ordinary bookkeeping, and which the closing phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER the verifier ran — flipped every phase that declared them to `stale` on zero implementation change, and from there `isPhaseComplete` → `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate. Fingerprint v2 leaves any direct child of a planning root out of the hash: `.planning/` itself, plus the phase's own planning root (the parent of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/` layouts are covered without the digest knowing what a workstream is — `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather than a name list so the set cannot drift; a root is accepted only when the phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated exactly as every other covered path (confined, present, a regular file — the fail-closed contract is unchanged); only its bytes are ignored, and a declaration made only of shared documents fails closed like an empty one. A stored digest names its version, and `readVerificationStatus` now recomputes under THAT version (`parseFingerprintVersion`, `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics until it is re-fingerprinted, so the upgrade alone stales nothing; a version this build cannot recompute fails closed. 2. `verification.fingerprint` received a raw positional slice, so `--files a`, `--files "a,b"` and `--files a --files b` all put the literal token into the covered set and failed closed as "a covered file is missing, unreadable, or escapes the project root" — the message that convinced the reporting project the digest was permanently unrecomputable. `parseFingerprintFileArgs` accepts every form (plus `--files=a,b`, freely mixed with bare positionals), treats any other `--flag` and an empty `--files` value as usage errors that say so, and the phase-dir argument must now be an existing directory: omitting it used to take the first covered file as the phase dir and print a plausible digest over the rest at exit 0. Regression tests (tests/verification-status.test.cjs, #4623 block): the cross-phase case from the report, the same-phase `requirements mark-complete` / `phase.complete` cases from the thread, a workstream-scoped root, v1-preserved / unknown-version-stale, the fail-closed cases (missing, directory, escaping symlink, all-shared), every `--files` form against the bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34 fail, the 7 that pass pin behaviour the fix must leave unchanged. Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's covered_files instruction (rewritten in place — the file sits 21 bytes under its LARGE hard cap), gsd-core/templates/verification-report.md. Fixes #4623 Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi * chore(#4623): set changeset fragment pr to 4749 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ad1477d659 |
enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap * enhance(260903-m7p): validate configured entrypoints before success * test(260903-m7p): require pre-success entrypoint validation * enhance(260903-m7p): gate install success on entrypoints * test(260903-m7p): cover configured entrypoints across runtimes * enhance(260903-m7p): cover emitted runtime entrypoints * fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number - finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process before the new configured-entrypoint assertion throws. Without a HOME + config-location-env sandbox that write resolved through the ambient environment and landed in the developer's live ~/.gsd (confirmed absent on origin/next baseline, present only on this branch — full-suite HERMETICITY WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the duration of the test, matching the existing in-process finishInstall/ install() pattern in tests/install.test.cjs (#2665). - .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the fork PR number until the upstream PR number is known. * fix(260903-m7p): repair cross-platform and pre-existing shape fallout - tests/configured-entrypoint-validation.test.cjs: the win32 branch of ensureCodexHooksJsonSessionStart writes a .cmd shim under <codexRoot>/hooks/; create that dir in the test (the real installer only calls this once hooks/gsd-check-update.js already exists) and assert the platform-common entrypoint shape instead of a fixed non-Windows array, since win32 legitimately emits two entries (cmd shim + script). - tests/install.test.cjs: finishInstall's shared settings-json return now carries configuredEntrypoints/rollbackInstallerMigrations for every runtime on that path (trae included, not just Claude/Cursor/Windsurf); update the trae install() exact-shape assertion to match. * fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash configuredEntrypointsForHook's shell branch dropped interpreterCandidates entirely when resolveBashExecutable returned null, unlike the sibling portableHooks runner entry a few lines below (which correctly falls back to the literal 'bash' token). Found via agy adversarial review; verified unreachable through the current call graph (buildHookCommand's own resolveBashRunner==null gate already short-circuits before recordConfiguredHookCommand runs), so this is a defensive consistency fix, not a live-bug patch — kept for the next caller that does not share that gate. * chore(260903-m7p): backfill changeset pr number to the opened upstream PR .changeset/quick-wasps-sing.md carried the fork PR number (16) as a placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now open, so record its real number per CONTRIBUTING.md's changeset pr-field convention. * fix(#4154): track already-registered hooks for entrypoint validation on update applySettingsJsonHooks registers each guard hook only if absent, so a hook already present from a prior install keeps its stale on-disk command. The new entrypoint tracker always records the freshly-computed command for it, which never matches what is actually persisted, so the exact-string filter in finishInstall silently dropped it from validation — the Blocker case this feature exists to catch (an already-installed entrypoint going stale between installs) was exactly the case it never validated. Match on the managed script's basename instead, which the persisted command carries either way, so an already-registered hook stays in the validated set. Regression test forces this path by mutating a freshly-installed hook's persisted command before a second install. * fix(#4154): distinguish an unreadable script from a missing one validateConfiguredEntrypoints folded an EACCES statSync failure into the same 'missing' reason as ENOENT, misreporting a real permission problem as an absent file. Check the error code and report 'unreadable' instead. * docs(#4154): document entrypoint validation's rollback and PATH scope CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite bin/install.js x CONTEXT.md being this repo's strongest co-change pairing. The update-gsd.md how-to overstated what a validation failure undoes: for Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/ config.toml inside install() before the aggregate validation call runs, so there is no rollback path for that write regardless of "where available" phrasing. Also note that interpreter resolution checks the installer's own PATH, not necessarily the PATH a hook fires under later (#2979 launchers). * chore(#4154): point changeset pr field at the fork PR while CI runs there Mirrors the branch's own prior backfill commit: pr: matches whichever PR number changeset-lint is currently validating against (fork PR #16 during the fork-first CI/review loop), flipped back to the upstream PR number right before the final push to open-gsd/gsd-core. * fix(#4249): address adversarial-review findings in entrypoint validation An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR found several real gaps beyond the human reviewer's Blocker, verified against source before fixing: - Codex's install() result bound rollbackInstallerMigrations to the narrow installer-migrations-only rollback instead of restoreCodexSnapshot (#3245), the full pre-install snapshot/restore Codex already owns for exactly this case — a validation failure discovered outside install() reverted nothing of the config.toml/hooks.json that call had already written. - The register-only-if-absent basename match from the prior fix used a bare substring, which an unrelated user command mentioning the same filename could false-positive into GSD's validated set — anchored on the `/hooks/<basename>` path segment instead. - nodeCandidates checked raw process.execPath (always true — we're running in that process) instead of normalizeNodePath's stable version-manager alias, the same one buildNodeRunnerChainToken bakes as its first choice — a false green regardless of whether that alias itself still resolves. - An entry with no interpreterCandidates (Cline's PreToolUse hook, or a Windows-Claude .sh hook invoked without a bash runner) runs via its own shebang; validateConfiguredEntrypoints checked only file-type, never the execute bit. Cline's writer also never reported an entrypoint at all. - Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor hook registered across several events) were validated once per duplicate. Each fix is covered by a new or extended test; the Codex one required inlining runCodexInstall's env sandboxing so the rollback closure — which re-resolves the $HOME-relative skills root live — runs before the sandbox is torn down, matching how installAllRuntimes' real aggregate gate calls it. * docs(#4249): document the round-2 entrypoint-validation fixes Runtime Hooks Surface Module and Installer Module entries now name ConfiguredEntrypoint's not-executable reason, the normalizeNodePath alignment, Cline's tracked hook, and which install() result the finishInstall/installAllRuntimes rollback path actually reverts per runtime (Codex's full snapshot vs. the others' narrow migrations-only rollback). * chore(#4249): point changeset pr field at the upstream PR now that fork CI is green * fix(#4249): address agy adversarial-review findings - validateConfiguredEntrypoints: statSync alone never detects a chmod-000 script (it only needs parent-dir search permission), so an interpreter-invoked entry with an unreadable script passed validation. Add an explicit R_OK check for the interpreterCandidates branch only — the candidate-less/shebang branch already has its own X_OK gate. - docs/how-to/update-gsd.md: the blanket "does not revert" claim was false for Codex, which reverts config.toml/hooks.json via its full pre-install snapshot; qualify it per runtime. - tests/codex-config.test.cjs: the #4249 rollback regression test asserted skills/ and VERSION were reverted but never asserted config.toml/hooks.json were too, despite the test's own stated intent. - CONTEXT.md: qualify which interpreterCandidates entries get normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's Windows .cmd shim as a third candidate-less case that relies on extension dispatch, not a shebang. * fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node', so it needs the execute bit (like any shebang-invoked entry), but its interpreter is looked up on PATH by 'env' at hook-fire time (unlike every other GSD JS hook, which bakes an absolute node path specifically to avoid that dependency). The candidate-less/interpreterCandidates fork treated these as mutually exclusive, so Cline's entry silently skipped interpreter resolution entirely — a completely missing 'node' on PATH would still validate successfully. Add an orthogonal selfExecutable flag so both checks run for entries that need them. (CodeRabbit finding on the fork rehearsal PR.) * fix(#4249): address second-round adversarial review findings (opus + agy) - validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry, not just interpreterCandidates ones — a self-executable shebang script is still opened and read by its kernel-invoked interpreter, so X_OK alone never proved it was readable. - selfExecutable is now the sole, explicit source of truth for the execute-bit check (every producer that needs it sets the flag) instead of being partly inferred from an absent interpreterCandidates, which Cline's hybrid entry also carries. - The execute-bit check now skips explicitly on win32 (matching resolveExecutableBinary's own carve-out) instead of relying on Node's accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows machine and not a test that simulates win32 on a POSIX runner. - bin/install.js: fixed a stale comment claiming no runtime's install()-time writes have a rollback path — Codex's does (restoreCodexSnapshot) — and added the omitted Cline to both that comment and CONTEXT.md's equivalent lists. - CONTEXT.md: fixed the Cline description left stale by the previous commit's selfExecutable addition, and rewrote the validation-mechanism paragraph for clarity (writing-for-agents pass). - docs/how-to/update-gsd.md: split an overloaded 4-clause sentence. - Removed a fault-injection integration test that could not reliably exercise the real installAllRuntimes -> finalize -> rollback wiring without fighting the installer's own pre-registration existence guards; the constituent pieces remain covered individually. * fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner X_OK is a POSIX-only concept, skipped entirely when an entry's platform is win32 (matching production). Two test entries omitted platform, defaulting to process.platform — on an actual windows-latest CI runner that silently skipped the very check they were meant to exercise, turning 'not-executable' into a false pass. Pin platform: 'linux' so these are deterministic regardless of which OS runs the suite. * fix(#4249): classify EPERM the same as EACCES in statSync error handling Windows raises EPERM (not EACCES) for a parent directory that couldn't be traversed into — was falling through to 'missing', misreporting a genuine permission problem as a nonexistent path. * docs(#4249): address final CodeRabbit doc-completeness findings - CONTEXT.md: install()'s documented result shape omitted configuredEntrypoints; the ConfiguredEntrypoint shape omitted selfExecutable. - docs/how-to/update-gsd.md: the failure-mode sentence omitted unreadable and lacks-execute-permission, which the installer also rejects. * fix(#4249): stop double-validating every configured entrypoint on install/update installAllRuntimes' finalize() already runs assertConfiguredEntrypoints once over the aggregate set; finishInstall then re-ran the identical check per runtime in the printSummaries loop right after, so every entrypoint paid its statSync/accessSync/interpreter-resolution cost twice on every install and update. Add entrypointsAlreadyValidated to skip the redundant pass specifically on that path, while leaving the check intact for any caller that invokes finishInstall directly. * chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there * perf(#4249): memoize interpreter candidate resolution across entrypoints resolveExecutableBinary walked PATH once per (entry, candidate) pair; a typical install has a dozen-plus entries sharing the same few candidate lists (process.execPath for JS hooks, bash for shell hooks). Cache by (platform, candidate) so each distinct pair resolves once per validation call instead of once per entry. * chore(#4249): point changeset pr field at the rebased rehearsal fork PR * fix(#4249): drop entrypoint tracking from the now-dead Codex event writer #2586 (landed on next after this branch forked) removed install.js's CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent no longer runs during install or update. The ConfiguredEntrypoint records this branch added inside it were therefore unreachable and untested. Restore the function to its upstream shape; the entrypoints it used to report were never collected by any caller. * refactor(#4249): drop the revalidation bypass flag and the candidate cache Both were this PR's own micro-optimisations over a set of roughly a dozen entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall gate off to save one statSync/accessSync pass; `resolvedCandidateCache` memoised resolveExecutableBinary across entries that are already deduped by (configPath, scriptPath). Neither is measurable, and the flag was the only way to reach finishInstall with validation disabled. finishInstall now always validates what it is given. * chore(#4249): point the changeset pr field back at the upstream PR * refactor(#4249): track settings.json entrypoints without the hooksSurface gate The install-surface writer only tracked configured entrypoints when the runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing asserts that axis agrees with `installSurface`, so a descriptor that broke the coupling would silently pass `configuredEntrypoints: undefined` and drop that runtime out of the validation this PR adds — reintroducing the exact 'reports Done! over a broken entrypoint' failure #4154 exists to close. Remove the dependence rather than test it: everything recorded on this path lands in settings.json by construction, and the registered-command filter already discards entries no persisted hook references. * chore(#4249): put the changeset body in the documented two-part format CONTRIBUTING.md and .changeset/README.md both show `**<bold change>** — <symptom-led explanation>.`; the fragment was a single unbolded sentence. * chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there * fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback #3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-* and gsd-core/VERSION. The install overwrites every other GSD-owned file too — hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest itself — before the entrypoint-validation gate runs, so a validation failure left the new payload sitting on top of the restored old config. Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims, before runInstallerMigrations so the bytes are the true pre-install state, and restore it from both Codex rollback closures ahead of the per-surface restores. Files only the failed install introduced are removed, read from the manifest now on disk. The manifest is already the authoritative record of what GSD owns, so no second hand-written list can drift out of sync, and user-owned files are never snapshotted or removed. Every path is confined through resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into an arbitrary-path write. Non-Codex runtimes are unaffected: the snapshot is gated on the same tomlConfigInstall + non-minimal condition as #3245's. * fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest Two follow-on defects in the previous commit's snapshot: - The capture was gated on `!isMinimalMode`, copied from #3245. A core/ --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the snapshot came back empty while the rollback still ran — and its removal pass would have deleted every file the new manifest lists. Gate on tomlConfigInstall alone, matching where the rollback actually reaches. - An unreadable or unparseable prior manifest was caught alongside ENOENT and treated as a fresh install. That is the same empty-snapshot state, so a failed update over a real install with a corrupt manifest could delete its prior payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means known-empty; any other read error or a parse failure means unknown, and the restore closure returns without touching anything, degrading to #3245's narrower rollback. Deliberately not fatal — a corrupt manifest has to stay repairable by reinstalling over it. Both paths are covered by red-checked regression tests. * fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too commit removed from the manifest snapshot. restoreCodexSnapshot is reachable for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-* skill dir and gsd-* agent file the snapshot does not claim — so with an empty minimal-mode snapshot a rollback deleted the whole skills/agents surface with nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via the ADR-1239 skills-kind home override, so this is also the reason manifest `skills/` keys do not resolve under configDir: that surface belongs to this snapshot, not to the manifest-driven one. Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal mode — doing nothing on an early failure is the non-destructive side. Covered by a red-checked regression test that plants bytes in an alternate-home skill file, reinstalls under the core profile marker, and asserts the rollback restores it. * fix(#4249): never remove on rollback unless a prior manifest proves what predates the install Three defects in the manifest-driven Codex rollback, all in its removal half: - ENOENT marked the snapshot usable, arming the removal pass on a FIRST install. GSD may have overwritten a user's file at a manifest-tracked path there, and no prior manifest records the difference — so rollback deleted it where before it merely left it overwritten. Absent, unreadable and malformed manifests now all leave the prior set UNKNOWN and skip removal entirely. - Membership was tested against the map of files whose pre-install read SUCCEEDED, so a tracked file that existed but was unreadable read as introduced-by-this-install and was removed. Track the prior manifest's paths in their own Set and test against that. - The unreachable "delete the manifest when there was no prior one" branch is gone: usable now implies a parsed prior manifest. Also adds the end-to-end test the aggregate gate was missing — the four Codex rollback tests drove the closure directly, proving the restore but not the wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes Cline's `env node` entry fail validation for real, and asserts Codex's payload comes back. Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper instead of six copies. Both new tests are red-checked. * test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its Windows-EBUSY retry budget. That budget is for directory trees; this removes a single file, which unlinkSync says more precisely and the rule does not flag. * chore(#4249): point the changeset pr field back at the upstream PR * fix(#4249): use an unambiguous dedup key and surface partial-restore failures trek-e's 2026-09-08 adversarial pass flagged two findings in the new entrypoint-validation/rollback code: - assertConfiguredEntrypoints' dedup key already used a raw NUL separator (introduced in ceebb65f2d), but git/Read render NUL as a space, so the key looked like a plain-space join to every reviewer that read the diff. Replace it with JSON.stringify([configPath, scriptPath]) so the separator is visible and unambiguous. - restoreManagedFileSnapshot's per-file restore catch block claimed to 'surface the original error' but only swallowed it, matching (and widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add an actual console.warn using the existing best-effort-warning convention, scoped to just this PR's new function. * fix(#4249): treat a files-less prior manifest as unknown, not known-empty agy's gemini-3.8-flash-high adversarial pass (round 5) found and I reproduced empirically: a structurally-valid manifest missing the files key (e.g. {"version":1}) parses without throwing, so Object.keys(undefined || {}) silently read as 'zero files predate this install' instead of the UNKNOWN state the malformed-manifest guard exists to produce. Rollback's removal pass then deleted every GSD-owned file the failed install's own manifest listed, including ones that predated it — the exact data loss the #4249 CodeRabbit malformed-manifest fix was supposed to prevent, reachable through a JSON.parse success instead of a failure. Route the shapeless case into the same catch-all UNKNOWN path via an explicit shape check. Regression test reproduces the deletion before the fix and confirms the file survives after it. Also extend restoreManagedFileSnapshot's removal-pass rmSync and final manifest-rewrite catches with the same real console.warn trek-e's round-4 review asked for on the per-file restore catch — same rollback function, same operator-facing-signal gap. * docs(#4249): correct which runtimes actually leave a written config on rollback agy's completeness audit (round 5, holistic pass) caught this new paragraph claiming 'for every other runtime, the configuration file(s) already written during that update are left in place' — false for Claude Code and other settings.json-based runtimes, whose write never happens on failure (assertConfiguredEntrypoints runs before finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline actually match that description, since they persist their config file inside install() ahead of the gate. Split the one sentence into the three actual outcomes; matches the PR body's own accurate Before/After wording, which this doc addition had drifted from. * fix(#4249): clean up doc/comment mismatches and dead fields from opus review Opus critical-code-reviewer + ponytail-review pass on the final diff: - assertConfiguredEntrypoints carried finishInstall's old docblock ("Apply statusline config, then print completion message") from before this function was inserted between comment and callee. finishInstall already has its own accurate #4249 comment, so the stale docblock is removed rather than moved. - checked: number on ConfiguredEntrypointValidationResult and error.configuredEntrypointValidation on the thrown error: the first had zero consumers anywhere in the repo, including its own defining file, and is removed. The second matches an existing repo convention (bin/install.js's installerMigrationRollbackFailures, #4249 predates this PR) of attaching structured diagnostic context to a re-thrown Error even before a consumer exists, so it's kept. - finishInstall's own assertConfiguredEntrypoints call is a redundant backstop on the real production path (installAllRuntimes's aggregate call already validates the superset first), but its comment read as though this call alone provided the before-the-write guarantee. Clarified rather than removed — it's the only gate for a caller that invokes finishInstall directly. * chore(#4249): split the manifest-driven rollback engine out into #4544 Issue #4154 asked the installer to consume a validation failure "through the existing rollback mechanism, without a second transaction mechanism". The manifest-driven rollback widening added during review (capture every path the prior gsd-file-manifest.json claims, restore those bytes, remove what only the failed install introduced) is that second mechanism on a plain reading. It is a real fix for a #3245-era gap, but an independent one, so it moves to its own bug report and PR. Removed here: - bin/install.js: the pre-install managed-file capture block and restoreManagedFileSnapshot, plus its call sites in _codexPreConfigRollback and restoreCodexSnapshot (99 lines). - tests/configured-entrypoint-validation.test.cjs: the five tests that exercise the manifest engine. - CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the widened restore. update-gsd.md again documents the #3245 surfaces only. Kept, because it is #4154's own scope: - the entrypoint-validation gate itself; - Codex's install() result binding rollbackInstallerMigrations to restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, gsd-core/VERSION); - the !isMinimalMode gate removal on that snapshot. Binding the closure to the result made it reachable for a core/--minimal install, where its pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot does not claim; an empty minimal-mode snapshot therefore deleted the whole surface with nothing to restore. The surviving aggregate-failure test now asserts on config.toml, a surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which only the manifest engine restored. Refs #4544 * test(#4249): cover configured entrypoints through the packed install path #4154's scope lists install smoke coverage alongside the installer gate — "assert representative configured entrypoints resolve for supported runtime profiles". The gate itself (assertConfiguredEntrypoints / validateConfiguredEntrypoints) is unit-covered by in-process install() calls; nothing proved the property survives npm pack -> npm install -g -> install.js. Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct config surfaces GSD writes launch paths into (settings.json, and hooks.json + config.toml) — run the tarball-installed installer into a throwaway HOME, then re-read that runtime's own written config and return the new ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check becomes a release gate on every matrix host without workflow changes. The scan re-derives paths from the written config instead of reusing the installer's own entrypoint list, and test I shows why that matters: a registration the installer never touched during a run is invisible to the in-process gate, so the install exits 0 and only reading the config back off disk catches the dangling launch path. * ci(#4249): pack a publish-shaped tarball in the install smoke lane `npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs after `npm ci` carries no hook scripts at all — the lane has been smoking a package that differs from the published one in exactly the artifacts the lifecycle smoke is supposed to launch. That went unnoticed because the lane's init runs `--local`, which registers no statusline and therefore registers no hook whose target is missing. A `--global` install on the same tarball exits 1 on #4249's own gate (`gsd-statusline.js (missing)`), which is what the new configured-entrypoint cycle performs, so without this step the cycle would report INIT_FAILED instead of checking anything. Build hooks before packing so the smoked tarball matches prepublishOnly. The CLI now reports 16 configured entrypoints for claude and 1 for codex instead of zero. * fix(#4249): scope Codex's full snapshot restore to entrypoint failures Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage exception un-install a Codex install that had already succeeded and already printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints` call. Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration` scopes finalize-stage rollback to installer *migrations* ("the executor uses the journal to restore modified paths"), and this PR's own operator-facing paragraph in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/ agents/VERSION revert to entrypoint-validation failures specifically ("If a script is missing, unreadable, ... For Codex, this reverts ..."). The wide behaviour is also incoherent as a transaction abort: the same doc says Cursor, Windsurf, Kimi and Cline keep the config they wrote inside install(). Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install bytes — on an update, silently downgrading a working Codex install to the previous version while the user had just been told it was Done. Select the rollback by error kind instead. `assertConfiguredEntrypoints` already tags its error with `configuredEntrypointValidation`, so the full snapshot restore runs for that error (and anything downstream of it, including finishInstall's per-runtime backstop) and the installer-migrations-only closure runs for everything else. The codex result now also exposes that narrow closure as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning the full restore, so the direct-call contract asserted by tests/codex-config.test.cjs is unchanged. Adds a regression test that installs codex+kilo together, injects EACCES on the Kilo permission write by monkeypatching node:fs (restored in a finally — never chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps the bytes the successful install wrote. Verified red against the pre-fix unconditional path. Cline cannot host this test: its plan is writesSharedSettings:false + finishPermissionWriter:null, so its finishInstall performs no write and has no non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally (unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded fs.writeFileSync. * docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing CONTEXT.md still described Codex's rollback as an unconditional bind to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint- validation failures via rollbackInstallerMigrationsOnly and the configuredEntrypointValidation error tag. Caught during the round-6 PR body pass. * fix(#4249): stop rollbackInstallerMigrations meaning its own opposite Codex's install() result bound `rollbackInstallerMigrations` to restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the actual installer-migrations-only closure behind `rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name meant the opposite of what it says, and CONTEXT.md had to concede as much in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure for every runtime, matching both its name and the meaning it already has on next, and the snapshot restore gets its own Codex-only field, `rollbackPreInstallSnapshot`. The selection in rollbackFinalizedInstallerMigrations collapses to one line and no longer needs a fallback chain. Also in this commit, all against the same rollback path: - Correct the rollbackFinalizedInstallerMigrations comment. It read as if the round-6 narrowing prevented any sibling-triggered revert of a Codex install the user has already seen "Done!" for. It does not, and is not meant to: `wide` is true for ANY entrypoint-validation error from ANY runtime, because the aggregate gate is all-or-nothing — an invalid Cline entrypoint reverts Codex's snapshot, which tests/configured-entrypoint-validation.test.cjs's 'an aggregate entrypoint validation failure rolls the Codex install back (#4249)' asserts directly. The discriminator is the error's KIND, not which runtime owns the failing path. Comment and CONTEXT.md now say that. - Name the runtime in the "Configured entrypoint validation failed" error. ConfiguredEntrypointInvalid already carries `runtime`; the message threw it away, leaving an operator of a multi-runtime install unable to tell whose entrypoint broke — which matters precisely because the failure can revert a runtime that was itself fine. - Set `configuredEntrypoints: []` explicitly on the copilot-instructions early return. Every other branch states the key; this one relied on installAllRuntimes' `(result.configuredEntrypoints || [])` defence. `[]` is correct, not a workaround: every Copilot hook is an inline printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no GSD-managed script or interpreter to resolve. No behaviour change beyond the error-message text. * docs(#4249): narrow the smoke scan's config-surface claim to what it checks RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is told to launch is registered in one of settings.json / hooks.json / config.toml, and that nothing else in a config dir is runtime configuration. Both halves are false as stated. Cline registers its hook at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native [[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi), a directory separate from Kimi's own GSD configDir — the same gap installer-migration 007 already documents as structurally unreachable. The scan is in fact correct for what it runs against: entrypointRuntimes defaults to claude + codex, whose launch paths do all live in those three top-level files. Restate the docstring at that scope, name the two known out-of-scope surfaces, and warn that adding either runtime to entrypointRuntimes without teaching scanConfiguredEntrypoints about its surface yields a scan that finds zero entrypoints and proves nothing. The entrypointRuntimes default comment carried the same overgeneralization ("every other runtime reuses one of them") and is corrected with it. Documentation only; no code change. * fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found that install()'s settings-json early return for an unparseable settings.local.json omitted configuredEntrypoints and rollbackInstallerMigrations from its result, unlike every other branch. rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations unconditionally, so this branch silently dropped its own installer-migration rollback on a later finalize-stage failure. Completed the return shape: configuredEntrypoints: [] (matching Copilot's equally-early no-entrypoints-yet return) and rollbackInstallerMigrations (already in closure scope). Red-then-green regression test added. * test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs Same internal adversarial review (round 8): this suite's before() packed the tarball directly, without the ensureHooksDist() guard every sibling install-test suite (install.test.cjs, install-minimal-hooks.test.cjs, mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run in isolation ahead of a suite that builds hooks/dist itself, this suite's pack would ship a tarball with no hook scripts and fail closed on SMOKE.INIT_FAILED instead of testing anything. * fix(#4249): refresh stale test-timings weight for the codex-config split next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the CI shard packer's weight table (tests/test-timings.json) was never updated: codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block this PR extends with its own #4249 install()-pipeline test — had no entry at all, so the packer would silently underestimate it at the table's median weight (roughly a 9x underestimate against its real cost). trek-e's most recent review flagged a Windows shard timeout in-flight on codex-config.test.cjs, plausibly aggravated by this PR's own addition to that file before the rebase moved it. Re-measured both files locally (node --test --test-reporter=tap, max of 3 runs, matching the table's own max-across-streams methodology) and patched just these two entries — not a full regeneration, which would need real multi-lane CI data this session doesn't have access to. * fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists next's platform-conformance-tier classifier (#4591/#4598) landed after this branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's gen-platform-conformance-tier --check and both the Linux and macOS conformance suites. * fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation next's #4733 (landed after this branch's last rebase) replaced the static ISOLATED_HEAVY_FILES set with a threshold derived live from tests/test-timings.json, and pins the current derived result in EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still named codex-config.test.cjs, whose own weight this PR already dropped from 127783ms to 189ms (after splitting its heavy install()-pipeline blocks into codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The live-computed set correctly no longer includes it; the pinned expectation is updated to match. * fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it trek-e's review flagged two Major gaps: the thrown error read identically regardless of which of three real outcomes a runtime hit (nothing persisted / snapshot reverted / config left broken on disk), and no test proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/ Kimi/Cline — only Codex's revert path was ever asserted. assertConfiguredEntrypoints now tags each invalid entry with its actual consequence, mirrored from docs/how-to/update-gsd.md's existing rollback-matrix disclosure. A new test drives the same aggregate failure through Cline (whose own entrypoint is the one that fails) and asserts its hook file is still on disk afterward. * fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR One review pass (gemini-3.8-flash-high via the antigravity review lane) against this PR's full diff against next, findings independently verified against source before fixing: - Copilot's install() return object was the only one of 6 runtime branches missing rollbackInstallerMigrations — reachable now that this PR's own aggregate gate runs rollback across every result on any runtime's entrypoint failure, not just Copilot's own. - buildHookCommand's unresolved-bash early return skipped track() entirely, so a win32 install with no Git Bash silently produced an unregistered .sh hook instead of the 'unresolved-interpreter' validation failure configuredEntrypointsForHook's own comment said it would. - release-tarball-smoke.cjs reported a Cycle 4 install failure under SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing SMOKE.INSTALL_FAILED. - SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's trailing args, which also truncated any configDir containing a space (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan. Anchored the match on the already-known configDir prefix instead of a generic absolute-path guess: removes the ambiguity outright rather than patching the character class, and stays a raw-text scan on purpose (it catches a writer that emits a path without registering it — a JSON.parse of the expected schema would miss exactly that case). One suggested finding (test-timings.json "missing" the new test file) was verified false — that table only holds measured CI timings, populated after a file's first real run — and one Ponytail suggestion (a JSON.stringify dedup key) was rejected as it would reintroduce a real, if narrow, key-collision risk for no benefit. * fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment changeset-lint requires pr: to match the PR it runs on (16 on the fork, not the eventual upstream number) — rehearsal-branch convention already established earlier in this PR's history. lint-docs-guard-registration's quote-pairing heuristic doesn't require the docs/ path itself to be quoted — it flags a file once ANY quote-delimited span containing "docs/" appears anywhere in it, alongside any real fs read call. A comment ending "...update-gsd.md's rollback-matrix paragraph" supplied the closing quote character (the possessive apostrophe) the heuristic paired with an unrelated single-quoted string earlier in the file. Reworded to avoid the unquoted apostrophe next to the path. * chore(#4249): point the changeset pr field back at the upstream PR Fork rehearsal (PR #16) is green; the real target for this changeset is upstream PR #4249. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
25d1cb916f |
fix(#4721): give worktree cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged (#4766)
* fix(#4721): give cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged `worktree cleanup-wave` ran `git merge --no-ff` under the module-wide DEFAULT_GIT_TIMEOUT_MS (10 s) that is sized for plumbing calls. The merge is the one call in the wave that runs user hooks, so a repo whose pre-merge-commit hook is a test-suite gate lost every code-bearing executor merge. Three things went wrong at once, each fixed here: 1. Budget. The merge now passes an explicit timeout — DEFAULT_MERGE_TIMEOUT_MS (10 min), overridable via deps.mergeTimeoutMs. Every other git call in the wave keeps the module default; the shared constant is untouched, because every other caller is exactly what its 10 s comment describes. 2. Reason. A merge that does time out blocks on `merge_timed_out`, and its stderr names the budget and says the hook may still be running, instead of `merge_failed` carrying whatever the hook had printed before git was killed — which made a healthy executor branch look broken. 3. Residue. A merge killed during its hook has already staged the merged tree into the primary's index but never wrote MERGE_HEAD, so `git merge --abort` finds nothing and repoRootStillMidMerge (#2852) reads the primary as clean while the executor's whole diff sits staged against the old HEAD; a `git commit` from that state squashes the executor's history into one parent. After any failed merge the wave now reads `git diff --cached --name-only`; anything staged is the merge's own (git refuses to start a merge when the index differs from HEAD), so it runs `git reset --merge` — restores exactly those paths, keeps unrelated unstaged edits — and re-reads. Restored paths are reported as WAVE_CLEANUP_WARNING.MERGE_RESIDUE_RESTORED and the wave continues; a still-dirty or unreadable index reports MERGE_RESIDUE_LEFT_STAGED and halts the remaining entries, the same repo-level carve-out an unfinished merge takes. Tests: five mock-driven rows (budget wiring incl. the deps override, the timeout classification with restore, the no-reset control for an ordinary refused merge, an unrestorable residue halting the wave, an unverifiable index failing closed) plus a real-git row that runs a sleeping pre-merge-commit hook under a 1 s budget and asserts HEAD unmoved, index and worktree clean, the executor branch intact — with the same fixture merging cleanly under the default budget as its negative control. Two existing #2852 rows gain a handler for the new post-failure index read. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * docs(#4721): add Fixed changeset Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * test(#4721): release the real-git fixtures with t.after, not try/finally The two real-git rows cleaned up their scratch repo in a `finally` block; this file's own convention for fixture teardown is the test context's `t.after(() => cleanup(dir))`, and the house PR ruleset flags `finally` in a test body. Behaviour-neutral. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): gate the residue restore on the timeout, re-apply a merge autostash, and correct the hook census Three findings from the pre-file adversarial review of the previous commit, each driven on real git before changing code: 1. A merge git REFUSED ("your local changes … would be overwritten") also leaves no MERGE_HEAD — and that refusal is exactly what a pre-existing dirty primary index earns. The residue restore read that index as the merge's own and `reset --merge`d the operator's staged work away (driven: a staged edit to an unrelated file was discarded and reported as "restored"). The restore now runs ONLY when the merge timed out; a refusal is an immediate exit, never a timeout, so on that path nothing is read or reset. 2. `merge.autoStash=true` lets a merge start on a dirty index by parking the work in MERGE_AUTOSTASH, which a killed merge never re-applies. `git reset --merge` moves that stash into the stash list; the wave now runs `git stash pop --index` afterwards (the outcome `merge --abort` gives an autostashed merge), and reports WAVE_CLEANUP_WARNING.MERGE_AUTOSTASH_UNRESTORED (path null) when the pop fails or the autostash state could not be read — the work stays in the stash, the index is clean, the wave continues. Because of this the reset runs on a timed-out merge even when the index reads clean. 3. The merge is not the only hook-running git call in the module: `worktree add` runs post-checkout and every ref update runs reference-transaction. It is the only call that runs the commit-family hooks, which is what the budget is for. Comments and docs say so now. Tests: the "ordinary merge_failed" control becomes the regression row for finding 1 (strict mock — a `diff --cached` or `reset --merge` on a refused merge throws), plus a mock row for the autostash pop (dirty and clean index, pop success and failure), and two real-git rows: a refused merge over pre-existing staged work leaves it byte-identical, and a killed merge under merge.autoStash restores the executor residue AND puts the operator's staged work back. The real-git hook now sleeps 4 s against a 1.5 s budget for margin on slow runners. The two #2852 handlers added earlier are removed — the residue read no longer fires on their path. Negative control: 4 of the 10 #4721 rows fail on the previous commit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): key the residue restore on a killed merge, and re-read the index after a failed autostash pop Two more findings from the continuation review, both driven: 1. An externally delivered SIGTERM leaves the same staged/no-MERGE_HEAD state as the timeout, and the seam reports it as exitCode null + signal with timedOut false — so the timeout-only gate skipped the restore on a state it was written for. The gate is now "killed": timedOut, or a null exit code with a signal. A refused merge still exits with a code and is still never touched. The reason stays merge_failed for a signal kill. 2. A failed `git stash pop --index` keeps the stash entry but can leave conflict entries (UU) and partially applied paths, after which the next merge fails on "you have unmerged files"; the code returned halt:false on the strength of the pre-pop recheck. The index is now re-read after a failed pop and a dirty result halts the wave as merge_residue_left_staged alongside the merge_autostash_unrestored warning. Also driven and now documented rather than changed: a kill that lands once MERGE_HEAD exists (inside commit-msg) is the ordinary #2852 abort path — `git merge --abort` restores the tree and re-applies an autostash itself, unstaged, as git does for any aborted autostashed merge. Tests: the pop-failure mock row now asserts the post-pop re-read and gains a conflict-leftover variant that halts; a signal-kill mock row; a real-git row with the sleeping hook moved to commit-msg (timed out, no residue warnings, MERGE_HEAD cleared, primary clean). 414 pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): key the kill gate on the seam's signal, not on a null exit code The shell projection seam normalizes a signal death to exitCode 1 and carries the signal alongside (`_spawnResult`: `result.status ?? 1`), so the previous `exitCode === null && signal` gate could never fire in production and the unit row that covered it modelled a shape the seam does not emit (caught in the round-3 review). The gate is now `timedOut || signal`; a refused merge exits with a code and no signal. The mock row uses the real shape, and a mocked spawnSync signal death driven through the compiled seam reaches `reset --merge` and reports the residue restored. Also: three comments that still said "at its budget" / "runs user hooks" / "the index is clean", and the CLI-TOOLS sentence that reserved `merge_failed` for refusals and conflicts, now name the signal case. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * chore(#4721): set changeset fragment pr to 4766 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0967358b8b |
enhance(#3638): render bracket phase IDs on progress, stats, manager and statusline surfaces (epic #612 PR-5) (#4111)
* enhance(#3638): render bracket IDs on display surfaces Gate progress, stats, manager, and statusline projections on the bracket convention; validate phase_id_convention and single-source the convention card. Forward note: the uat.cts bracket co-change remains deliberately deferred to its owning slice. * chore(#3638): point the changeset at PR #4111 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3638): close bracket display review gaps * docs(#3638): register phase display modules * chore(#3638): re-trigger CI after macOS shard SIGTERM `full test (macos-latest, 24, shard 3/3)` failed on 20ce98cd1 in `tests/lint-compiled-artifact-sync.test.cjs` — the spawned `scripts/lint-compiled-artifact-sync.cjs` was killed at 60024ms (`exited null (signal SIGTERM)`, stdout and stderr both empty), 24ms past the test's own `TSC_COMPILE_TIMEOUT_MS`. That is the failure mode the constant's comment already documents ("under CI shard load that compile can exceed the budget, dying to a SIGTERM with empty piped stdout"). No content change; this empty commit exists only to re-run the matrix, since re-running a job needs write access on the upstream repository. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
2f0e99f9e0 |
fix(#4377): opt in to project-relative includes for local installs (#4425)
* enhance(#4377): opt-in project-relative includes for local installs A local install wrote the includes that point at GSD's own files as absolute paths — whatever the installer resolved at install time. For one checkout that is invisible. Across git worktrees it is not: each worktree gets its own .claude/ copy, but all of them point back at the checkout that ran the installer, so a worktree runs its own gsd-tools.cjs while reading workflow prose from a different checkout. Update that one checkout and every other worktree is running new instructions against an old engine, with nothing to stage the update with. --relative-includes (or GSD_RELATIVE_INCLUDES=1) makes a local install emit `@.claude/gsd-core/...`. Opt-in, and staying opt-in: absolute works for a single checkout, which is most people, and flipping the default would change every existing local install to solve a problem those users do not have. The prefix is the runtime's own localConfigDir descriptor value, never a literal — the same value resolveScope joins onto the cwd to produce the install target, and the same one the rewrite engine already uses for its ./.claude/ -> ./<dir>/ substitutions. Copilot and Antigravity have shipped this shape for local installs since they were added, with hardcoded .github/ and .agents/; this is that behavior, derived rather than written down. Six seams compute a path prefix and all six had to be threaded, which is why the opt-in travels through the environment the way --portable-hooks already does: one variable they all read cannot fall out of sync the way six signatures can. The launcher shim deliberately keeps its ABSOLUTE fallbacks. It probes gsd-tools through ${CLAUDE_CONFIG_DIR:-$HOME/.claude} and one such default per runtime; those are shell word expansions, not includes, and a relative value there resolves against the shell's cwd rather than the project. Trading an include that points at the wrong checkout for a path that points at nothing is not a fix. All three rewrite paths mask ${VAR:-default} spans before substituting and restore them after, and the mask only runs when the prefix is relative, so an absolute install is byte-for-byte unchanged. Every unexpressible case falls back to absolute: no opt-in, a global install, a missing dir name, the configHome.kind === 'none' sentinel, an absolute descriptor value, or one climbing out of the project with '..'. * chore(#4377): add changeset for project-relative local includes * fix(#4377): compare against POSIX-normalized roots in the install e2e arms The emitted prefix is POSIX-normalized by design — it is substituted into markdown @-references, which use forward slashes universally, so a backslash would leak into shipped content (#1615). The e2e arms compared against the raw temp root, which on Windows is `D:\a\...` and appears in no emitted file. That reddened the control arm on the windows shard, and it was worse than a red: the negative arm ("nothing references the checkout") was passing VACUOUSLY there, because a string that cannot occur is trivially absent. Both now go through the same normalization, so the Windows lane asserts what the Linux lane does. * fix(#4377): tolerate a resolved temp root, and make the e2e diff self-diagnosing Two changes, one confirmed and one to stop guessing. Confirmed: the emitted content carries the RESOLVED root, not the spelling mkdtemp handed back. Reproduced on Linux with a symlinked install root — 236 emitted files carry the realpath, zero carry the link path. macOS has this structurally, since /var is a symlink to /private/var. Comparisons now go through both spellings, or the negative arms pass vacuously: "nothing references the checkout" is trivially true when the string being searched for cannot occur. Not confirmed: the macOS shard reported ~every workflow file differing in the "differ ONLY" arm while the five arms around it passed, and the assertion printed a list of filenames — which says a difference exists somewhere across 236 files and leaves the reader to guess which bytes. I cannot reproduce that platform locally, and guessing turns one CI round-trip into four. The assertion now reports the first divergence as text: the file, the byte offset, and a bounded window of both sides. * fix(#4377): strip the longest root spelling first in the install e2e diff The macOS failure was my test corrupting its own comparison, not a product defect. /var/folders/…/X is a SUBSTRING of /private/var/folders/…/X, so stripping the unresolved spelling first matched inside the resolved one and left the /private prefix glued to what followed: @/private/var/…/X/.claude/gsd-core/… -> @/private.claude/gsd-core/… a string present in neither install, which is why all 236 files "differed". Sorting the spellings longest-first consumes the whole occurrence, and the short form then has nothing left to match. Proven in isolation on the exact macOS shapes: short-first yields @/private.claude/…, longest-first yields @.claude/…. The self-diagnosing assertion added in the previous commit is what found this — it named the file, the byte offset, and printed both sides, so the corrupted string was visible rather than inferred from a list of 236 filenames. Keeping it. * fix(#4377): address review findings * test(#4377): scan nested shell defaults without regex backtracking * fix(#4377): close relative include review gaps * fix(#4377): preserve root-target runtime includes * fix(#4377): guard project-root relative includes * test(#4377): normalize Cline fallback roots * fix(#4377): persist relative include style --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
06845717fe |
feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition Failing-first coverage for the Loop Host Contract role partition. At this commit crossCheckRoleFamilies does not exist, so the rows throw "crossCheckRoleFamilies is not a function" -- the RED proof they bind to behavior rather than restating it. ADR-894 section 3 assigns roles per step but parenthesises the assignment as "(illustrative roles)", and nothing enforced it. The only thing standing in the way was a single deepEqual in this same file, which is editable prose. Rows cover: each step's own family accepted; a strict subset accepted; a foreign role rejected at every step; an unknown role rejected; an unknown step failing CLOSED; capitalization not silently matched; every offending role reported rather than only the first; and purity, because buildContract puts the same array into the generated contract. Two rows exist because an earlier cut of this suite was vacuous. The purity fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot fail an in-place sort(), and the mutant was being killed by three unrelated rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the exact same role-name domain, both directions: they are parallel constants over one domain, so divergence is the generative-fix class CLAUDE.md names. Every negative row asserts the offending ROLE NAME and the STEP NAME appear in the message. A count-only assertion survives a mutant that reports the wrong role, which the 80% Stryker gate would surface only after a full CI round-trip. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#4740): reject a cross-family agent-role declaration Orchestration and execution are distinct functions of the loop and must not drift into one another. That partition was real but unenforced: ADR-894 section 3 calls its own role assignment "illustrative", and the generator accepted anything. Adding orchestrator to execute-phase.md's agent-roles line compiled, --check passed once regenerated, and capability-validator.cjs then began accepting into:"orchestrator" at every execute point. ROLE_FAMILY maps every role to one of orchestration, planning or execution. EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family. crossCheckRoleFamilies rejects a cross-family role, a role outside the vocabulary, and an unknown step. It reports every offender, not the first. It fails CLOSED on an unknown step, deliberately diverging from assertPointsCoverage's "unknown step -- caught elsewhere". For points that is true: the canonical-set and duplicate checks catch it. For roles there is no second net, so failing open would leave an unknown step as the one input that bypasses the gate. crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles to agent FILES and the orchestrator is the host, owning none -- admissibility and agent-file presence are separate concerns with separate checks. Additive to section 3's existing rule that contribution.into must be a member of the step's agentRoles, which is unchanged. That governs what a CAPABILITY may target; this governs what a WORKFLOW may declare. No capability is affected, and all five workflows already declare single-family sets, so the gate is green on the commit that introduces it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4740): make the ADR-894 role assignment normative Section 3 parenthesises its per-step role assignment as "(illustrative roles)". That word was accurate about the list's PURPOSE -- it illustrated the shape of a generated contract entry -- and wrong about its STATUS, because the assignment was load-bearing from the moment the generator consumed it. Read literally it makes the partition an example rather than a rule. Appended as a dated in-place section per docs/contributor-standards.md, which records that an accepted ADR is never rewritten and names this the default pattern. Section 3's original body is untouched. The amendment states the three disjoint families, the one family each step admits, that a step may declare a strict subset but never outside it, and why this is a clarification rather than a new decision: the contract is generated from the workflow markers "so it cannot drift into a lie", and all five workflows have always declared single-family sets. What was absent was any statement that it is required, and any check that it holds. It also pins the distinction that is easy to re-merge: contribution.into being a member of agentRoles governs what a CAPABILITY may target and is unchanged; the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary entry for the Loop Host Contract records the same, beside the agent-reference drift guard it already documented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): backfill changeset pr number Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both go from invalid_pr(0) to ok. Without that env both report success without evaluating the branch at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): stop injecting the orchestrator procedure into executors claude-orchestration declared a contribution at execute:wave:pre with into:"executor". loop-hook-dispatch.md defines a contribution as "inject fragment.inline verbatim into the context for the role named in into", so its 267 lines were injected into EXECUTOR prompts whenever the capability was enabled. Those lines are orchestration end to end -- construct a wave manifest, resolve the dispatch backend, invoke the Workflow tool to spawn executors, bridge per-agent results into the merge chain. An executor can act on none of it. Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT carries no orchestrator entry by design: the orchestrator IS the host, and the host's procedure lives in execute-phase.md. A step's agentRoles enumerates agents a capability may inject context INTO, so adding orchestrator there would model the host as an injectable agent -- the same category error pointed the other way, and it would need an exception carved into the partition the same issue just made normative. So the defect is the mechanism, not the label. A contribution injects into an agent's context; "replace step 3's inline dispatch loop" is a change to what the HOST does. The contribution channel was serving as a host-behaviour directive because it was the only channel available at an execute point. The entry is removed. plan:post into:"planner" is correct and untouched. The procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the capability -- it is the only copy in the repo -- and is no longer injected anywhere. Consequence, not softened: the Workflow backend now has no loop wiring. Detection, emission and config remain and the design is intact, but nothing dispatches it. Under the separation ADR-1143 itself asserts it never had a legitimate channel; ADR-1143's own audit already records the end-to-end path has never been exercised. Wiring it properly needs a host-level mechanism that does not exist today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): invert the stale execute:wave:pre registry assertions Removing the contribution left four surfaces asserting or describing the old state. Caught by an isolated review before a verification run was spent, which is the point of reviewing first: the first of these was a guaranteed CI red. execute-wave-post-gate-pipeline-e2e asserted against the REAL generated registry that byLoopPoint['execute:wave:pre'] held exactly one contribution with capId claude-orchestration. It now holds zero. Inverted to assert exactly 0 -- not a vague >= 0 -- and the #2285 comment above it now explains the current state rather than the one it was written for. CONTEXT.md's Claude Orchestration entry claimed two contributions at wired points. It is now one, and the entry's execute:wave:post label was already wrong before this change: the manifest said execute:wave:pre. Rewritten to one plan:post contribution, why the execute-point one was removed, and where the procedure now lives. One assertion in claude-orchestration.test.cjs could not fail. It tested for the prose "(into the executor)" while the doc says "(`into: executor`)", so no plausible wording matched it and the paired plan:post assertion was carrying the row. Replaced with a check on the structural claim, and proved RED by restoring the two-contribution wording before reverting. The moved procedure keeps section headings that speak as a live contribution -- "When this contribution is active", "Why execute:wave:pre". Preserving the body verbatim was deliberate, so the headings stay and an editor's note under the header explains why they read that way. A sweep of all 17 files referencing byLoopPoint found no further siblings: the remaining hits are a synthetic capability fixture and an empty-points test that already expected no active hooks, both correct before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0763326ced |
fix(#3780): serialize WINDOWS.md ledger mutations on a cross-process lock (#4681)
* test(#3780): regression tests for parallel ledger-writer loss * fix(#3780): serialize WINDOWS.md mutations on a cross-process ledger lock * fix(#3780): keep the ledger-unavailable degrade contract intact under the lock wrapper * chore(#3780): backfill changeset PR number (4681) --------- Co-authored-by: sim <sim@local> |
||
|
|
bbdf7e8e84 |
chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold. THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser, not grep) found what the epic never enumerated: ADR-4650 named seven containment implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13 files, several guarding a write or an `fs.rmSync`. Two verified by reading rather than pattern-matching — `research-store.cts` comments its own as "ensure the resolved file path stays inside the store dir" immediately before a write, and `capability-lifecycle.cts` gates `fs.rmSync` with one. So the epic's Done-when "one containment predicate, used at every site" was FALSE when Phase 3 reported it satisfied. It is true now: the rule is clean across src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist. WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose first argument is a managed root and whose later arguments derive from argv. That is a taint analysis over 2046 call sites, in ESLint, without type information; "derives from argv" is not locally decidable. Any approximation either floods or is trivially evaded, and a rule that fires on hundreds of correct sites earns an allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is the COMPARISON, not the join, and that has one recognizable shape. Arm 1 X.startsWith(Y + sep) the hand-rolled containment idiom Arm 2 a containment predicate called as a bare statement, answer discarded Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper was called". Its example `validatePath(x, root).resolved` is already structurally impossible — Phase 3 un-exported `validatePath` — so the remaining expressible failure is ignoring the answer, which is the defect that recurred five times in this epic. The census found exactly one live instance (`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers stop re-deriving the path the comment above it was extracted to stop them re-deriving. The rule deliberately does NOT try to catch validate-one-path-use-another where the answer is used but a different variable flows onward. That needs flow analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the two are complementary. PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a lexical site onto the realpath family broke four tests and was caught only by the matrix. Every migrated site was triaged individually. The six installer-migrations tree-walks and the six capability-lifecycle gates take the LEXICAL family because their operands are already realpath-resolved and they deliberately treat the final component as a link; boundary sites take realpath. TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken. `installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT `target === root` by contract, while the canonical comparison ACCEPTS it. Swapped naively, a migration could `rmdir` the user's config root and a third-party descriptor could write at configHome itself. Both keep `=== root` as an explicit additional arm alongside the predicate call — the predicate decides containment, the call site keeps its own extra condition (ADR-4650 decision 6). ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was byte-identical to `isContainedIn` and said so in its own docstring. `isContainedIn` is now exported for callers that have already resolved both operands and need only the comparison, with a doc note that a caller which has NOT resolved them must use a full predicate instead. THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable reason. Two justifications: (a) not a containment decision — an ancestor-walk loop condition, sub-repo grouping, worktree identity matching, declared-path coverage; (b) it IS containment but the canonical predicate is unreachable — `capability-validator.cjs` is a committed pre-build `.cjs` and the compiled `security.cjs` is untracked build output, so requiring it would break a fresh clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure. The marker was renamed from `allow-lexical-prefix-match` mid-phase because that name asserted only (a) and would have stated something false at the (b) sites. A marker suppresses BEFORE the violation counter increments, so a file whose every occurrence is marked still reports `staleAllowlistEntry` — otherwise a drained entry lingers and silently re-permits the site later. DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/` file made `npm run lint` fail with the rule's full guidance message; removing it returned the tree to clean. Both halves recorded — red alone proves nothing, since a rule red for an unrelated reason looks identical. DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two distinct rejection messages and two manual realpath calls; routing it through `tryWithinRoot` collapses them to one message, and a missing module now surfaces as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The security property is preserved and slightly strengthened — the candidate is realpathed and containment re-checked, and the dangling-symlink oracle closure comes along with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4654): record the containment ratchet in CONTEXT.md and the security model Both entries previously described the seam without the thing that keeps it a seam. They now state what the rule bans, and — more usefully for whoever reads this next — what it deliberately does NOT attempt: deciding per path.join call whether an argument came from user input. That question is not locally decidable, and an approximation across ~2000 join sites would earn an exemption list of hundreds, which is the opposite of a ratchet. Also records the marker's two legitimate justifications and that its reason is mandatory, so the escape stays reviewable rather than becoming a mute button. Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json regenerated and confirmed byte-identical rather than assumed — which also confirms eslint-rules/ is not a shipped path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): close review findings and the two matrix failures MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my evidence for collapsing it was wrong. I searched tests/ for the literal string "resolves outside its install root", found nothing, and reported that no test asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module pointing OUTSIDE the install root is not loaded" — the test guarding the exact property I claimed was preserved. defaultRequireFromInstallRoot now does both checks again with both messages byte-identical, each routed through the canonical predicate, which is better than the original since that hand-rolled both comparisons. MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments, so a suppression marker inside a plan body drifts the baseline exactly as an edit does. Measured: with markers in place, two of the four still differed from their committed checksums. The four shipped bodies are now byte-identical to next, and the rule's config excludes those four paths BY NAME rather than by a directory wildcard, so a NEW migration is still covered. Six containment comparisons stay un-ratcheted there; that gap is recorded in the rule's Known gaps, in CONTEXT.md and in the security model rather than left implicit. Justification (c) is removed from the marker's documented reasons, because a marker was proven unable to express it. ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT shape while permitting the incorrect one: startsWith(root) with no separator is the genuinely unsafe form, since it accepts a sibling such as root-evil, and my own test blessed it as valid. Flagging every bare startsWith would swamp the rule, so that stays a STATED gap rather than a silent one. Closed for real: the template-literal spelling, which the census never saw because it only inspected plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated. A separator reached through a const alias is now resolved via scope analysis. And isContainedIn, exported in Phase 3, was missing from the discarded-result set, so a bare no-op call went unflagged on the one function the epic funnels through. SECURITY REVIEW — the marker could over-suppress two ways: a block comment worked identically to a line comment, and one marker silently covered every violation sharing its line. It now requires a Line comment positioned after the flagged node ends, so it anchors to the node it trails. Four sites had dropped an unreachable-but-deliberate equality rejection against the root; each is restored as the call site's own arm. eslint.config.mjs still documented the OLD marker token, which my rename missed — it would have sent the next author in circles. A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a stale eslint cache while twelve real violations existed. Every lint check here now clears the cache first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4654): anchor a suppression marker to the violation it actually trails The matrix caught this; my own test caught it, on its first execution. The case "two violations on one line: trailing marker suppresses only the one it trails" expected 1 error and got 0 — both were suppressed. ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose range started at or after the node's end. A trailing marker at the END of a line sits after EVERY node on that line, so that condition held for all of them. "After the node" does not identify WHICH node the marker trails. The fix reads as correct and is not. FIX: deferred reporting. Violations accumulate during traversal instead of being reported immediately; at Program:exit each marker claims exactly ONE pending violation — the one on its line whose end is nearest before the marker begins — and every unclaimed violation is then counted and reported. One marker, one suppression. An earlier violation sharing the line is still reported, which is the property the security review asked for and the previous attempt only appeared to deliver. The counter now increments at flush time rather than during traversal, so a suppressed occurrence still does not keep an allowlist entry alive. AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test` is hard-blocked here, so this rule's test file could only ever be executed on the remote matrix — which is why a broken anchoring shipped into a run. ESLint's programmatic Linter API is not a test runner, and exercising the rule through it verifies every case locally in seconds. All 24 now pass locally, including the two-on-one-line case that failed remotely. That loop should have been built before the rule was first sent to the matrix rather than after it failed twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs now carries both, with the how-to quadrant skipped for a stated reason rather than an empty field. The audience for this deliverable is a contributor who trips the rule, and the task-oriented guidance reaches them in the ESLint message itself — which names the correct predicate, says how to choose between the realpath and lexical families, cites the Phase 3 regression caused by choosing wrong, and gives the marker syntax. A docs/how-to page would be a second, driftable copy read by nobody at the moment of failure. enablementSequence is recorded as what it actually is: a VERIFICATION sequence, not an enablement one. The rule is never off, so there is no off-to-on transition to describe. scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not evaluate against the mandated pr:0 placeholder. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d1d9ee85a5 |
feat(#4142): thread convention through the completion-path membership seam (#3644)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
bbc3f131be |
refactor(#4653): make containment ONE decision, resolved two ways
Satisfies #4653 DW1 and DW9, which were the phase's outstanding acceptance criteria: every other implementation must be deleted or route its containment DECISION through the canonical predicate, and no surviving wrapper may decide WHETHER a path is contained. Three implementations were being retained with their own comparisons, on the argument that each needs LEXICAL resolution — a realpath-based predicate is the wrong tool wherever a symlink must be preserved rather than resolved. That argument is correct about RESOLUTION and was being used to justify owning the DECISION too. Those are separable, and separating them is what closes the criteria honestly rather than by reinterpretation. isContainedIn(resolvedTarget, resolvedRoot, pathImpl?) module-internal is now the single place this repo decides containment. It is separator-aware, so a sibling merely sharing a prefix (`<root>-evil` against `<root>`) is still rejected. Two exported families sit on it and differ ONLY in how a candidate is resolved before the decision: assertWithinRoot / tryWithinRoot realpath-resolving assertWithinRootLexical / tryWithinRootLexical path.resolve only, no I/O The lexical pair carries `opts.pathImpl`, so win32 separator semantics stay testable off Windows — that seam already existed in isPathConfined and would have been lost by a naive collapse. The three call sites now take their decision from the predicate and keep only what is genuinely theirs: external-descriptor-trust isPathConfined delegates outright; pathImpl forwarded installer-migrations ensureInsideConfig delegates; keeps its own message and its LEXICAL fullPath, which callers consume for existsSync and journal rows gsd-tools.cjs isInsideDir delegates; keeps its own `target !== root` condition, and the separate symlink refusal above it stands DW5 is not weakened by this. That criterion binds the symlink oracle and the ancestor canonicalization; both are untouched. The only change inside validatePath is three comparison lines becoming one call, and the rejection string `Path escapes allowed directory: <resolved> is outside <base>` stays byte-identical because it is an observable CLI contract. What this does NOT do, stated plainly: the lexical family still cannot see a symlink. That is a property of lexical resolution, not a gap in the seam, and the three callers that need it are the three that must pair it with their own symlink refusal — which is exactly what the fix earlier in this phase added at the install sites. The doc comment says so at the definition, and CONTEXT.md and docs/explanation/security-model.md are corrected: they previously described these three as deliberately NOT routed through the predicate, which is no longer true. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6f0e5ccf85 |
fix(#4636,#4653): close the symlink hole, revert a wrong collapse, fix six review findings
The RED checkpoint and two orthogonal reviews found eight defects. All fixed here.
THE COLLAPSE THAT WAS WRONG — installer-migrations. Routing ensureInsideConfig's
containment decision through the realpath-based canonical predicate broke four
tests, and the failure message says it plainly: "migration path escapes
configDir: extensions/gsd.cjs". That module's entire contract is that a
symlinked managed path is snapshotted, restored and backed up AS A LINK and
never dereferenced. The canonical predicate dereferences, then rejects the
result for escaping configDir — so it destroys exactly the thing the module
exists to preserve. Reverted to lexical, with the ruling recorded above the
function so it is not collapsed a third time. normalizeRelPath is the real
pre-gate there; it throws on absolute paths and '..' before this check runs.
That makes THREE deliberately-retained implementations, not two, and they share
one shape worth naming: a realpath-based predicate is the wrong tool wherever a
symlink must be PRESERVED rather than resolved. CONTEXT.md and
docs/explanation/security-model.md are corrected — both previously described
ensureInsideConfig as collapsed.
THE MISSED CONSUMER. tests/security-prompt-injection.security.test.cjs
destructures validatePath from the compiled lib; un-exporting it turned five
tests into TypeError. It appeared in my own earlier search output and I did not
follow it up. Translated under the same rule as the rest: assertions on the
rejection REASON go through assertWithinRoot, boolean-only through
tryWithinRoot.
VALIDATE-ONE-PATH-USE-ANOTHER, FOUND TWICE MORE. This is the fourth and fifth
occurrence in this epic of the exact defect it exists to prevent.
- scripts/check-glossary-refs.cjs decided containment on `token` and then
stat'd a separately re-joined path.join(ROOT, token). The ContainedPath is
now carried through to the probe, so the validated value is the probed one.
- src/init.cts computed skillPathContained and DISCARDED it, re-joining from
the raw input for the existsSync and read. The branded type exists to make
that a type error and here it was inert.
AND THE OVER-CORRECTION OF THAT FIX, caught before it shipped. The first attempt
also substituted the validated value into the EMITTED `ref` for a global skill.
That value is a display token, not a path anything reads through — the only fs
access in that branch runs on the lexical path beforehand — so substituting it
changed emitted output two ways: it is realpath-resolved, so a symlinked global
skills directory would have emitted its resolved target instead of the user's
own path, and it came from path.join, so Windows would have emitted a backslash
where the template has a literal '/'. Restored, with the distinction recorded:
the containment check there is a GATE, not a path producer.
A TEST THAT COULD NOT FAIL. The first symlink regression planted its symlink
from inside a hooked fs.readdirSync and never asserted the planting happened —
if the hook did not fire, the "nothing was written outside" assertion passed
trivially, green against vulnerable code. It now asserts the plant, matching its
sibling. The other two were re-checked: one already asserted its equivalent, the
other plants synchronously and cannot silently no-op.
THE SYMLINK FIX ITSELF, now that the tests are proven red on the matrix.
isPathConfined is lexical by design and structurally cannot see a symlink; three
callers relied on it with no defense of their own. install-engine.cts:1608 and
install-profiles.cts:880 refuse to mkdir/write through a link — mkdirSync with
recursive:true does NOT throw on an existing symlink-to-directory, so a planted
link redirected the SKILL.md write outside the install root.
install-profiles.cts:755 refuses to read through one — statSync FOLLOWS links,
so an outside file's contents were returned and installed as a skill body. Each
mirrors the guard retired-artifact-cleanup.cts:77 already uses.
Severity stated accurately rather than dramatically: only the read at :755 needs
no race. _removeGsdEntries sweeps a pre-planted link at :1608 before the write
loop, and :880's stageDir is a fresh mkdtemp, so both of those require winning a
window. They are fixed as defense-in-depth, not as live exploits.
ALSO: the Changed changeset claimed "every command's observable behavior [is]
unchanged". Three rejection messages are reworded. It now says so, and says that
none of them reveals a host path it previously hid. A stale comment in
verify.cts still named validatePath; an init.cts warning hardcoded "resolves
outside the project directory" for a check that also rejects absolute paths, NUL
bytes and empty strings; and the rationale deleted with check-glossary-refs'
retired helper is restored, noting honestly that a rejected token is now
realpath-resolved before rejection rather than rejected by string comparison.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5967f1939a |
docs(#4653): record the path-containment seam in CONTEXT.md and add the changeset
The glossary had no entry for the containment predicate at all, which is the epic's actual deliverable. The new entry states the engine/export split, the branded type, the named acceptance policy, the preserved message contract, and — the part most likely to be undone by a later cleanup — the two implementations deliberately NOT collapsed and why each is stricter or narrower rather than duplicative. Glossary gate re-run: 269 refs, exit 0. Install-tree goldens regenerated and confirmed byte-identical rather than assumed unchanged; lint:ci exits 0, so no conformance-tier drift either. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
374300da17 |
fix(#4652): confine every boundary that joins argv to a managed root
Phase 2 of epic #4636, absorbing #4327 and #4354. Implements ADR-4650 decision 3: containment is a boundary concern — the predicate runs where external input enters, not at whichever interior call site remembered. Four boundaries now validate against their managed root and reject with a USAGE-shaped error before touching the filesystem: todo complete <name> -> todosDir(cwd) check predicate --phase-dir <dir> -> projectDir check decision-coverage-plan <dir> -> projectDir (via resolvePath) check gap-analysis.plan-post <dir> -> projectDir #4327 understated its own severity. It reports that a traversal name "resolves outside the todos root", which reads as an information leak. Measured, it was destructive: the command exited 0, MOVED the outside file into completed/, and unlinked the original. cmdTodoComplete ends in fs.unlinkSync(sourcePath), so an unconfined name consumed across the boundary rather than merely reading across it. Validation now precedes every fs call — existsSync, readFileSync, ensureDir, writeSync, unlinkSync — and both halves of the move are confined, so neither source nor destination can land outside the root. --dry-run is rejected on the same terms; a preview must not leak a resolved outside path either. #4354 reproduces exactly: a BLOCKING gate returned block:false sourced entirely from a SECURITY.md in a caller-chosen directory outside the project. THE HARDER HALF, found by the isolated adversarial review of the first attempt: validating a path and then using a DIFFERENT one closes nothing. The first fix validated `--phase-dir` joined against `--cwd`, then passed the RAW unjoined value into the predicate context. gate-predicate-evaluator uses it as-is and findPhaseArtifact resolves a relative path against the REAL process cwd — so validation and the read used two different roots whenever process.cwd() differed from --cwd. Reproduced: running from a directory holding a plan with `secret_field: LEAKED_VALUE`, a predicate declared against an empty --cwd project exited 0 and returned "actual":"LEAKED_VALUE". The rule now applied at all three router sites: **use the validated resolved path, never the raw input.** Independently re-verified after the fix — the lookup resolves in the --cwd project and no value leaks. gate-predicate-evaluator.cts is untouched and still imports no fs. Confining in the router is what keeps that pure-leaf contract intact AND covers ${PHASE_DIR} interpolation into command-exit-zero, which an evaluator-local fix would have missed entirely. Also fixed, same review: `todo complete .` and `..` passed containment (they resolve to the pending dir, which IS inside the root) and then threw an uncaught EISDIR with an absolute-path stack trace. Now a clean USAGE rejection naming the real reason — "todo name is not a file" — rather than borrowing the escape message, which would have stated something false. Ripples discharged BEFORE the verification checkpoint rather than after, per the Phase 1 retrospective: docs/reference/gate-predicates.md and docs/CLI-TOOLS.md document the new constraints, CONTEXT.md records why containment lives at the router rather than the evaluator, the changeset is written, and the install-tree goldens were regenerated to confirm unchanged (no new shipped file) rather than assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
38e4ce5f62 |
fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
03738824de | enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) | ||
|
|
c3e2da153b |
fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names A first-milestone ROADMAP.md H1 that puts the version after the name (# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's own version token, which the ADR-3180 §7.2 pinned name rule returns as a COMPLETE-scope milestone name. Failing-first coverage: - getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only fallback) must yield TRUNCATED {version, name: null}, never ')' - the refusal is level-agnostic (H2/H3) - punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only) - listMilestoneHeadings enumerates the heading with name: null - init manager CLI reports milestone_name: null and no lone ')' anywhere - property (seed 20260905, 300 runs): a word-char remainder is always a name, a punctuation-only remainder never is - negative space: canonical delimiter forms, parenthetical names (#3171), trailing markers, digit-only names, CRLF headings, version-last-no-parens control * fix(#4134): refuse punctuation-only milestone heading names extractMilestoneHeadingName returns everything after the heading's own version token as the name (ADR-3180 §7.2 pinned rule), which assumes version-then-name. A name-then-version heading — the H1 a first-ever ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves exactly ')' after the token, and that fragment was returned as a COMPLETE-scope milestone name, propagating into init.* JSON output and buildStateFrontmatter's STATE.md writes. A remainder with no letter or digit anywhere (any script) is heading structure, not a curated name: refuse it as name: null so callers report the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names that merely contain punctuation are unaffected — '(' stays an ordinary name character (#3171) — and digit-only names qualify. Also closes the template gap that lets the shape occur: the roadmapper agent's output_formats now templates the version-free canonical H1 ('# Roadmap: [Project Name]', per templates/roadmap.md) instead of leaving a first milestone's title line to invention. The new section shifts the file's existing bare-gsd-tools prose mention from line 647 to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134) * chore(#4134): add changeset * chore(#4134): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
4c60879b5d |
fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources * chore(#4132): record PR number in changeset * test(#4132): cover rejected commands source alias * fix(#4132): reject aliased package fallback * test(#4132): cover rejected agents source alias * test(#4132): cover partially aliased marker provider * fix(#4132): reject partially aliased source providers * test(#4132): cover routed source identity probes * fix(#4132): route installed source identity probes * refactor(#4132): tighten installer source metadata * test(#4132): cover corpus trust boundary attacks * fix(#4132): close installed corpus trust gaps * refactor(#4132): keep installer authority private * fix(#4132): preserve private installer fallback * test(#4132): preserve fixture source authority * fix(#4132): reject overlapping source fallback * fix(#4132): avoid redundant installed corpus reads * refactor(#4132): simplify provider resolution * test(#4132): sync install tree fixtures after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
70f22e4643 |
fix(#4213): keep STATE.md progress surfaces synchronized (#4231)
* fix(#4213): keep STATE.md progress surfaces synchronized * fix(#4213): clamp the shared progress bar and keep bold-first priority, changeset + property tests - formatProgressMachineSegment clamps through clampPercentFromFraction (ADR-3180 Decision 7 kernel) with a 0 floor, so a hand-edited out-of-range persisted percent renders a clamped bar instead of throwing RangeError on repeat() inside the write seam - stateReplaceProgressPercent restores the #2177 bold-first priority: **Progress:** anywhere in the body wins; a plain ^Progress: line is the fallback, so free text starting with Progress: cannot capture the rewrite ahead of the real status line - cross-reference comment names the three consumers and the cmdStateSync sanctioned exception (ADR-3408 §8.3) - CONTEXT.md: applyPostSyncPreservation reconciliation documented in the STATE.md Transition Module entry - property tests (never-throws/well-formed, idempotency, round-trip, bold-first) + two regression rows through the CLI --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5869febb16 |
enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change readVerificationStatus() now recomputes a deterministic sha256 fingerprint over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY, requirements, implementation files in the verified change set) and returns stale on any mismatch, fail-closed when a covered file is missing, unreadable, or escapes the project root. Legacy reports with no fingerprint metadata keep the prior SUMMARY-mtime staleness check unchanged. The verifier computes covered_digest via the new verification.fingerprint CLI command rather than by hand, since a digest is deterministic math, not an LLM-estimated value. * chore(#4155): backfill fork PR number in changeset * fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap * fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior Partial fingerprint metadata (one of covered_files/covered_digest present, the other missing or malformed) now fails closed to stale instead of silently downgrading to the legacy mtime-only check. computeCoveredDigest also canonicalizes with realpathSync before re-confining, so an in-root symlink whose target escapes the project root can no longer produce a matching digest. gsd-verifier.md restores the completeness requirement and checklist item trimmed by the earlier size-budget fix, within the LARGE tier byte cap. * chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap * fix(#4155): address gemini adversarial review findings computeCoveredDigest now threads the caller-supplied opts.fs seam through its confinement and read paths instead of always using raw node:fs — a caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP 2, #2790 follow-up) was silently bypassed for covered-input reads. The project-root anchor itself still canonicalizes through real fs (it is a trusted value the caller derived, not attacker-influenced covered-input data); only per-file candidate reads go through the injected seam. Covered-file paths are now canonicalized (./ prefixes, redundant slashes, internal .. segments) before becoming dedup/sort/hash keys or confinement subjects — closes both a spurious-stale false positive (two spellings of the same file hashing differently) and a confinement gap (an internal .. segment that doesn't start the string). gsd-verifier.md now states covered-file paths are project-root-relative, not phaseDir-relative, closing an ambiguity that would have made a real verifier agent's first fingerprint invocation fail closed. defaultFsImpl's methods now late-bind through fs.<method> rather than capturing function references at module load — the earlier direct-capture form was invisible to existing tests' t.mock.method(fs, 'statSync', ...) seams, a real regression caught by the full suite (not the reviewer). * fix(#4155): catch a plan/summary added to the phase dir after verification but never declared The content digest only recomputes hashes for paths the verifier actually declared in covered_files — it had no way to notice a plan or summary added to the phase directory after verification if that new file was never declared, silently regressing behind the legacy mtime check it replaces (which scans the live directory, not a declared list). findUncoveredCurrentArtifact re-scans the live phase directory for every current *-PLAN.md/*-SUMMARY.md and requires each to be represented in covered_files, closing that gap; a directory scan failure fails closed to stale rather than silently skipping the check. CONTEXT.md's Verification Module entry corrected to describe the fingerprint path's stricter fail-closed FS-error contract (routes to stale) instead of the module's original degrade-to-safe one (missing / not-stale), which only the legacy path still keeps. * refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/ sorted covered_files independently — one shared helper now backs both (gemini review's ponytail-lens finding). Adds one CLI-to-readVerificationStatus test against a genuine .planning/phases/NN-x/ project with an implementation file outside .planning/ entirely, closing the review finding that prior #4155 unit fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot falls back to phaseDir itself there) and never exercised real multi-level path resolution. * fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/ Two independent review rounds (opus critical-reviewer + opus ponytail + agy, run twice) found two instances of the same fail-open class: - computeCoveredDigest's per-file reads routed through the caller's injected fsImpl. planning-inspect.cts passes a `.planning/`-confined containment fs into readVerificationStatus's opts.fs, so any covered implementation file outside `.planning/` (mandatory per the issue) made the confinement wrapper throw, which was caught and turned into a stale digest -- reporting every fingerprinted phase permanently stale via `planning.inspect`, regardless of actual drift. Per-file reads now always use real node:fs, matching the pre-existing treatment of root canonicalization; the realRel-vs-realRoot check is the real confinement boundary for this data and needs no seam. - allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans reports readdir failures via a `scope` field, it never throws), so an unreadable nested plans/ dir was silently treated as "zero artifacts, all covered" instead of failing closed. Now branches on scope !== SCOPE.COMPLETE. Also, per ponytail's second-round findings: reverted an unwarranted FINGERPRINT_VERSION bump and digest length-prefix from the first fix (no v1 digest has ever existed -- the feature is unreleased -- and the prefix closed a collision that grants no capability beyond what a writer of covered_files already has more cheaply); removed a verifier-facing escape-hatch instruction whose own example was a case that should trigger staleness, not bypass it; corrected CONTEXT.md references to the renamed allCurrentArtifactsCovered and a stale "unconditional" rescan claim; simplified the isStale derivation, removed dead FsLike members, and tightened test coverage. Regression tests for both fail-open bugs are included and were each confirmed to fail against the pre-fix code before the fix landed. full test suite: 2558/2560 pass, 2 skipped, 0 fail * fix(#4155): trim gsd-verifier.md under the LARGE size cap Fork CI caught what my local runs missed: the superseded/nested-plans instruction added earlier pushed gsd-verifier.md to 49299 bytes, 147 over the LARGE tier's 49152-byte hard cap (tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's wording and dropped a redundant inline comment tag; no content lost. * chore(#4155): point changeset at the upstream PR number pr: 19 was the fork PR opened for internal review-lane CI; now that open-gsd/gsd-core#4290 exists, the changeset field must match it per CONTRIBUTING.md's release-notes convention. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
91ed46882a |
feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives Adds the full behavioral (tests/quick-batch.test.cjs) and property-based (tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core primitives per the phase's 35-row test matrix — task-list parsing (inline + --file, with path-confinement/symlink-escape/non-regular-file rejection), collision-safe quick-id preallocation under withPlanningLock, BATCH.json schema/validation/resume, dependency-DAG + partitionByFileOverlap wave construction, and exactly-once STATE.md completion (including the STATE-row-written-but-manifest-not-yet-updated crash window). The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a not-yet-written src/quick-batch.cts) does not exist yet — every test in both files fails at the top-level require() before any assertion runs. Five fast-check properties cover collision-freedom under lock contention, resume idempotency, exactly-once STATE completion, wave totality, and DAG-respecting wave order, per the design doc's property-based-coverage requirement. * feat(#3675): implement quick-batch core primitives Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core primitives per the phase design lock — pure/state primitives and CLI-testable core operations only, no agent dispatch, no worktree creation, no user-facing command (Phase 4/#3676's job): - parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list parsing (>=2 items required) and a --file variant strictly confined to the planning workspace root via requireSafePath, rejecting non-regular-file targets. - allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id preallocation under withPlanningLock, checked against both on-disk .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json manifests (never on-disk-only, which would miss another in-flight batch that hasn't dispatched any real quick directory yet) — replicates cmdInitQuick's own grammar rather than delegating to it (that function's 2-second granularity is not batch-safe). - computeWaves: deterministic wave construction combining dependency-DAG layering with partitionByFileOverlap (#3674), called per DAG layer over path-separator-normalized planned_files — normalization happens at this module's boundary, never inside the Phase 2 helper. - loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated JSON, wrong types, missing fields, out-of-batch dependency references, dependency cycles, a worktree path absent from disk). - resumeBatch: skips complete items, never auto-retries failed items, propagates/reverses blocked status along the DAG to a fixed point, and detects a STATE.md row that already exists for a non-complete item (the "STATE written, BATCH.json not yet updated" crash window) — completing it without re-appending. Idempotent across repeated calls. - completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion — appendQuickTaskRow (unmodified) is called at most once per quick id, gated by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks Completed" table, since appendQuickTaskRow itself carries no idempotency. BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling of .planning/quick/ — never inside it, so scanQuickTasks never misreads a batch manifest as a broken quick task. * docs(#3675): register the new quick-batch module New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row plus the regenerated docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary entry matching the convention set by the sibling File Overlap Partitioner Module (#3674) entry it sits beside. NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically sorted single-entry insertion into families.cli_modules, matching the existing file's structure) rather than via `node scripts/gen-inventory-manifest.cjs --write` — this session's MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script path from Bash, with no available Memtrace tool to route through instead. The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs --check` to confirm this hand-edit is byte-identical to the generator's own output before merging. * fix(#3675): resolve lint findings in quick-batch primitives and tests Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color array, two unnecessary `as string[]` casts TS 5.5's inferred type predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch` import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a CONTEXT.md glossary illustration that looked like a real file reference. * feat(#3675): close acceptance-criteria gaps found in review Standards- and spec-axis review (plus a self-caught race) surfaced real gaps against issue #3675's own acceptance criteria and this repo's test conventions: - BATCH.json was missing options, base_revision, per-item wave, and per-item commit — the issue's AC explicitly lists all four as things the manifest must track. Added them: createBatch persists caller-supplied batchOptions/baseRevision verbatim and assigns each item its computed wave index; completeQuickItem now persists the commit onto the item, not just the STATE.md row. All four are backward-tolerant on load (an older/hand-built manifest without them still validates). - resumeBatch had no "incompatible base divergence" check at all, despite the AC and the ADR's own "Base divergence" section requiring one. Added an opt-in currentBaseRevision comparison that fails closed with a recoverable diagnostic on mismatch, and touches nothing on refusal. - resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the only durable write path in this module that wasn't lock-protected, a real lost-update race against a concurrent completeQuickItem or another resume. Now runs inside the same lock createBatch/ completeQuickItem use. - loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no size cap (security review, Low/informational); switched to the existing safeJsonParse (1MB cap) for defense-in-depth. - Parser (parseTaskList) had only example-based tests; CLAUDE.md requires a fast-check property test for parsers. Added one plus a companion reject-property for <2 items. - The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for direct limit-1/limit/limit+1 testing without needing 46k fixture dirs. - Issue AC explicitly asks for prompt-injection-payload test coverage, distinct from the existing shell-metacharacter test; added one. - Test row 9 (FIFO skip) silently returned instead of calling t.skip(), so an unsupported platform would report a pass rather than a documented skip; fixed to bind the test-context param and skip properly. - Extracted toWaveInput to remove a 2-site production duplication of the QuickBatchItem -> computeWaves reshape (Standards-axis smell). - Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/ is user-facing even though the compiled .cjs is gitignored). * fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason gsd-test caught this: switching loadBatch to safeJsonParse changed the parse- failure message shape ("... parse error — ...") without preserving the "not valid JSON" substring row 27's own test asserts on. Re-wrap safeJsonParse's error into the original diagnostic phrasing regardless of which of its three failure modes fired. * docs(#3675): backfill changeset pr number to 4190 * fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one CI caught this on windows-latest: row 9's platform-skip only caught mkfifo throwing (command not found). On this runner mkfifo resolves to something that exits 0 without creating a file (NTFS has no FIFO concept), so execution fell through to parseTaskListFromFile against a path that doesn't exist, producing an ENOENT stat error instead of the expected "not a regular file" rejection. Check the artifact actually exists before trusting a zero exit code, and skip with a documented reason either way. --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
b0572c0108 |
feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage golden script) as a regression safety net ahead of extracting partitionStages into a standalone module. Also adds the new module's unit and property tests (test matrix rows 1-11) against its expected public API, which does not exist yet and is added in the next commit. * feat(#3674): extract file-overlap partitioner into a shared, generic module Moves partitionStages' greedy first-fit file-overlap algorithm into a new, dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap), generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own Plan[] shape onto the generic input and back — behavior-preserving, no dependency ordering, no path normalization, no filesystem access moved or added. Enables a future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive without pulling in orchestration internals. * docs(#3674): register the file-overlap-partitioner module bookkeeping New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row (regenerated via gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry matching the convention set by similarly-scoped leaf modules (text-lines.cts, plan-dependency-graph.cts, spec-section.cts). * fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct - docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write, produced no diff (manifest is keyed by content, not row order) - tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is: tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier. * fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction The `no two plans in the same stage share a modified file` property reconstructs which physical item produced each output id via `remaining.findIndex(r => r.id === id)`. Under duplicate ids (an explicitly-supported input shape for `partitionByFileOverlap`) that reconstruction can pick the wrong physical occurrence, producing a false-positive overlap failure (observed counterexample: p0(f1), p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but misread by id-order as [[p0,p208#1],...], which do overlap). Properties (a) determinism and (b) totality already exercise duplicate ids correctly and are left unchanged; only this property's generated items are now constrained to unique ids via `fc.uniqueArray`, where the reconstruction is unambiguous. --------- Co-authored-by: sim <sim@local> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
7c116b1c17 |
fix(#3697): warn when the phase-complete Requirements-line tokenizer under-selects REQ-IDs (#3744)
* fix(#3697): warn when the Requirements line under-selects REQ-IDs
`cmdPhaseComplete` tokenizes ROADMAP's `**Requirements**:` line by splitting
on `[,\s]+` and keeping tokens matching the anchored REQ-ID shape. That is
correct for the canonical comma list the template ships, and silently wrong
for every other form:
`RANGE-01 … RANGE-05` -> the two ENDPOINTS only; the interior IDs are
never considered, yet `requirements_updated`
reports true with zero warnings
`RANGE-01…05` -> ZERO IDs; the whole line is inert
The silence is structural: the only cross-check, `ghostReqIds`, is itself
`citedReqIds.filter(...)`, so an ID the tokenizer dropped is invisible to it
by construction — and to `traceabilityWriteMisses` and `requirements_updated`
with it.
Warn on both paths. This does not add range support: the selected set is
unchanged, so no existing ledger write changes. The trigger is ID-SHAPED
EVIDENCE only — an ID-shaped substring the tokenizer did not select, or a
range operator joining two IDs — with parenthetical citations and HTML
comments stripped before the scan, so the #2334/#2339 over-warning on
`None`, on the shipped `<!-- brackets optional -->` template comment, and on
annotated lines cannot return.
Regression tests extend the #2316/#2334 fixture family in tests/phase.test.cjs
(10 cases: 4 defect, 2 canonical controls, 4 negative-space controls).
Fixes #3697
* fix(#3697): rework under-selection detection onto tokens, not a free-text scan
Round 2, driven by the P4.6 cross-AI review (codex, gpt-5.6-sol) of b3ce71cb.
That review refuted 5 of 9 claims; three were false-positive classes in exactly
the category #2334/#2339 had to REMOVE:
`RANGE-01, RANGE-02 - 3 points` the bare-hyphen alternative read
`RANGE-02 - 3` as a range
`REQ-01, REQ-02 — locked per ADR-7.` the trailing period kept `ADR-7.` out
of the anchored filter, so the
unanchored substring scan reported it
as unparsed
`REQ-01, REQ-02 (see (ADR-7), then ADR-8)`
nested parens left `ADR-8)` behind
Replaces the free-text substring scan + loose range regex with three narrow,
token-based rules (R1 range-shaped token, R2 pure range operator flanked by two
selected IDs, R3 zero-selection with ID-shaped text). Also fixes the review's
CLAIM 9: the warning said IDs were "marked complete" when a ghost range marks
nothing — it now says "selected".
Side effect: the two false NEGATIVES the same review found are now covered —
`RANGE-01 through RANGE-05` and a parenthesised `(plus RANGE-02..RANGE-05)`.
NOT YET DONE (see the handoff prompt): regression tests for the four false
positives, the two new true positives, and the #3697-4 tightening the review's
CLAIM 8 asked for (it currently filters on the warning's phrasing rather than
asserting silence). Verified so far: tsc clean, the 10 existing #3697 tests
green, and a 20-case standalone harness covering every case above.
* test(#3697): pin the v2 token-detector boundary end-to-end
Six new cases + two hardenings for the review findings against v1:
- #3697-1 gains the worded spaced range (`RANGE-01 through RANGE-05`) —
the operator set's `to|thru|through` arm was previously untested.
- #3697-5 (new): a tight range hidden inside balanced parentheses
(`RANGE-01 (plus RANGE-02..RANGE-05)`) warns, names the range token,
and ticks exactly RANGE-01 — the paren shave must not hide it.
- #3697-4 gains the four false-positive classes a free-text detector
produced: numeric estimate (`- 3 points`), date annotation, em-dash
citation with trailing period (`— locked per ADR-7.`), and nested
parenthetical citations.
- #3697-3 and #3697-4 now assert the ENTIRE warnings channel is empty,
not that one phrase is absent — a re-worded over-warning cannot pass.
Negative control: against the merge-base with its lib rebuilt, all 6
defect tests fail and all 10 controls pass.
* fix(#3697): close round-2 review findings — annotation false positives
Round 2 of the adversarial review (against 822a72a04) refuted five
claims; this closes the false-positive class and the cheap misses:
- R1's bare-hyphen arm now demands a full ID on BOTH sides
(`REQ-01-REQ-05`): `LETTERS-\d+-\d+` is also a date-like annotation
(`FY-2026-08`) and a sub-numbered ID, and warning on those is the
expensive class. Tight hyphen shorthand with a live selection is the
disclosed false negative; at zero selection R3 still catches it.
- R2 requires the endpoint pair to imply an INTERIOR (same prefix,
gap > 1): `REQ-02 - REQ-03` selects both endpoints and can drop
nothing, so an annotation hyphen between adjacent IDs stays silent.
- R3 skips placeholder-led lines: `None (per ADR-7)` is a declared-empty
line citing its rationale, not unparsed residue.
- Token shave: quotes/backticks now shaved from alphanumeric tokens
(`` `RANGE-02..RANGE-05` `` warns); punctuation-only tokens get a
bracket-only shave so `(..)` surfaces its operator.
- 256-char token cap bounds the quadratic unanchored substring test.
- Warning text mentions range expansion only when a range rule fired.
Tests: 6 new cases (22 total). Negative control against the merge-base:
8 defect tests fail, 14 controls pass.
* fix(#3697): close round-3 review findings — half-spaced ranges, cross-prefix annotations, markdown wrappers
Round 3 of the adversarial review (against 2eb92dd0e) refuted four
claims; this closes them:
- Half-spaced ranges (`REQ-01 -REQ-05`, `REQ-01- REQ-05`) split at the
tokenizer before R1's `\s*` can see them and under-selected silently.
A glued-fragment rule warns when an operator is glued to a full ID
with an ID-shaped neighbour on the open side and the endpoint pair
implies an interior.
- Cross-prefix pairs around a separator no longer read as ranges:
`REQ-02 - (ADR-7)` and `REQ-02 (...) (ADR-7)` are annotations, and
real ranges are same-prefix by nature. `impliesInterior` now returns
false on prefix mismatch and computes the gap with BigInt (parseInt
lost precision past 2^53).
- The token shave now removes markdown emphasis markers and curly
quotes, so `**None** (per ADR-7)` reaches the placeholder gate and
`**RANGE-02..RANGE-05**` reaches R1.
- The unanchored-substring cap rises to 2048 (a markdown-link range
with a long URL cleared 256); the anchored range regexes scan
linearly and drop their cap.
Tests: 6 new cases (28 total; 377/377 file-wide). Negative control
against the merge-base: 11 defect tests fail, 17 controls pass.
* fix(#3697): round-4 review finding — word operators excluded from glued-fragment rule
`TOREQ-05` is a valid prefix-agnostic REQ-ID, and the glued-fragment
rule read it as `to` + `REQ-05`, warning on the canonical two-ID list
`REQ-01, TOREQ-05`. Glued fragments are now SYMBOL-operator-only
(`..`+, ellipsis, dashes): a word operator glued to an ID is an ID,
not a range spelling.
Tests: word-operator-prefixed ID control (misparse channel silent; the
fixture's ghost-ID warning legitimately fires, so the whole-channel
assertion stays with the registered controls) and an underscore-wrapped
tight-range defect case. 30 targeted cases; 379/379 file-wide; negative
control: 12 defect tests fail on the merge-base, 18 controls pass.
* fix(#3697): round-5 review findings — trailing word-op glue, dot shave, honest wording
- The glued-fragment TRAILING arm takes the word operators back: an ID
must end in digits, so `REQ-01through` can never be an ID — the
round-4 TOREQ collision was leading-arm-only, and symbol-only on both
arms lost the `REQ-01through REQ-05` typo class.
- A trailing run of 2+ dots survives the punctuation shave: `REQ-01..`
is a glued range operator, not sentence punctuation, and the shave
was silently eating the `REQ-01.. REQ-05` form.
- The warning now says the line "could not be parsed as" a
comma-separated REQ-ID list: `**REQ-01**, **REQ-05**` IS such a list
— the selector just cannot parse decorated tokens — and a warning
that misstates the input teaches readers to distrust it.
Tests: two new trailing-glue defect cases (32 targeted; 381/381
file-wide). Negative control: 14 defect tests fail on the merge-base,
18 controls pass.
* test(#3697): use t.after for cleanup per CONTRIBUTING test ruleset
CONTRIBUTING bans try/finally inside test bodies (it masks failures);
the approved shape is `t.after(() => cleanup(tmpDir))`. All seven
converted tests are this PR's own additions; the file's pre-existing
instances are untouched.
* chore(#3697): add changeset fragment for the Requirements-line under-selection warning
changeset-lint fails on this PR (fail_missing_fragment): src/phase.cts is a
user-facing surface and the branch carried no .changeset/*.md. Adds the Fixed
fragment via `npm run changeset -- --type Fixed --pr 3744`, symptom-led per
the house format, with the (#3697) backlink.
* refactor(#3697): extract the Requirements-line detector to a testable surface
Round-3 review Blocker 1 requires a fast-check property test over this
detector (`RULESET.TESTS.property-based-testing`: modules implementing
parsing contracts must include at least one), and Blocker 2 requires
limit-1/limit/limit+1 fixtures on its 2048-char token cap
(`RULESET.TESTS.boundary-coverage.fixtures`). Neither is expressible while
the logic is a closure inside `cmdPhaseComplete`: every existing #3697 test
reaches it by spawning the CLI, and a property test cannot pay a subprocess
per generated case.
So the selector and the three detection rules move to module scope as
`analyzeRequirementsLine` (pure, exported) plus
`formatRequirementsLineWarning`, and `cmdPhaseComplete` calls them. This
commit changes NO behaviour: `tests/phase.test.cjs` is untouched here, and
the pre-round suite passes against it unmodified (403/403).
Two things the move makes explicit rather than incidental. The selector and
the detector tokenize the SAME line DIFFERENTLY — the selector strips only
`[` and `]`, the detector also shaves quotes, emphasis and trailing sentence
punctuation — and that gap is deliberate: it is why `ADR-7)` is not selected
while `ADR-7` is still nameable in a warning. They now sit adjacent with the
reason written down, so they cannot drift apart silently.
And the stale citations in the moved comment are corrected. It pointed at
src/phase.cts:833,920,1078 for the `**Requirements**: TBD` seeds, which had
drifted to 1132/1237/1413, and at `templates/roadmap.md:32`, which is
`gsd-core/templates/roadmap.md:32`. Both are now anchored by content.
* fix(#3697): stop the warning claiming a misparse that did not happen
Round-3 review Major 3 and Minor 4. Both are the same defect: the warning
asserted more than the evidence supported.
MAJOR 3 — a correct comma list such as `RANGE-01, RANGE-02 — RANGE-05
deferred` warned "could not be parsed ... Range forms are not expanded;
rewrite the line". Reproduced: it selects RANGE-01, RANGE-02 AND RANGE-05,
i.e. every ID written on the line. Nothing was dropped, and the pinned
control only stayed silent because its pair was ADJACENT (gap == 1), so the
control was passing by accident of the fixture rather than by the rule.
The obvious fix — go silent — is not available. `RANGE-02 — RANGE-05` as a
range and as an annotation separator are textually identical, and no
token-level rule separates them; staying quiet re-opens the exact silent
under-selection #3697 is about. Deciding the ambiguity by assertion in
either direction is wrong. So it is DISCLOSED: the warning now has two
channels, chosen by whether any ID-shaped token was actually left unselected
(`droppedIdShaped`).
* something was dropped (tight range, glued fragment, inert residue)
-> "could not be parsed as a comma-separated REQ-ID list", as before.
* nothing was dropped (only the spaced-operator rule fired)
-> "contains what reads as a range between two cited REQ-IDs", stating
both readings and saying explicitly that an annotation separator means
the line is already correct.
This retires the "could not be parsed" wording for the four #3697-1 spaced
cases too, and that is a deliberate expectation change rather than a fix
counted twice: those lines never failed to parse either. They still warn,
still name the selected IDs, and still assert the endpoint-only marking is
unchanged; #3697-1 now also asserts the misparse channel stays SILENT.
MINOR 4 — `Deferred (see ADR-7)` reported `Unparsed text: ADR-7`, naming a
citation as requirement content it had failed to read. The trigger is
correct and stays: #3697's acceptance criterion asks for a warning "when it
selects zero IDs from a line that is non-empty and is not the `TBD`
placeholder", and inferring placeholder-ness from arbitrary prose is the
free-text heuristic this detector exists to avoid. What was wrong is the
wording, so the non-range arm now says "ID-shaped text that was not
selected" and names the escape the author actually has (`TBD` / `None`).
Tests: #3697-9 (three spaced forms — must warn, must NOT claim a misparse,
must offer both readings) and #3697-10 (`Deferred (see ADR-7)`, `N/A
(tracked in ADR-12)` — must warn, must not say "Unparsed text", must not
diagnose a range, must name the placeholder escape).
Reversion control: reverting the ambiguous channel fails #3697-9 (3 named
tests); reverting the R3 wording fails #3697-10 (2 named tests).
* fix(#3697): cap every token predicate, complete the dash set, cover the boundary
Round-3 review Blocker 2 and Nit 6, plus one self-found finding. All three
are about the detector's own predicates, so they land together.
BLOCKER 2 — the 2048-char budget had no boundary coverage.
`RULESET.TESTS.boundary-coverage.fixtures` requires limit-1 / limit /
limit+1 for any budget parameter. #3697-B1 and #3697-B2 now exercise 2047 /
2048 / 2049 against BOTH predicate families the cap guards, and each asserts
its fixture's exact length before asserting behaviour, so a mis-built
fixture fails loudly rather than passing at the wrong size. Clause (d) of
that rule — an input pushed within reserve-distance of the limit — has no
referent here: this is a hard cap with no reserve constant beside it, and
the test comment says so rather than leaving the omission to be re-derived.
NIT 6 — the cap guarded only the unanchored ID-substring regex. The
anchored range regexes were left uncapped, justified by a comment asserting
they scan linearly. The finding is right that this is informational (they
are anchored; the input is a local ROADMAP.md), but an asserted property is
cheaper to enforce than to defend, so all three predicates now share one
`short()` guard. #3697-B2 is what pins it: at 2049 the anchored scan must
now decline to classify.
SELF-FOUND (RV4 guard-shape census) — the range-operator set is a list this
code fixes at author time over a domain that grows without it, so the round
owes a census of what the enumeration reaches.
reached: `..`+, U+2026, U+2013, U+2014, ASCII `-`, to/thru/through
NOT reached: U+2010 hyphen, U+2011 non-breaking hyphen, U+2012 figure
dash, U+2015 horizontal bar, U+2212 minus sign
consequence: a range spelled with any of those is SILENTLY under-selected
— #3697's own defect, in the code that exists to fix it
Those five close. They are the same operator at a different codepoint and
carry none of the ASCII hyphen's collision risk, because they are not the
REQ-ID separator: `FY-2026-08` is date-shaped only with ASCII hyphens, so a
U+2010 never reaches the ID shape. They therefore join the NOHYPHEN arm
beside `—` and `–`; the strict full-ID-both-sides shape the bare hyphen is
held to is untouched, and #3697-12 pins that.
Still NOT reached, declined with reason rather than left unstated: `→`, `~`,
`..=`, `..<`, `until`, and `up to` (two tokens, so never one operator
token). Each is a symbol or word with an independent non-range use between
two REQ-IDs — the over-warning class #2334 cost three rounds.
Reversion control: reverting the uniform cap fails #3697-B2 (limit+1);
reverting the dash set fails #3697-11 (5 named tests).
* test(#3697): add the fast-check property coverage the parser rule requires
Round-3 review Blocker 1. `RULESET.TESTS.property-based-testing` (CONTEXT.md)
requires modules implementing parsing contracts to carry at least one
fast-check property test asserting a domain invariant, and the round-2 diff
had zero occurrences of `fc.` across its +311 test lines. Five properties,
1,900 generated cases:
P1 soundness of silence (boundary containment) — for ANY canonical comma
list of well-formed REQ-IDs, the selected set EQUALS the written set
and nothing warns. This is the #2334 over-warning invariant and the
#3697 under-warning invariant asserted as one statement, over
generated IDs rather than hand-picked ones. It generalises #3697-4b:
a prefix beginning with a word operator (`TORANGE-05`) is an ID, and
P1 covers that class rather than the single example.
P2 completeness — a same-prefix pair with an interior between them,
separated by any of the nine spaced operators, ALWAYS warns.
P3 the #2334 invariant — an ADJACENT pair around a separator can drop
nothing, so it stays silent however it is annotated.
P4 totality + idempotency — total over arbitrary strings, deterministic,
and the formatter agrees with the analysis on whether there is
anything to say (a warn with no text, or text with no warn, is a
channel that can go silent or noisy on its own).
P5 containment — every selected ID is ID-shaped and appears verbatim in
the input.
Honest scoping, since a property test is easy to overclaim: P1, P3, P4 and
P5 hold against the round-2 code as well as this one — they are regression
guards, not bug-finders, and their value is that the invariants are now
stated and generatively checked rather than implied by examples. P2 is the
one that would have failed before the dash enumeration was completed.
fast-check v4 removed `fc.stringOf`, so the ID-prefix tail is built from
`fc.array(...).map(join)` with the alphabet pinned to the selector's own
`[A-Z0-9]` class.
These live in tests/phase.test.cjs rather than a new
`phase.property.test.cjs`: `lint-test-file-count` caps a production module
at 2 test files and phase.cts is already at its allowlisted entry, so a new
file would trade one gate for another.
* docs(#3697): document the ROADMAP Requirements-line grammar
Round-3 review Minor 5 — the change adds net-new user-visible warning output
for a grammar constraint documented nowhere under docs/. `type: Fixed` is
docs-exempt so this does not block, but a warning about a rule the reader
cannot look up is not actionable, and that is worth fixing whether or not a
gate demands it.
Added as a subsection of `phase complete` in docs/CLI-TOOLS.md, beside the
existing SUMMARY artifact-check advisory it is a sibling of: the supported
comma-list form, why ranges are deliberately not expanded, that `TBD` and
`None` are the entire placeholder vocabulary, and what each of the two
warning voices means — including that the range/annotation one may be
reporting a line that is already correct.
Existing file rather than a new one, deliberately: docs/ carries generated
indexes and zh-CN / ja-JP trees, and a new top-level page invites a parity
or index gate this change has no reason to touch.
* fix(#3697): rule-scope the warning-channel discriminator
Self-found at the round's pre-push review, against the Major 3 fix two
commits back. That fix chose the channel from a LINE-GLOBAL question — "was
any ID-shaped token left unselected?" — while the rules that produce the
warning are not line-global. The two disagree as soon as the line carries an
ID-shaped token no rule fired on:
`RANGE-01, RANGE-02 — RANGE-05 deferred per (ADR-7)`
`(ADR-7)` survives the selector's bracket strip, so the global test called it
a drop and sent the line to the assertive channel — putting the false "could
not be parsed ... rewrite the line" claim back on a correct line. That is
review finding Major 3 returning through a side door, and it directly
contradicts #3697-4, which pins a parenthetical citation as NOT unparsed
residue.
The discriminator is now rule-scoped: R2 is the only ambiguous rule, so the
ambiguous channel requires that R2 fired, that no other rule did, and that
every endpoint R2 fired on was actually selected. The last conjunct is not
redundant — the detector shaves brackets and the selector does not, so R2 can
fire on a `(RANGE-02)` that was never selected, and that IS a drop:
`RANGE-01 (RANGE-02) — RANGE-05` -> assertive, correctly
`droppedIdShaped` is replaced by `spacedRangePairs` (R2's hits, so the
channel can ask about the endpoints the rule fired on) and the
`rangeReadingOnly` verdict.
Reversion control: against the line-global rule, #3697-9b fails. #3697-9c
passes under both rules — there the dropped token IS the R2 endpoint, so the
two agree; it is a regression guard, not a bug-finder, and is recorded as
such rather than counted as a second control.
* fix(#3697): hold every dash to the strict range shape, not just ASCII
Self-found at the round's pre-push review, and it CORRECTS a claim made two
commits back. That commit widened the range-operator set by five Unicode
dashes and asserted they "carry none of the ASCII hyphen's collision risk,
because they are not the REQ-ID separator". That reasoning was wrong. The
collision is a property of the SHAPE — `PREFIX-\d+ <dash> \d+` is also a date
(`FY-2026-08`) and a sub-numbered ID (`API-2-01`) — and the shape does not
care which dash sits in the operator slot, because the ID's own separator is
still ASCII either side of it. Measured:
RANGE-01 (target FY-2026-08) silent <- pinned by #3697-4
RANGE-01 (target FY-2026‐08) WARNED <- same line, U+2010
So the widening reintroduced the #2334 over-warning class on a date
annotation. It also exposed that the inconsistency PREDATES this PR: U+2013
and U+2014 were already in the loose arm at ce71dd399, so the en- and em-dash
forms of that same date annotation warned before round 3 ever ran.
One rule for every dash: a tight range spelled with any of the eight must
carry a FULL ID on both sides, exactly as the bare hyphen already had to.
`..`, `…` and the word operators stay loose — no date or sub-number reading
exists between two numbers, so the strict shape would cost them coverage for
nothing.
The cost is a false negative, and it is one the design already accepts:
`RANGE-01, RANGE-02-05` is silent today, deliberately, and now
`RANGE-01, RANGE-02–05` is too. That removes an inconsistency rather than
opening a gap, and a bare `RANGE-02–05` still warns — it selects nothing, so
R3 catches it.
Tests: #3697-13 (date annotation AND sub-numbered ID silent for all eight
dashes), #3697-13b (full-ID tight range still warns for all eight),
#3697-13c (loose operators keep their numeric endpoint), #3697-13d (the
accepted false negative is symmetric, and the bare zero-selection line still
warns).
* fix(#3697): close the round's own pre-push review findings
An adversarial cross-AI review of this round refuted 4 of its 10 claims. All
four were real. Every fix below is to code THIS round introduced.
1. THE SOFT VOICE CLAIMED TOO MUCH (refuted CLAIM 1).
`REQ-01, (REQ-02), REQ-03 — REQ-05` took the range-reading voice and told
the author "the line is already correct and nothing needs to change" — while
`(REQ-02)` had been dropped by the selector, which does not strip
parentheses.
The channel choice is still right, and deliberately so: `(ADR-7)` and
`(REQ-02)` are the SAME shape, so routing on "was anything unselected?" puts
the false "could not be parsed" claim back on a line carrying a citation —
the misroute fixed two commits ago. No rule can adjudicate this; the author
can. So the voice stops asserting the line is correct (it now speaks about
the SEPARATOR, which is all it has evidence about), and BOTH voices gained a
factual clause naming ID-shaped text the selector skipped, with the reason
(brackets are not stripped) and no verdict attached.
2. THE CAP SILENCED A LINE THAT USED TO WARN (refuted CLAIM 2).
A 2049-char range token warned before this round and went silent after it:
the "uniform cap" commit bounded the predicate and, with it, the warning.
That is #3697's own defect, introduced by the fix for a nit.
The cap bounds the WORK, not the warning. An over-cap token carrying `-` is
now recorded as unclassified (a linear `includes`, never the unanchored
regex the cap exists to keep off it) and gets its own voice: "could not be
checked ... the REQ-ID selection on this line is unverified". Unclassified
is reported, never treated as clean.
3. THE CAP WAS NOT UNIFORM (review MISSED finding).
R2 capped the operator token but not its neighbours, so
`<2049-char ID> .. <2049-char ID>` still ran REQ_ID_SHAPE_RE and BigInt over
both endpoints unbounded. The glued rule had the same hole. Every
participant is capped now.
4. PROPERTY P5 WAS VACUOUS (refuted CLAIM 5).
It drew from a bare `fc.string()`, which over 500 samples produced max
length 10 and ZERO inputs containing a REQ-ID — the loop body never executed
an assertion. A containment property that never contains anything is a green
test measuring nothing. The generator now interleaves real IDs with noise
and the property ASSERTS it saw them (>50/500), so it can never silently go
vacuous again. The free-form coverage it was actually providing survives,
honestly labelled, as #3697-P6.
The same finding refuted this round's claim that P2 distinguishes pre-round
behaviour: every operator P2 uses was already in the pre-round operator set.
P2 is a regression guard, and its comment now says so.
Also: docs/CLI-TOOLS.md repeated the broken channel claim verbatim (review
MISSED finding) and is corrected with the code.
Tests: #3697-9d (soft voice names the skipped ID, never claims the line is
correct), #3697-9e (over-cap token reported as unclassified, still warns),
#3697-9f (R2 and the glued rule cap their neighbours). #3697-B1/B2 now key the
boundary on the PREDICATE's verdict with `warn` asserted true at every length —
asserting `warn === false` at limit+1 was itself finding 2.
* docs(#3697): describe the third voice and the dash rule
Follow-on to the review-findings commit: that commit corrected the docs' claim
about the soft voice but left two things the code now does undescribed.
- There are THREE voices, not two. The over-cap voice ("could not be checked
... unverified") arrived with the fix for the review's CLAIM 2 and had no
entry.
- Dash spellings require a full ID on both sides, and `..` / `…` / the word
operators do not. That asymmetry is deliberate and load-bearing —
`PREFIX-<digits><dash><digits>` is date- and sub-number-shaped — so a reader
hitting `REQ-01-05` and getting silence has no way to find out why. The
accepted cost (`REQ-01, REQ-02-05` unreported, bare `REQ-02-05` still
reported) is stated rather than left to be discovered.
Documentation only; no behaviour change.
* fix(#3697): close the continuation review's findings
A continuation of the same adversarial reviewer, run against the reworked
round, refuted 6 of 7 claims. Four were real defects in this round's own work
and are fixed here; the other two are answered rather than changed, below.
1. THE SKIPPED-TEXT CLAUSE WAS ON ONE VOICE, NOT BOTH (refuted CLAIM A).
The previous commit's message said both voices gained it. Only the soft
return appended it. The assertive voice now carries it too — and, because
that voice already names range tokens and inert residue under its own
clauses, the note is filtered to what those did not already name. A warning
that says the same token twice is one readers learn to skim.
2. THE CLAUSE'S WORDING WAS FALSE (also CLAIM A).
It read "brackets and parentheses are not stripped". Square brackets ARE
stripped by the selector — `[REQ-01, REQ-02]` is the documented form — so
only parentheses qualify. Corrected in the message and in docs/CLI-TOOLS.md,
which had inherited the same error.
3. THE OVER-CAP RULE STILL SILENCED A LINE (refuted CLAIM B).
`oversizedTokens` filtered on `includes('-')`, which misses an over-cap
OPERATOR: `REQ-01 <2049 dots> REQ-05` warned before this round, R2 declined
to classify it once capped, and nothing reported it. That is the exact
regression the field was added to close, one input over. Any token past the
cap now counts — what it contains is irrelevant when we could not read it.
4. AND THEN OVER-REPORTED ONE (review MISSED finding).
With (3) in place, a 2049-character CANONICAL REQ-ID was selected by the
uncapped, fully-anchored selector AND flagged "REQ-ID selection on this line
is unverified" — a contradiction inside one warning. A token the selector
took was examined end to end, so it is excluded.
Two findings are answered, not changed:
CLAIM C — the selector's own `REQ_ID_SHAPE_RE.test` is uncapped. True, and
deliberate: this round does not touch what gets MARKED, and the pattern is
anchored at both ends with no nested quantifier, so it is linear. The claim
that "all predicate paths are capped" was too broad; the DETECTOR's are.
CLAIM E — `REQ-01, REQ-02<dash>05` is silent for every dash. That is the
documented, deliberate cost of holding dashes to the strict shape, and it is
symmetric with ASCII, which behaved that way before this PR. The reviewer is
right that "without losing a range spelling that should be detected" was too
strong; a bare `REQ-02<dash>05` still warns.
Tests: #3697-9g (clause on the assertive voice, no repetition, bracket claim
true), #3697-9h (over-cap operator does not silence the line), #3697-9i (a
selected over-cap ID is never called unverified). #3697-9f is rebuilt — its
first version used the SAME id twice, so R2 could not have fired even uncapped
and it proved nothing; it now uses endpoints with a gap and fails when the
neighbour cap is removed.
Reversion control: all four fixes fail a named test when reverted in isolation
(#3697-9h, #3697-9i, #3697-9g, #3697-9f).
* fix(#3697): scope the over-cap exemption to what could actually pair
A second continuation of the same reviewer, against the reworked round,
confirmed the two claims that matter most and refuted three. This closes the
one real defect; the other two are answered below.
CLAIM J / CLAIM K (one defect, found from both directions). The previous
commit exempted EVERY selector-accepted token from `oversizedTokens`, on the
reasoning that the selector is uncapped and anchored so it examined the whole
token. True of that token's SELECTION — and not the same as "no rule was
suppressed by it". Two over-cap valid IDs either side of `..` are both
selected, so both were exempted, and R2 is capped: a line that warned before
this round went silent.
That is the third appearance of one class in this round — the cap suppresses a
check, and the suppression is not reported. Each fix for it over-corrected in
the opposite direction, which is why the rule is now stated in terms of what
was actually suppressed rather than in terms of the token: an over-cap token is
exempt only when it was selected AND nothing beside it could have paired with
it into a range (no range operator, no glued fragment, no second over-cap
token). Everything else is unexaminable and says so.
Two findings are answered, not changed:
CLAIM M — the reviewer demonstrated, with driven evidence, a contextual rule
that catches `REQ-01, REQ-02-05` while leaving `FY-2026-08` and `API-2-01`
silent: recognise `PREFIX-a<dash>b` only when another SELECTED id on the line
shares that prefix. That refutes this round's claim that the strict-dash
trade was FORCED, and the claim is withdrawn — it is a design choice. The
choice stands for this PR: the conservative rule is what ASCII already did
before #3697, adopting a new contextual heuristic unreviewed at the end of a
round is how the last three defects in this round were made, and #3697 asks
for a warning rather than better range inference. Named here so the
alternative is on the record rather than lost.
Docs MISSED — CLI-TOOLS said every token over 2,048 characters "is not
classified at all" and warns. Selection is not bounded; only range detection
is. Corrected.
Confirmed by the same pass, and worth recording because they are the PR's
load-bearing promises: a 20,000-input comparison of the pre-extraction selector
against HEAD found `mismatches=0` (nothing about which REQ-IDs are MARKED has
changed), and the uncapped selector regex was measured linear from 100k to 800k
characters.
Tests: #3697-9f now asserts the range case is reported rather than silent, and
#3697-9j pins the exemption's scope in both directions. Reversion control:
restoring the blanket exemption fails both.
* fix(#3697): warn on zero selection, as the acceptance criterion asks
`Deferred`, `N/A`, `Pending`, `TBA` and `-` selected no REQ-IDs and stayed
SILENT, while three shipped artifacts said they warned: `docs/CLI-TOOLS.md`,
the `placeholderLed` census comment, and the advice string the command emits
to the user. The asymmetry was the tell — `Deferred (see ADR-7)` warned,
because the citation supplied the ID-shaped residue R3 required, while bare
`Deferred` did not. The claim was written into three places and never
executed once.
This is also #3697's AC-1b/AC-4 verbatim: "warn when `citedReqIds.length ===
0` while the raw capture is non-empty and not `TBD`".
R3b keys on the SELECTION being empty, never on what the prose means, so it
adds no free-text heuristic. It is deliberately not gated on ID-shaped
residue the way R3 is, and the negative space is what settles that: all
fifteen #2334/#2339 fixtures are held silent by non-zero selection or by
`placeholderLed`, and not one of them by the ID-shape gate — measured, not
argued. The gate was buying no negative space while costing the acceptance
criterion.
`tokens.length > 0` keeps an empty line and a comment-only line silent: the
tokenizer strips `<!-- ... -->` before splitting, so the shipped template's
own comment cannot reach the rule.
Selection behavior is unchanged. This warns; it never invents an ID.
Also extracts `warn` to a named const (round 3 review Minor 3) — this commit
adds a disjunct to exactly that predicate, and in the return literal a later
reordering would be a TDZ ReferenceError rather than a reader-visible error.
Tests: #3697-14 (six zero-selection lines warn and tick nothing, and the
warning names the TBD/None escape), #3697-14b (five placeholder spellings
stay whole-channel silent), #3697-14c (comment-only line stays silent).
Fail-first controls: all six #3697-14 cases fail against the pre-fix tree;
-14b and -14c pass at both ends, which is correct — they pin silence the
widening must preserve.
* fix(#3697): name the REQ-ID a glued delimiter dropped
`RANGE-01; RANGE-02` selects only RANGE-02 and marks only RANGE-02, with
`requirements_updated: true` — #3697's own half-success failure mode, reached
by one wrong delimiter, and silent before this rule. It is the issue's AC-1a
("a warning whenever the line contains ID-shaped content that the tokenizer
did NOT select") at the shape most likely to be typed by accident.
Round 4 review rated this Major rather than Blocker on the ground that the
case is indistinguishable from a parenthesised citation, since `(ADR-7)` also
shaves down to a bare ID. At the RAW token level it is distinguishable, and
that is what makes the rule shippable: `REQ-01;` is shaved of a trailing
DELIMITER, `ADR-7)` of a citation wrapper. R4 keys on that shave class and
requires the token to sit outside any parenthetical.
Measured before implementing: 0 false positives and 0 false negatives across
21 probes, including all fifteen #2334/#2339 negative-space fixtures. A first
cut without the parenthetical test scored 3 false positives — every one of
them a colon inside a citation (`(see ADR-7: section 3)`) — which is why that
test is the rule's boundary rather than an optimisation.
Adds the delimiter census the module did not have. The range-operator domain
was already censused; the comma-substitute domain was not. Swept 26
spellings: exactly two produce a silent under-selection, `; ` and `: `. Every
other spelling either selects both IDs or selects none and already warns. The
review hand-listed the semicolon; the colon is the sibling that sweep found,
and it fails identically.
`rangeReadingOnly` now excludes an R4 hit — the ambiguous voice claims nothing
was dropped, and must not speak for a line where something demonstrably was.
Tests: #3697-15 (four delimiter shapes warn, name EVERY dropped ID, and tick
exactly the unchanged selection), #3697-15b (three citation forms stay
whole-channel silent). Fail-first control: all four #3697-15 cases fail
against the previous commit's tree; -15b passes at both ends, pinning the
boundary the widening must not cross.
* fix(#3697): give the Requirements-line warning a stable machine kind
The warning's kind existed only in the prose of its message, so every consumer
and every test had to regex an English sentence — and rewording a message
silently un-asserted the tests that pinned it. Round 4 review Major 3.
The repo already had the settled seam for exactly these semantics.
`CONTEXT.md` records `diffLiveConfig` emitting `kind:'unverified'` for a
truncated scan, which is precisely this module's third voice; and
`WAVE_CLEANUP_WARNING` in `src/worktree-safety.cts` carries codes for the same
reason. ADR-3473 Decision 3 ("failure is a value") points the same way.
`formatRequirementsLineWarning` now returns `{ code, message }` instead of a
bare string, which also settles round 4 Nit 3 — `null` still means CLEAN, a
legitimate value, but the success arm is no longer a naked string one field
away from the shape the ADR standardises on.
The kind is carried ALONGSIDE the prose, never instead of it. `warnings[]` is
a documented `string[]` in `phase complete`'s JSON output, rendered by
execute-phase.md's "If has_warnings is true" step, so re-typing its elements
would be a breaking output-contract change for a shipped command. The code is
emitted as its own additive `requirements_line_warning` field, absent
entirely when the line is clean.
Vocabulary, exported so tests key on it rather than on string literals:
`req-line-misparse`, `req-line-range-reading`, `req-line-unverified`.
Tests: channel ROUTING in #3697-9/-9b/-9c/-9d/-9e/-9g/-10 now asserts the code;
message-content assertions stay where the user-visible wording is itself under
test. #3697-16 pins the code end-to-end through the CLI's JSON for four line
shapes and asserts warnings[] is still a string[]; #3697-16b pins that a clean
line emits no kind at all, because a field present on every run carries no
information. #3697-P4 holds kind-and-message-appear-together and
kind-is-in-the-declared-vocabulary over arbitrary input, so a channel added
later cannot ship without one.
* test(#3697): pin the divergence against the second parser of the same line
CLAUDE.md, KNOWN DEFECTS & ANTI-PATTERNS: "Generative Fix Divergence: when
sharing constants/arrays/parsers between parallel surfaces, add a parity
assertion test that fails if they diverge." Round 4 review Major 2.
`normalizePhaseReqIds` (src/gap-checker.cts) parses the SAME ROADMAP
`**Requirements:**` value — its own docblock says callers "may pass the
roadmap value through verbatim" — and diverges on four axes. Measured, not
inferred:
line phase complete gap-checker
RANGE-01..RANGE-05 [] 5 IDs
None (per ADR-7) [] ["ADR-7"]
(REQ-02) [] ["REQ-02"]
REQ-01a [] ["REQ-01a"]
REQ-01, REQ-02 both both
This pins the divergence rather than removing it, which is the review's
second option and the correct one here: unifying the two would change what
`phase complete` MARKS, and "the ledger-writing set is byte-identical to base"
is the one invariant this PR holds fixed. Every axis is now asserted in BOTH
directions, so drift on either side fails here instead of widening silently.
The range axis is a DELIBERATE disagreement and is labelled as such — #3697
declines range expansion in terms ("I am not asking for range syntax to be
supported") while gap analysis adopted it under #1269.
The placeholder axis is the one worth reading twice: `None (per ADR-7)` is a
declared-empty line to `phase complete`, which reads the lead token, and a
one-requirement line to gap-checker, which strips parentheses first so the
citation survives its ID-shape filter. That is a citation being reported as a
requirement.
#3697-17b states the cost concretely: one line, five requirements in scope to
gap analysis and zero to phase complete. This PR is what makes that
contradiction visible, by finally giving the silent side a voice.
* fix(#3697): stop the skipped-text rider reporting a date, and close the 4b channel gap
Two round 4 review minors, both about a warning saying something it cannot
support.
MINOR 2 — false rider content. `REQ_ID_SUBSTRING_RE` is unanchored, so
`FY-2026-08` matches as `FY-2026` and lands in `unselectedIdShaped`.
`REQ_RANGE_TOKEN_RE`'s entire strict-dash arm exists to keep that shape
silent, and #3697-4 pins `RANGE-01 (target FY-2026-08)` as producing no
warning at all — but whenever some OTHER rule fired on a line that also
carried a date annotation, the rider told the author to "check whether any of
it is a requirement" about a date. Not a false warning, since the line was
warning anyway; false CONTENT, in the #2334 voice, through the side door.
Filtered at the MESSAGE rather than in the analysis: `unselectedIdShaped`
stays a faithful record of what the selector skipped — it is documented as a
fact that never routes — while the user-facing clause declines to assert
requirement-ness about a shape the design already ruled unadjudicable.
#3697-18b is the other half, so the filter cannot become a silencer: a
genuinely dropped REQ-ID is still named.
MINOR 1 — `#3697-4b` asserted only that the ASSERTIVE channel stayed silent,
so a regression routing `RANGE-01, TORANGE-05` into the AMBIGUOUS channel
would have passed. Whole-channel silence is not available on that fixture (the
pre-existing ghost-ID warning legitimately fires on the unregistered
`TORANGE-05`), so the precise assertion is that no Requirements-line warning
of ANY kind was emitted. The machine code added earlier in this round is what
makes that statable; before it, "both channels" could only have meant a second
prose regex.
* docs(#3697): record the Requirements-line seam in CONTEXT.md
CLAUDE.md names the CONTEXT.md glossary as a PR gate, and
`get_cochange_context(src/phase.cts, 45d)` ranks CONTEXT.md 4th at 25
co-changes — above src/init.cts and src/roadmap.cts. This PR introduced a
named seam, three warning kinds, a bound, a rule taxonomy and a deliberate
cross-parser divergence, and recorded none of it. Round 4 review Major 4.
The precedent is explicit rather than inferred: the directly analogous seam
is already there as `LIVE-CONFIG.GUARD.SEAM.truncation`, including its bound
and its boundary obligation — and that entry is the one this module's third
voice was modelled on.
Eight predicates, in the machine-oriented section beside it:
.module the two exported functions and the code vocabulary
.selector-identity citedReqIds is byte-identical to base and is the
only thing reaching the ledger — a change to what
phase.complete MARKS is outside this contract
.rules R1 / R2 / R2' / R3 / R3b / R4 / over-cap
.kinds the three codes, and why they ride beside
warnings[] rather than inside it
.cap 2048, neighbours included, and the boundary rule
.placeholder the gate that actually holds the negative space
.census-domains both open domains with their NOT-reached members
.gap-checker-divergence the four axes, pinned not unified
The changeset type is `Fixed`, which exempts this PR from the docs/
co-change requirement — but the glossary gate is separate from that
exemption, and the 2048 cap in particular is a machine-canon-shaped fact
that until now existed only inside a source comment.
`docs/CONTEXT-INDEX.json` regenerated (269 predicates); lint:generated-sync
confirms all six targets in sync.
* docs(#3697): document what the command now does, in one changeset sentence
DOCS. The grammar section predated this round's two new rules, so it
under-described the behaviour it exists to make lookup-able:
- The placeholder paragraph enumerated three words; the rule is a DEFAULT.
Any wording that selects no REQ-IDs warns, and the placeholders are matched
as the LEAD token, so `None (per ADR-7)` and `**None**` are declared-empty
too. The comment-only line is called out, because "any other wording" would
otherwise read as covering the shipped template's own `<!-- ... -->`.
- The comma rule was implicit. `REQ-01; REQ-02` marks only REQ-02, and it is
the quietest way to lose a requirement on this line — `requirements_updated`
reads `true` either way — so it gets its own paragraph, with the
parenthetical exemption stated beside it.
- The machine kind is documented where a consumer would look for it, with the
instruction to key on the kind rather than the wording.
- The skipped-text note no longer implies it reports date shapes; it
deliberately does not, and silently omitting that left the doc promising the
behaviour this round removed.
CHANGESET (round 4 review Minor 4). CONTRIBUTING.md's format is
`**<Bold user-visible change>** — <symptom-led explanation>.` and both
canonical examples are one sentence; this fragment ran three. Now one, and
covering what the round actually delivers rather than only the range shape it
started from.
* test(#3697): keep phase.test.cjs off the docs-guard exemption fingerprint
A comment added earlier in this round named `docs/CLI-TOOLS.md` by path. The
docs-guard exemption ratchet (#3753 FIX 3) fingerprints literal `docs/`
references in exempt test files and fails when a new one appears, so that
comment turned four green gates red — `ci-docs-guard-registry` and the
registration lint — for a file that reads no documentation at all.
Caught by diffing the full suite's failing-name set against the same suite run
at `upstream/next` in a probe worktree: 33 of 37 failures reproduce at base
(install / config-home / shadowing tests under the sandbox HOME), and exactly
these 4 did not.
Rephrased rather than baselined. Adding the path to
DOCS_GUARD_EXEMPT_DOCS_PATHS is the sanctioned response when a test genuinely
starts READING a new docs path — the violation text asks the author to
re-confirm the exemption still holds. Nothing here reads documentation; the
guard matched prose. Baselining would have recorded a coupling that does not
exist and made the next reader wonder what phase.test.cjs does with
CLI-TOOLS.md. The comment still names where the contract is written, just
without planting a path string.
* fix(#3697): close four defects this round's own pre-push review drove
An adversarial cross-AI review of this round, run before the push, returned 6
CONFIRMED and 4 REFUTED. Every refutation was driven against the built tree,
and every one was a shape the author had not probed — the rules were correct
across the probe set and wrong just outside it.
(1) R4 FALSE POSITIVE, and it is the #2334 over-warning class arriving through
the rule added to close a different hole. `REQ-01, see ADR-7: section 3` fired:
`ADR-7:` is the same shave class as `REQ-01;`, and the parenthetical test does
not reach a BARE citation. The FP probe that scored this rule 0/0 only ever
tested the parenthesised form.
Fixed by requiring the dropped id's prefix to agree with a SELECTED id — the
module's own idiom, not a new heuristic: `reqEndpointsImplyInterior` already
demands an agreeing prefix for the same reason. Cost, stated in the census: a
dropped id whose prefix is on no selected id (`REQ-01, FOO-02: x`) stays
silent. Same trade the strict-dash rule takes — under-report a rare shape
rather than over-report a common one. Pinned as a declared blind spot by
(2) R4 FALSE NEGATIVE, on the DOCUMENTED form. `[REQ-01; REQ-02]` dropped
REQ-01 silently: the selector strips square brackets and R4's raw scanner did
not. The bracket spelling the shipped template recommends was the one shape the
rule could not see.
(3) The rider filter suppressed a REGISTERED requirement. `API-2-01` is a legal
requirement id — gap-checker's `parseRequirements` accepts it from
REQUIREMENTS.md — so a `\d+-\d+` filter hid a genuinely dropped requirement
behind a rule meant only to hide dates. Narrowed to a four-digit year segment.
The earlier #3697-18 case asserting `API-2-01` should be suppressed is REMOVED,
and the removal is recorded in place: its premise was refuted, it was not
inconvenient.
(4) An INVISIBLE line warned. A lone U+200B carried a token to the parser while
reading as empty to the author, so R3b fired with nothing on screen to explain
it. Zero-width and format characters are now stripped — stripped rather than
treated as delimiters, because splitting on one would fabricate two fragments
out of one ID.
Also corrects the documentation the same review found overstated: the line is
split on commas AND whitespace, and the ID shape is matched case-insensitively,
so `REQ-01 REQ-02` and `req-01, req-02` both select and neither warns. That was
pre-existing selector behaviour; this round is the one that asserted the docs
were true of it.
Tests: #3697-19 (four invisible-only shapes), -19b (embedded zero-width is
stripped, not split on), -19c (three citation forms), -19d (both bracket
spellings), -19e (both halves of the rider boundary), -19f (the declared blind
spot), -19g (the two documented tolerances). 500 tests in phase.test.cjs, 0
failures; lint:ci clean.
* fix(#3697): generalise the drop rule, and stop the invisible fix hiding a drop
The pre-push review's continuation refuted six of seven follow-up claims. The
first one is the one that mattered: the invisible-character fix committed in
c7dce173a INTRODUCED #3697's own defect. Stripping zero-width characters from
the detector wholesale made `REQ-01<ZWSP>, REQ-02` go SILENT — the selector
really does drop REQ-01, and the strip removed the only evidence of it. The
test written alongside asserted the tokens and the empty R4 result and never
asserted `warn`, so it DOCUMENTED the bug rather than catching it; that
omission was the reviewer's own MISSED finding.
An invisible is two different questions about one character, and the fix is to
stop conflating them: absence-of-content for the empty test, DECORATION on a
token for the drop rule. Neither is a reason to delete it from the line.
R4 is generalised accordingly, because the continuation drove four more shapes
a trailing-delimiter-only regex could not see — `REQ-01 ;REQ-02`,
`REQ-01 :REQ-02`, `**REQ-01;** REQ-02`, the backticked form — plus
`**REQ-01**, REQ-02`, where emphasis alone defeats the selector. These are one
class: decoration on a token the selector then cannot take. One rule, not four
patches; patching them individually is how a list stays short and wrong.
PARENTHESES ARE NOT DECORATION, and the suite caught me learning that: shaving
them made `REQ-01, (REQ-02), REQ-03 — REQ-05` report a glued delimiter that was
never there and broke #3697-9d's channel routing with it. A parenthesis is this
rule's citation marker.
The rider stops adjudicating an undecidable shape. `API-2-01` is a legal
requirement id and `API-2026-08` is too, while `FY-26-08` and `FY-2026-08-15`
are dates — no regex separates them, and both filters this round tried scored a
miss in each direction. It now NAMES the token and states the ambiguity, which
is the same thing the two warning voices already do about a range separator.
Filtering hides a real dropped requirement; reporting it bare asks the author
whether a date is a requirement; saying "this may equally be a date" does
neither.
The census and the docs are corrected to what the code does, including the part
that is NOT complete: the prefix gate does not stop a citation that SHARES a
selected prefix (`ADR-01, see ADR-7: sec 3` fires), and nothing at token level
separates that from a real drop. A prose heuristic on "see" is the free-text
detector this module exists to avoid, so the honest move is to say so.
CLAIM 17 — the invariant that actually matters — came back CONFIRMED on a
20,000-run fast-check property over arbitrary Unicode: `citedReqIds` is
identical to upstream/next's for every input, and marking is untouched.
513 tests in phase.test.cjs, 0 failures; lint:ci clean.
* fix(#3697): gate the drop rule on evidence, and stop an unmatched paren swallowing the line
Third pass of the round's own pre-push review, scoped to regression-hunting
rather than further polish. Three findings, all driven, all mine.
R4 OVER-WARNED on markdown styling. `REQ-01, see **REQ-7** for context` claimed
a dropped requirement: the previous cut treated any shaved decoration as
evidence, and emphasis is not evidence. Nothing separates that line from
`**REQ-01**, REQ-02` meaning to list one, so the rule now requires a positive
signal — a glued `;`/`:` (a list separator was INTENDED) or an invisible (the
token is CORRUPTED; nobody types one on purpose). Emphasis alone falls back to
the skipped-text rider, which names the id without asserting a drop, exactly as
`(REQ-02)` is handled. That is the #2334 class caught one cut before shipping.
R4 UNDER-WARNED on `**REQ-01**; REQ-02` — one shave pass cannot reach a wrapper
sitting behind a delimiter. Shaves to a stable point now.
The range OPERATOR lost its invisibles handling. `REQ-01 <ZWSP>..<ZWSP> REQ-05`
went silent, because the previous commit removed the invisible strip from BOTH
the tokenizer and R4 when only R4's was wrong. An invisible is two questions
about one character: for the classification rules it is noise and is stripped
from the token; for the drop rule it is the evidence and must survive on the
raw line. Stripping in both places hid a dropped id; stripping in neither hid a
range. The reviewer's MISSED finding named the missing control — regression
tests covered invisibles inside ids and not beside operators — and #3697-19i is
that control.
UNBALANCED PARENTHESES swallowed the line. `REQ-01, (note REQ-02; REQ-03`
reported nothing: a running-depth counter left the unclosed `(` open through
end-of-line, so every genuine drop after it inherited citation immunity. A
parenthesis confers that immunity only as part of a MATCHED span now — an
unmatched one is a typo, not a citation.
CLAIM 22 re-confirmed on a fresh 20,000-run property over arbitrary Unicode:
`citedReqIds` identical to upstream/next, marking untouched, warnings appended.
522 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite carries zero
head-only failures against a probe worktree at upstream/next.
* fix(#3697): make delimiter ADJACENCY the rule, and delete matched citations outright
Fourth and final pass of the round's own pre-push review. Three findings, and
they shared one root cause, so this is a narrower rule rather than a longer list
of shapes.
TOKEN-WIDE PAREN IMMUNITY LEAKED. `REQ-01, REQ-02;(note) REQ-03` is a single
whitespace token, so a matched parenthetical inside it conferred immunity on the
`REQ-02;` sitting OUTSIDE the parens, and the drop went silent. Matched spans
are now deleted from the line outright — which states what is actually meant,
that for this rule a citation is not on the line — and an UNMATCHED paren is a
typo that confers nothing. That also retires the running-depth counter whose
previous bug was the mirror image: an unclosed `(` swallowing the rest of the
line.
DECORATION WAS TESTED TOKEN-WIDE, so `REQ-01, see **REQ-7**; next topic` was
reported as a dropped requirement. It is a citation with sentence punctuation.
The rule is now ADJACENCY: styling is stripped, then the `;`/`:` must be
touching the id. `REQ-01;`, `;REQ-02` and `**REQ-01;**` qualify;
`**REQ-01**;` does not, because outside the styling that character is
punctuation. An invisible needs no adjacency test — nobody types one on
purpose, so anywhere in the token it is corruption rather than intent.
`**REQ-01**; REQ-02` therefore goes silent, and the test row asserting
otherwise is inverted rather than deleted quietly: it was added one commit ago
on the reasoning this pass refuted, and nothing distinguishes it from
`see **REQ-7**; next topic`.
Worth recording plainly: three successive cuts of this rule fired on a
citation, and each fix was a narrower definition of EVIDENCE, never a longer
list of shapes. The list-lengthening instinct is what produced the bug each
time.
The review's last MISSED finding named the missing control — the paren tests
all surrounded matched spans with whitespace, so none covered a span sharing a
token with an id outside it. #3697-19j carries both directions now.
526 tests in phase.test.cjs, 0 failures; lint:ci rc=0; full suite zero
head-only failures against a probe at upstream/next; the invariant that
`citedReqIds` is identical to upstream re-confirmed on 20,000 arbitrary
Unicode inputs.
* fix(#3697): state R4's real boundary, and stop the over-cap voice masking a drop
Two defects, both found by this round's own pre-publication body claim-audit.
1. A DEMONSTRATED drop was discarded by the unverified voice. On
`REQ-01, REQ-02: <2049 chars>` the analyzer names REQ-02 in
delimiterDroppedIds and the formatter then reported `req-line-unverified`,
whose message never mentions it — the one actionable finding masked by the
token beside it. The over-cap channel now excludes a line carrying an R4
hit, exactly as rangeReadingOnly already did and for the same reason: that
voice's whole claim is that nothing could be checked, and R4 has already
checked something. The assertive channel still carries the over-cap rider,
so nothing about the cap is traded away. Pinned by #3697-19l, which fails
against the pre-fix build and nothing else does.
2. Three shipped artifacts asserted behaviour the code does not have — the
same class as this PR's round-4 blocker, re-committed. CONTEXT.md's rules
predicate, the CLI tools reference, and the warning's own advice string all
listed markdown emphasis as an R4 trigger. It is not: styling is shaved
BEFORE the test and tolerated around an id, never a trigger on its own, so
`**REQ-01**, REQ-02` and `**REQ-01**; REQ-02` are both silent. The trigger
is exactly a glued `;`/`:` or an embedded invisible.
The census predicate was wrong in a second way. Its 26-spelling separator
sweep found only `;` and `:` because the sweep was SYMMETRIC-ONLY and
therefore biased: one-sided attachment drops silently for every punctuation
outside the set — `/ | & + . > \` and the full-width and non-ASCII forms
`; , ؛` all measured silent. The domain is wide open and R4 covers two
characters of it. Said plainly in all three places rather than widened
here: every previous widening of this rule first fired on a citation, so it
is not done blind at the end of a round.
Both blind spots are now PINNED as tests (#3697-19m styling-only, #3697-19n
one-sided separators) so the documents and the code cannot drift apart again —
which is what the round-4 blocker asked for.
tests/phase.test.cjs: 542 tests, 542 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* fix(#3697): re-sweep the separator census properly, and say what it really found
The round-4 census in src/phase.cts concluded "exactly two — `; ` and `: `"
from a 26-spelling sweep. That conclusion was forced by how the sweep was
built, not by the code: it swept the ONE-SIDED form (`REQ-01; REQ-02`) for the
semicolon and colon, and only the BARE and SYMMETRIC forms (`|`, ` | `) for
every other separator. Different members of the domain were tested in
different shapes, so no other answer was reachable. Caught by this round's
pre-publication claim-audit of the response comment, reading the census
comment against its own swept list.
Re-swept fully crossed and driven through the built artifact: 21 separators x
{bare, trailing-space, leading-space, both-spaces} = 84 combinations. 26 select
both ids, 24 under-select and already warn, and 34 UNDER-SELECT SILENTLY. All
34 are one shape — a separator glued to exactly one of the two ids, e.g.
`REQ-01/ REQ-02` or `REQ-01 /REQ-02` — for every punctuation except `,` and
the `;`/`:` that R4 covers.
So R4 covers TWO CHARACTERS of a wide-open domain. That is now what the census
comment, the CONTEXT.md census-domains predicate and the CLI tools reference
all say. The set is deliberately not widened here: three successive cuts of
this rule fired on a citation, and a fourth at the end of a round with no
adversarial pass is how each of those got in.
Second false passage in the same block: styling-only decoration was described
as "left to the skipped-text rider, which names the id". A rider only exists
inside a message, and a message only exists once some rule sets `warn` — so on
a line where nothing else fires, `REQ-01, **REQ-02**` is wholly silent.
Describing it as handled reads as coverage. #3697-19m already pins the silence.
so the test matches the documented claim.
tests/phase.test.cjs: 552 tests, 552 pass, 0 fail, 0 skip. lint:ci rc=0.
Both CONTEXT-INDEX consumers regenerated.
* chore(#3697): regenerate both CONTEXT-INDEX.json after rebasing onto next
Rebased onto next @
|
||
|
|
1a358ce0fd |
feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
dd4f179672 |
feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3a6c0412a9 |
enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) Extends ADR-1703's portability rule catalog with a second production-runtime rule: it flags an exact-case read of a Windows case-varying environment variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any receiver that is not process.env itself, matched via an env-shaped-receiver check to avoid colliding with ordinary `.path`-named properties elsewhere in the tree. Exports the seam's private `_envGet` as `envGet` so the rule's remediation message names a real helper, and fixes the one pre-existing violation the tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read). Closes #3624 * fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case Review findings from the code-review + isolated-adversarial passes: - extractStaticName only matched non-computed Identifier keys, so a destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8 acceptance case) silently evaded the rule. Widened to accept a Literal key regardless of computed, which is safe for MemberExpression too (its non-computed property is always an Identifier by grammar). - Added the missing RuleTester valid case for "a case-insensitive accessor call" (envGet(env, 'PATH')) from the issue's Done-when checklist. * docs: backfill changeset PR number for #3624 (PR #3976) --------- Co-authored-by: sim <sim@local> |
||
|
|
355c943b08 |
enhance(#3626): make CONTEXT.md seam claims checkable via a SEAM.*.enforced-by gate (#3975)
* feat(#3626): make CONTEXT.md seam claims checkable via SEAM.*.enforced-by gate Adds SEAM.<id>.owns / SEAM.<id>.enforced-by=lint-rule:<name>|test:<path> predicates to CONTEXT.md, generalizing the existing WORKTREE.SEAM.* shape, plus scripts/lint-seam-enforcement.cjs (wired into lint:ci) which fails when a declared single-owner seam names no existing, registered enforcement mechanism. Backs all six current module-level single-seam/ single-canonical-owner claims found in CONTEXT.md, including the Shell Command Projection Module's Windows-binary-resolution claim via #3619's local/no-private-binary-resolution rule. Scope is resolves-only per maintainer decision: the gate proves an enforcement pointer exists and is registered, not that its surface covers every file the seam claims. See docs/adr/3626-context-md-seam-claim-gate.md. Closes #3626 * fix(#3626): back the Package Identity Module's seam claim too Isolated adversarial review caught a miss in the "no grandfather list" sweep: the Package Identity Module also declares itself "Single seam owning GSD's published-package coordinates" and already names its real enforcement (scripts/lint-package-identity-drift.cjs). Backs it with SEAM.package-identity.owns/enforced-by=test:tests/package-identity.test.cjs, bringing the total to 7 backed seams. Also makes explicit, in the design doc and ADR, that function-level "single owner" sentences inside already-covered modules (STATE.md Document Module, etc.) are deliberately out of scope — a seam claim is about a module's boundary, not every function inside it. * docs(#3626): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
fa41bfec5c |
enhance(#3942): the emitted-drift ack is PR-lifetime data — move it to a commit trailer (#3954)
* test(#3942): failing-first suite for the emitted-drift ack commit trailer Binds 37 input classes from the phase test matrix to the behavior ADR-3942 specifies, before any of it exists. Stubs return benign empty values rather than throwing, deliberately: several rows assert that something DOES throw (cap overflow, uncomputable commit range), and a throwing stub would turn those green for the wrong reason and destroy the red. The two rows that carry the design's load: - merge-base semantics. The range is $(git merge-base base HEAD)..HEAD, not base..HEAD, because changedPaths comes from `git diff base...HEAD` (three dot). Two-dot would let the ack set and the change set disagree about which commits are this PR's. The fixture forks a topic branch, puts a trailer on each side, and asserts only the topic-side trailer is in range. - fail-closed on an uncomputable range. With fragments a depth-1 checkout passes VACUOUSLY, every fragment reading as brand-new. With trailers the range cannot be computed at all, and returning an empty set would silently disarm the gate, so it must throw. The fixture builds a genuine shallow clone rather than simulating one. Also covers the self-inflicted case: this change's own documentation quotes the trailer syntax, so an example landing at the end of a commit message would arm a live acknowledgment keyed on the literal placeholder text. Keys carrying angle brackets or whitespace are rejected. Authored per the phase artifacts 40-design.md and 50-test-matrix.md. Not yet run on the remote runner — this commit exists to be tested. Refs #3942 * chore(#3942): move the emitted-drift ack to a commit trailer Implements ADR-3942, superseding ADR-2719 section 3 and its #2789 amendment. Sections 1, 2 and 4-7 are retained: the conservation law is unchanged, only the storage of its escape hatch moved off the working tree. An acknowledgment explains one PR's ripple, and the moment that PR merges the ripple is in the base, so it can never clear anything again. It was stored in permanent shared state anyway, and every consequence of that mismatch had to be built and then maintained. The chain is #2789 -> #2914 -> #3078 -> #3842 -> #3823 -> #3875, each fix generating the next defect, ending in a scheduled sweeper whose own first PR could not merge itself. Added parseAckTrailers + renderAckTrailer (pure) and readAckTrailers (IO shell), reading Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: trailers over the merge-base range. tests/emitted-ack-trailer.test.cjs, 37 cases, written failing-first and confirmed red before any of this existed. Changed diffEmitted takes two structurally distinct key-space maps instead of one shared paths map. That closes a latent defect: the spaces were separated by convention only, so a growth key satisfied a hash lookup by naming coincidence. staleAcks now reports which space a key was declared in. REMEDIATION teaches the trailer, per space, with its example rendered through renderAckTrailer so the taught grammar cannot drift from what the parser accepts. Removed the sweep workflow, the guard-no-ack-on-next job, the standalone linter and its lint:ci entry, the fragment directory and its three spent fragments, the legacy single-file union, and the baseAck/spentAcks mechanism -- spentness is now structural, not computed. Two range properties carry the design and are pinned by tests rather than asserted: the range is merge-base scoped, matching git diff base...HEAD, so an already-merged trailer is out of range by construction; and an uncomputable range throws instead of reading as zero acknowledgments, which is the inverse of the fragment guard's vacuous pass. Three deliberate observable changes, each disclosed in the changeset: the unread runtime field is gone, the legacy file is no longer read, and cross-space excusal no longer works. Ten open PRs carry fragments and will meet a modify/delete conflict. Measured before landing and accepted deliberately; the one-line migration is in the PR body. Verified: lint:ci exit 0. Remote runner to follow on this exact sha. Refs #3942 * fix(#3942): silent trailer collapse, lost coverage, and an unbounded cap Six findings from the orthogonal review round, all fixed in place. BLOCKER -- two trailers of the same name on one commit collapsed silently. readAckTrailers built `separator=1d` where git needs `separator=%x1d`: the `separator=` value inside a %(trailers:...) placeholder is itself a pretty-format string, so the bare hex was emitted as two literal characters and the split on \x1d never matched. Two same-name trailers therefore joined into one value with errors empty -- the first reason absorbing the second entry's key. Silent truncation, the exact class MAX_ACK_TRAILERS throws to prevent. Confirmed with od -c against real git output before and after. The failing-first matrix did not catch it because its "both spaces coexist" row uses Hash plus Growth -- different trailer NAMES -- so the value separator was never exercised. Two regression tests now cover same-name trailers directly. Coverage recovered: normalizeAckReason and INVISIBLE stayed on the live path via parseAckTrailers but lost every test when the old suite was pruned. Back under test against the current surface -- all six invisible codepoints individually, whitespace collapse, trim, CRLF, and two seeded fast-check properties. Dropping any single codepoint now fails. MAX_ACK_TRAILERS counted raw trailers before de-duplication, so one trailer carried forward across rebased commits counted once per commit and could throw on a legitimate branch. Now counts distinct entries; 100 identical repeats dedupe to one. diffEmitted validated baseline, current and changedPaths but not the new ackHash/ackGrowth, so a bad shape raised an unhandled TypeError instead of an error verdict -- the same defect shape this file documents for #2778. Docs: CONTRIBUTING and TESTING-SUITES were rewritten only in their first sections; the later passages still taught fragments, git rm and the deleted guard, contradicting the new text directly above them. Finished. Also extends lint-removed-but-needed to exempt docs/adr and docs/research. That gate fails on any docs mention of a file deleted in the same diff, which makes it impossible to document a deletion in the PR performing it -- an ADR's whole job is naming what it retired. Exemption is narrow and comes with a test proving the gate still fires for a live consumer elsewhere under docs/. A guard that cannot fail is worse than no guard. Maintainer-approved. CONTEXT.md names the retired machinery by role rather than by filename: its generated projection lands in docs/, which that gate does scan. Adds docs/how-to/acknowledge-emitted-drift.md. The required docs set is Reference and Explanation, so the task quadrant can be empty with every gate green -- and this change has a real multi-step journey, including the fragment migration ten open PRs now need. lint:ci exit 0. Refs #3942 * docs(#3942): correct the duplicate-trailer rule in CONTRIBUTING Both axes of the code review independently flagged the same passage, without seeing each other's output. It claimed two declarations of the same key are always "a hard, loudly-reported error, not a silent last-wins". That is only half true, and the missing half is the one contributors hit: identical declarations -- same key, same reason -- dedupe silently, because a trailer legitimately survives a rebase and reappears on every rebased commit. Failing there would red a branch for doing nothing wrong, which is exactly why the dedup exists. Only a same-key/different-reason pair errors, and that one is a genuine ambiguity about which explanation holds. As written, the paragraph told a contributor that a rebase-carried trailer breaks the gate -- the opposite of the behavior. CONTEXT.md's parallel entry already stated it correctly; this brings CONTRIBUTING into line. Doc-only, root-level markdown. Refs #3942 * chore(#3942): backfill changeset PR number to 3954 --------- Co-authored-by: sim <sim@local> |
||
|
|
39673ae9ff |
fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config Regression tests (RED first): --skills-root and gsd-tools query surfaces, install-plan dest dirs, and converter skills-path rewrite. * fix(#3738): antigravity global skills/agents install to ~/.gemini/config Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents}; the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the ADR-1239 skills/agents 'home' override on the antigravity global layout — the same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references in converted global content to ~/.gemini/config/skills/. configHome, settings, probe/migration semantics, and the local .agents layout are unchanged. * fix(#3738): retire deprecated configHome artifacts via installer migration 010 Next install converges an existing antigravity install: manifest-managed skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does not scan) are removed — modified files backed up first, unmanifested and non-gsd entries preserved — and now-empty containers retired. Global scope only; the local .agents surface is live. Docs + inventory updated. * fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline - bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/ rewrite as src (ADR-1508 dual copy must stay in sync). - Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so antigravity's emitted skills/agents stay differential-visible at their new install root; install-tree fixture regen confirms an unchanged key set. - skills-from-commands rule declares the antigravity converter as a runtime-scoped transform; one ack fragment covers the identity-classed workflow whose antigravity copy embeds the old skills path. - Migration 010 checksum baseline + home-override set doc updated; existing tests updated to the #3738 contract (global dest, golden parity via layout dest, integration expectations). * fix(#3738): tolerate an absent extra emit root on baseline-side measurement The base tree's installer predates the home override, so <HOME>/.gemini/config does not exist there; walk() threw ENOENT and the in-job baseline build failed. An absent extra root is the legitimate pre-override shape — skip it. * fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment - writeManifest resolves the agents-kind home override (_kindDestDirSafe), so the manifest records agents at their actual install root and drift detection keeps working (isolated review finding 1, major). - Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing slash) divert to ~/.gemini/config/skills instead of falling through to the retired configHome path (finding 2). - real-home-guard comment updated: antigravity's global agents kind is the first agents-kind home override (finding 3, doc-only). - Regression tests for both behavioral findings. * chore(#3738): changeset fragment (pr number backfilled after PR creation) * chore(#3738): backfill changeset PR number (3921) * fix(#3738): sandbox HOME in tests that install antigravity global artifacts antigravity is the first home-override runtime in the golden-parity and skills-wrapper suites (codex is not in their runtime lists), so those tests never needed HOME sandboxing — the real-home guard now (correctly) refuses their un-sandboxed global installs on CI, where HOME is the passwd home. * fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans K3's two back-to-back sandboxHome calls leave HOME pointing at the first sandbox once the after-hooks restore (each call saves the env as it found it, so the second saves the first's sandbox as 'original'). On the windows matrix that leaked gsd-k3-qwen-* home into the L2 property, whose antigravity/global run then (correctly) refused via the #3712 real-home guard — antigravity is the runtime that made L2's plan escape into os.homedir(). K3 now manages the env with a single restore; L2 sandboxes HOME per run, mirroring L1. * fix(#3738): L2 property's HOME sandbox must exist on disk The #3712 guard's sandbox exemption fails closed when identify(effectiveHome) is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir under the real home) the antigravity/global run refused even with HOME sandboxed. Create the per-run sandbox dir and clean it up. --------- Co-authored-by: sim <sim@local> |
||
|
|
e20744eacb |
enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
878f25025c |
enhance(#3905): the exit-code registry — one number, one meaning, enforced at build (#3920)
* feat(#3905): the exit-code registry — one number, one meaning, enforced at build A generated registry replaces locally-invented exit codes. Every entry records code, name, meaning, owning module and the decision that authorized it. The generator refuses to build a table where two entries claim one code, two claim one name, a code falls in a range Node or the shell reserves, 2 is claimed by anything but the hook adapter, or an allocation carries no justification. exitCodeFor is pure and total: it throws rather than returning undefined, including for prototype-chain names. Inert by design — nothing emits a registered code until #3906. Every registered code is non-zero, asserted over the whole table, so a caller testing for failure behaves identically for pass and trips for everything else. * feat(#3905): make the registry generator's failures machine-readable Adds a --json mode carrying {ok, reason, context, detail}, where context is a typed payload naming the specifics the prose embedded - which code collided and under which names, which band rejected a code, which field was missing. The tests now assert on that structure instead of regex-matching the generator's stderr, which CONTRIBUTING prohibits, and the CONTEXT.md glossary gains the entry the issue's scope requires. * test(#3905): refresh the install-tree fixtures for the new declaration The registry declaration ships in the install tree, so all 19 golden fixtures needed regenerating. Caught by the remote matrix, not by lint:ci - the install-tree goldens are verified by a test rather than a lint, so a newly shipped file clears every local gate and fails only under the suite. * chore(#3905): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
7e9d33c378 |
enhance(#3915): stryker on the official tap runner, 29m to 12m on the critical path (#3919)
* enhance(#3915): stryker on the official tap runner, per-test coverage The 'command' runner is the one runner Stryker excludes from coverage analysis, which forced coverageAnalysis:'off' and made mutation cost strictly linear in (mutants x whole-shard test time). The frontmatter shard measured 1751s against 212s for the next slowest. Swap to @stryker-mutator/tap-runner (official plugin, peer-pinned to the @stryker-mutator/core@9.6.1 already installed) and turn coverage analysis on, so Stryker re-runs only the test files that cover each mutated line. Per-shard injection moves from a MUTATION_TEST_CMD command string to MUTATION_TEST_FILES, read by a new fail-closed resolveMutationTestFiles() that mirrors resolveMutationBreak. It stays derived from scripts/mutation-matrix.cjs and now has a single owner: the union moves behind allCoveredTests() and stryker.config.mjs no longer imports COVERED. The resolver existence-checks every entry, because the tap runner resolves tap.testFiles with glob() and a non-matching pattern yields an empty list silently - a fast, confident, meaningless run. tap.forceBail is off by measurement, not preference: a structural AST audit found 3 of 26 shard test files spawn subprocesses, and bail fires on every killed mutant, so leaving it on would kill processes mid-spawnSync and orphan their children. The matrix 'isolation' field is removed - per-file process isolation is inherent to the tap runner, so the field had no consumer left. Score arithmetic is unchanged and now pinned: mutationScore counts NoCoverage in the same denominator as Survived, which both Stryker's thresholds.break and check-mutation-score-ratchet.cjs read. A new non-vacuity test proves it diverges from mutationScoreBasedOnCoveredCode, so the gate cannot be quietly swapped to the field that would make every floor trivially satisfiable. Refs #3915 * fix(#3915): enforce the resolver's documented containment contract Review findings, all fixed in place. resolveMutationTestFiles claimed to verify each entry exists 'relative to the repo root' but used a bare fs.existsSync(path.join(...)), which accepts an existing DIRECTORY and lets ../ segments escape the root (path.join('/repo/root','../../etc/passwd') resolves to /etc/passwd). Not reachable from PR content - the value only ever comes from the static COVERED registry via CI env - but a fail-closed contract that overstates its own guarantee is a defect in the contract. Each entry is now resolved, rejected if path.relative puts it outside the root, and required to be a regular file. Three hostile-input tests added; the original missing-file wording is preserved so the existing assertion still binds. Also: removed a stale buildResult comment still naming the isolation field this branch deleted; hoisted one top-level path require in place of three inline ones; de-duplicated the derived-test-list expression in the tests to a single const, deliberately still re-derived from COVERED rather than calling allCoveredTests() so the assertion cannot become a tautology; tightened the workflow-parity assertion to exact equality; changed the tap-runner range from an exact 9.6.1 to ^9.6.1 so it tracks the caret-ranged core its peerDependency pins exactly. Regenerated examples/dynamic-context-management/CONTEXT-INDEX.json, which the earlier CONTEXT.md edit left stale - lint:ci was red on lint-example-parser-parity until it was refreshed. Refs #3915 * test(#3915): kill model-catalog survivors, set frontmatter budget from measurement From mutation run 33026833181 (all 13 shards, dispatched on this branch before any PR). WALL TIME: frontmatter measured 713s (11m53s) vs 1751s (29m11s) on the command runner, a 59% cut. timeoutMinutes 60 -> 20 (1.68x measured). Not deleted outright: the shared 15-minute default would leave only 21% headroom, and this module's mutant count grew 1.8x in one change. MODEL-CATALOG: came back 57.91 against its floor of 58. Diagnosed from the report JSON, not assumed - all 24 of its new RuntimeError mutants are in the load-time catalog bootstrap, so mutating them makes the module throw at require. Under node --test that is a failed test file (Killed); under the tap runner the process dies before emitting TAP, which Stryker classifies RuntimeError and excludes from the denominator. Add them back as killed and the shard is 248/416 = 59.62, the pre-change number exactly. Detection did not regress, classification changed. The floor is NOT lowered to absorb that. 11 new behavioural tests target ~42 genuinely surviving mutants: exact Set equality on the EFFORT_RENDERING/EFFORT_ARGV supported sets, an exact-string table render that makes the column-width arithmetic observable (padEnd never truncates, so the old substring assertion could not see a too-small width), clampEffortForHost's full null matrix plus a spoofed-toString host, and a prototype-pollution guard driven through Object.fromEntries. Mutants judged equivalent were skipped rather than papered over; reasoning is in the phase artifacts. All 15 new assertions verified against unmutated code. lint:ci exit 0. Refs #3915 * test(#3915): ratchet model-catalog's floor to 74 on measured 75.26% CI run 33029755081 measured model-catalog at 75.26% (295 killed / 49 survived / 48 no-coverage / 24 runtime-error, totalValid 392), so check-mutation-score-ratchet.cjs correctly failed the shard for unclaimed headroom: 17.26 points above the declared floor of 58, well past the 5-point slack. Floor raised 58 -> 74 (floor(75.26)-1, this file's documented convention), with RATCHET_BASELINE updated in the same diff as the equality assertion requires. Worth recording why the number moved this far. The shard first came back at 57.91 under the tap runner and the temptation was to lower the floor to match. The drop was not a regression: all 24 of its RuntimeError mutants sit in the load-time catalog bootstrap, and adding them back as killed reproduces 248/416 = 59.62, the pre-swap figure exactly. Rather than absorb a reporting artifact by weakening the gate, 11 behavioural tests went after the genuinely surviving mutants and killed 68 of them - carrying the module from 59.62 past its old ceiling to 75.26, within reach of the ADR-456 target of 80. The #3007 measurement is kept as clearly-labelled prior context rather than deleted, so the entry does not read as carrying two current numbers. Refs #3915 --------- Co-authored-by: sim <sim@local> |
||
|
|
6b7df61938 |
enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises ADR-3473 §8.1 carries a blocking open question with a forcing function: it must be answered before any implementation PR for the rule opens. Answered here as (a), a string-coercing adapter, with the measurement that settles it. The sequencing note bet that §8.8's schema would make (b) tractable. Measured against merged reality it does not: only 33 of extractFrontmatter's 78 non-test call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists survive real types, leaving ~31 lines across 3 call sites as the actual prize. Also corrects three claims verified false while answering it. §8.1's justifying sentence names #3349 and #3360 as defects a real parser would fix; both are already fixed on next, confirmed by executing the compiled parser rather than reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an expected casualty of this rule, but it guards shell grep idioms in workflow bash fences and never touches our parser. The same roster calls lint-vendored-deps.cjs reusable as-is; it is hardcoded to re2js throughout. The last two were caught by applying the rule this amendment records -- a factual claim in this ADR is a hypothesis until the implementing phase executes it -- on its first use. Refs #3881 * docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable An adversarial pass on the Phase 4 design established by execution that extractFrontmatter is not a YAML parser but a line-oriented scanner whose output is a function of raw source text. Four spellings of the same value collapse to one js-yaml tree but produce four distinct legacy strings, one of them mangled. No adapter over a tree can choose among outputs the tree does not distinguish, so fork (a) -- keep a string-coercing adapter so the existing contract holds -- cannot be built. For any document with a non-scalar value, (a) collapses into (b); about 26 percent of frontmatter-carrying documents have one. Also records three design defects and one new attack surface, all confirmed by execution: catching a parse failure and returning {} would delete the frontmatter block on the next write at eight call sites that conflate empty with unparseable; an empty value yields null where legacy yields {}, and reconstructFrontmatter omits null-valued keys, so the shipped state template's empty progress key would vanish; the #1882 truncation probe is parseYamlRegion itself rather than a pre-parse heuristic, so it cannot both stay unchanged and survive that deletion; and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB. The rule is not deferred. The measurement is the deliverable and the re-scoping is recorded as an open question with a forcing function, per section 8's own rule. Refs #3881 * test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed. Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states. B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today. B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today. B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today. Refs #3881 * chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496). gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own. src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare. scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored. docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds. Refs #3881 * feat(#3881): parse .planning frontmatter with the vendored js-yaml ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and parseQuotedScalar are deleted (not patched); parsing now goes through the vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA, json: true }. Everything js-yaml does not do is layered on top, in one place, carrying the seven design-doc consequences: 1. Empty value: a null js-yaml value is coerced to {} (matching legacy's own empty-value contract) so reconstructFrontmatter — which omits null-valued keys — still round-trips a bare `key:` line instead of deleting it. Verified live: progress: with no value survives parse -> reconstruct -> re-parse. 2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS channel, is carried on the {} returned for malformed/refused YAML. Invisible to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that never inspect it are unaffected; wiring the 8 hasFrontmatter sites to consult it is a separate change, not done here. 3. Non-scalar object-list items (the four spellings of `- test: a b` that js-yaml collapses into one tree shape) are rendered as a canonical `key: value[, key2: value2]` string per item, keeping the existing array-of-strings value SHAPE. A full corpus differential over all 1702 tracked markdown files found 11 residual divergences from the legacy parser (enumerated in the PR/report), most of them the parser now being MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key defect, a dropped Unicode key). 4. The #1882 truncation probe still runs the one real parser, but derives its key count from js-yaml's own thrown error and mark.line when the whole region doesn't parse cleanly (the dominant real truncation shape: fence opened, well-formed keys, no closing fence). Verified against both the clean-parse and the exception-fallback path. 5. The #3257 comment channel now attributes each pending column-0 comment against js-yaml's own parsed top-level key list (matched by literal key text, in document order) instead of the legacy ASCII-only key regex, so a comment above a Unicode key attaches correctly. 6. Anchors, aliases and merge keys are refused outright (a raw-text pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences today: zero. A 7-line billion-laughs fixture is verified refused rather than expanded. 7. A literal U+0000 is swapped for a private-use sentinel before the parse and restored in every resulting string afterward, since js-yaml rejects NUL unconditionally under every schema. escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump() (forced double-quoted style), with control-char hex escapes lowercased to keep serialized output byte-stable (#1779 emitted lowercase); it keeps its exported name and signature for its two other call sites (commands.cts, runtime-artifact-conversion.cts), which need no change. frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments, regenerateFrontmatterKey's guard, noOpObjectListSetError and parseMustHavesBlock are all unchanged — retiring them is fork (b) and is not this phase. Refs #3881 * fix(#3881): quote template placeholders and preserve unparseable frontmatter SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax under the vendored js-yaml parser, not the literal placeholder text intended. Quote them so they parse as strings. Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the 8 call sites in state.cts/state-transition.cts that compute hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and reassemble the document without a frontmatter block when false. That check conflated 'no frontmatter' with 'unparseable frontmatter' (both parse to {}), so a document with a merge-conflict marker or refused alias in its frontmatter had that block silently dropped on write. Each site now preserves the exact raw bytes stripFrontmatter removed when the marker is set, leaving the genuinely-empty case unchanged. Refs #3881 * test(#3881): consequence and boundary coverage for the js-yaml migration Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix. Refs #3881 * docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to Refs #3881 * docs(#3881): correct the frontmatter glossary entry Two errors in the entry as first written: it named parseYamlRegion as part of the read path when that function is deleted, and it recorded the eight hasFrontmatter call sites as unwired follow-on work when they were wired in e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the frontmatter block independently, so the marker binds at the transform layer. Refs #3881 * docs(#3881): record the semantic-migration decision and the counted guard ledger The maintainer chose the full semantic migration over splitting the rule into its own epic or patching the scanner, so section 8.1 is answered as "the fork was ill-posed and the migration is semantic" rather than as (a) or (b). Also replaces the pre-implementation guess that this phase would shrink the guard surface with the counted result: excluding vendored third-party lines the hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines despite four functions being deleted, because the compatibility layer over js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is therefore not delivered as written; what improved is the kind of code maintained, not the amount. Decision 6 requires recording that rather than netting it away. Refs #3881 * chore(#3881): changeset for the vendored YAML parser migration Refs #3881 * test(#3881): golden parity, round-trip property and packaging coverage Refs #3881 * fix(#3881): refuse anchors structurally and fold in review findings ADR-3473 §8.1 review findings, addressed inline: Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping ({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor mechanics while never matching that line shape, so the exact expansion the guard exists to stop went straight through unrefused (a 303-byte quoted-key bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener` callback, which reports `state.anchor` for every event belonging to an anchored node in every spelling, and throws from inside the callback to abort before any expansion (~1-2ms vs full expand-then-discard). A merge key with an alias is still refused (merge always requires a previously anchored node, so the alias itself trips the listener); a bare merge key with NO alias is no longer separately refused, documented as intentional: FAILSAFE_SCHEMA never resolves `!!merge`, so it carries no expansion risk. Table-driven tests added for all four bypass spellings + merge key, plus a quoted-key-spelled billion-laughs fixture registered in the adversarial matrix and README. Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/ aliases were "simply UNREACHABLE from typed code" through the twin. Corrected to state the truth: anchor/alias resolution is document-level `load` mechanics reachable through exactly the declared surface, and refusal is enforced at RUNTIME (Finding 1's listener), not by the type surface. Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree back to NUL, including one the document author legitimately wrote, silently corrupting it. Now refuses outright whenever the raw region already contains U+E000 (consistent with the existing anchor/merge-key refusal path), making the substitution provably injective. Tests added for a real NUL alone (preserved), a pre-existing U+E000 alone (refused, not corrupted), and both together (refused, not merged into one byte). Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead for a hand-authored row (only read inside the upstream-verbatim branch) — exactly how Finding 2's stale docblock drifted unnoticed. Added checkHandAuthoredTwin: every value-level export the twin DECLARES must be an actual own property of the vendored runtime module at require-time. Tests added, including a sensor that a declared-but-nonexistent export IS caught. Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/ unparseableFm/reassemble preamble, copy-pasted at 7 sites in state-transition.cts plus a sixth hand-inlined copy in state.cts's cmdStateCompletePhase, is now one exported helper (beginFrontmatterReassembly) every site routes through, including the hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore) keep a literal `body = stripFrontmatter(content)` assignment alongside the helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward scan (which does not chase aliases) still sees the strip; stripFrontmatter is pure/idempotent so the extra call changes nothing observable. Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a separate change" claim (the 8 call sites are wired on this branch) and the changeset's backlink from (#3473) to (#3881). Finding 7: fixed the lint:ci failures blocking the gate — an @typescript-eslint/only-throw-error violation from throwing a bare Symbol as the anchor-detected signal (now a real Error subclass), unused-var warnings left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by two migration-specific test files (allowlisted with justification), and the lint-state-write-path-drift false positive from Finding 5's helper (fixed above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already carried an explicit timeout; no change was needed there. Golden fixture: added a golden entry for the new anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line scanner would also produce, since it independently dropped every quoted top-level key). No other corpus document diverges: real .planning/ documents carry zero anchors/aliases/merge keys/U+E000 today. Refs #3881 * fix(#3881): fold in second-round review findings Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration ASCII-only key regex for the Unicode fixture; updated to require the 相 key's value now that js-yaml has no such restriction. Audited the rest of the file for other pre-migration pins (block scalars, quoted keys, flattened values, empty values, duplicate keys, unclosed blocks, null bytes) by execution against real fixtures; found none regressed. Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts so no function still answers to the deleted hand-rolled scanner's name (ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three external call sites (src/commands.cts, src/runtime-artifact-conversion.cts) updated in the same change — a mechanical rename, not an ADR-amendment matter. Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found by execution. A top-level key named constructor/__proto__/toString/ valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read resolving an inherited Object.prototype member); a key literally named __proto__ was silently DROPPED entirely (bracket assignment on an ordinary {} invoked the inherited __proto__ setter instead of creating a data property). Fixed by building every parsed Frontmatter object with Object.create(null), and replacing an `in` check with hasOwnProperty.call in propagateCommentChannel. Added round-trip tests for all five hostile keys, each with its own leading comment. Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed full byte-stability across the migration. Verified by execution: BEL/NUL/ NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw- literal forms. Proved round-trip equivalence (each escape re-parses to the exact source codepoint) and corrected the docstring. Found and fixed a related real defect while verifying: a lone UTF-16 surrogate was emitted BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely unparseable YAML that silently collapsed to {} on re-read — extended scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped path. Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real truncation shapes (unquoted colon, open flow collection, mis-indented sibling key, refused anchor). Root cause: the mark-based prefix recovery excluded the very line whose key needed counting, and a mark-less refusal never entered the recovery branch at all. Fixed by taking the max of two lower bounds: the longest parser-verified line-prefix, and a raw-text count of key-shaped lines (reusing the same key-shape pattern this file already uses for isFrontmatterShaped). Extended test-matrix row A5 table-driven over all 4 regressed shapes. Finding 6: the design doc's claim that no test owned the #3594 adversarial fixture corpus was false — consolidation epic #1969 had already folded it into tests/frontmatter.test.cjs. An earlier commit on this branch re-created a standalone duplicate under that false premise; folded its genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures, block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the ADR's §8.1 note, including the roadmap-sibling claim (no such file exists). Finding 7: the golden serializer sorted object keys, making it structurally blind to the key-order-parity invariant ADR-3473 §8.1 actually claims. Made it order-preserving and regenerated the golden fixture from a standalone compile of the legacy (pre-#3881) parser at ddde001af; the current parser matches it with zero undocumented divergences, confirming key-order parity genuinely holds. Extended row A2 table-driven across 6 of the remaining 7 transitionCore kinds (all pass) plus documented, by execution, a newly-discovered 8th-site regression: state.cts's cmdStateCompletePhase calls the same preservation helper but its result is clobbered by a later unconditional resync — filed as a distinct finding rather than fixed here (touches syncAndPreserveStateMd, outside this change's verified scope). Refs #3881 * fix(#3881): preserve unparseable frontmatter through the CLI write path Characterization (executed, before/after shown): case (b), not (a). The frontmatter FENCE survives — `state complete-phase` on a conflict-marked STATE.md returns success and a well-formed, freshly-derived frontmatter block, not a document with no frontmatter at all. But the block's actual content (the merge-conflict markers, and with them any signal to a human that the document was in conflict) is silently discarded and replaced. Root cause was two clobber sites, not one: 1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved `transformedContent` from readModifyWriteStateMd, finds {} + the FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh frontmatter block from the body anyway. 2. Even after (1) is fixed, applyPostSyncPreservation's own postFm/applyStatePreservation/authoritativeFm-reassertion machinery re-extracts frontmatter from syncedContent, restores curated fields from the pre-write snapshot, and reconstructs a NEW block again — confirmed live via `state begin-phase`, which still lost the markers after fixing (1) alone. Both are now guarded by the same predicate (isUnparseableFrontmatter, checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not parse and the caller is not on ADR-3408 §8.3's closed "body wins" list, both functions return their input content unchanged rather than re-deriving over it. The closed list (cmdStateSync #905, /gsd-health --repair's REGENERATE_STATE, both routed only through writeStateMd, which never reaches applyPostSyncPreservation and passes sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is untouched — neither widened nor narrowed; verified by execution that `state sync` still overwrites the conflict-marked block exactly as before. Other verbs sharing the same readModifyWriteStateMd path were checked and were equally affected before this fix: state update, query state.patch, and state begin-phase all lost the conflict markers (RED, shown by execution), and all three now preserve them (GREEN). Covered table-driven in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe block, which drives the real CLI verbs via runGsdTools — not just the pure transitionCore layer the earlier A2 rows exercised — plus a control asserting state sync's body-wins contract is unchanged. Refs #3881 * fix(#3881): restore the parse surface's prototype and fix remote-runner failures Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched. Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write. Refs #3881 * chore(#3881): backfill changeset PR number Refs #3881 * test(#3881): make golden parity resistant to unrelated tree churn A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus. Refs #3881 * test(#3881): make the parser golden hermetic instead of tree-keyed This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md, docs/*.md) the prior golden pinned by tracked path. Any PR editing one of those files' frontmatter for reasons unrelated to the parser (an argument-hint addition, an allowed-tools tweak) turned the suite red, and the reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to catch a real regression. Excluding .changeset/** was not enough; the design itself was wrong: a regression fixture must not be keyed to mutable repo paths, and a single 376-entry JSON every such PR touches is also a guaranteed merge-conflict surface. Rebuilt the fixture to carry its own documents: each of 51 entries stores a stable id, literal documentText (shrunk from a real ddde001af-era corpus document), and an expectedParse captured independently from the pre-migration legacy parser (git show ddde001af:src/frontmatter.cts, compiled standalone against its byte-identical sibling modules). The test reads no tracked path, shells out to no git command, and enumerates no tree -- a PR editing commands/gsd/help.md cannot affect it. Every entry's reconstruction was verified at capture time to reproduce both the current and legacy parser's output on the original document; 0 of 51 candidates were dropped by that check (1, the deliberately-unterminated unclosed-block.md adversarial fixture, has no closing fence to truncate at and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now diverges:true entries) and the D2 order-preserving structural serializer that keeps the comparison from passing vacuously; dropped the tree-enumeration helpers, the coverage floor, the post-capture-skip logic, and the vanished-file check -- all artifacts of the path-keyed design. Refs #3881 * fix(#3881): resolve vendored-deps paths independently of cwd shape Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on windows-latest CI: the test passed absolute scratch-file paths into compareFiles()/checkRow(), whose helpers joined every input onto ROOT via path.join(ROOT, rel), producing garbage when the input was already absolute. It surfaced on windows-latest specifically because GitHub's Windows runners checkout the repo on a different drive than TEMP, so path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged (no relative traversal is representable across drives) rather than the relative form the test assumed. The remote gsd-test runner this repo gates pushes on is Linux-only and could never have caught this; GitHub CI's windows-latest job is the only signal that does, and it did. Fixed the helper itself (scripts/lint-vendored-deps.cjs's new resolvePath()) to treat an already-absolute input as absolute-in, absolute-out instead of silently mis-joining it, and updated the test to pass the scratch file's absolute path directly rather than relying on a relative conversion that is not always representable. Kept every mutation-sensor assertion intact and added coverage proving resolvePath is a no-op for relative inputs and correctly passes absolute ones through unchanged. Refs #3881 * fix(#3881): warn when state sync regenerates over unparseable frontmatter state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly overwrites an unparseable frontmatter block per its 'body wins' contract — that overwrite behavior is unchanged here. The defect was the silence: synced:true/exit 0 gave no signal that the existing block (including git merge-conflict markers) could not be parsed and was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be reported as authoritative when the derivation dropped input it could not resolve') and §8.4 ('failure is a value'). Adds a gsd: warning — ... (#3881) line on stderr, matching the existing #3573 precedent, and surfaces the same disclosure in the JSON result's existing changes[] array so a machine consumer sees it too. Exit code and synced:true are left unchanged — sync did what its contract says. REGENERATE_STATE (/gsd-health --repair's sibling on the same sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally refused by applyRepairs's dispatcher before runRepairAction ever runs (src/health-diagnostic.cts), so it is not a live path today and is not in scope for this fix. Refs #3881 * fix(#3881): exit non-zero when a state command returns an error Refs #3881 * chore(#3881): changeset for the state exit-code fix Refs #3881 * fix(#3881): honor the documented --project-dir flag Refs #3881 * revert(#3881): restore exit-0 result envelopes for state errors Reverts 9638f2936 and its changeset. The change was wrong and the revert is the correction. This repo distinguishes two error mechanisms deliberately. error() in src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure path. output({error: ...}) writes a JSON result envelope to stdout and returns normally with exit 0. The reverted commit converted 23 result-envelope sites into hard failures, which is a different contract, not a bug fix. tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope contract directly -- a failing command exits 0 with a JSON error envelope and must not publish state.json -- and the remote matrix run caught it along with four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that the original commit rewrote were encoding that real contract, not the bug it claimed; they are restored. Whether an error envelope on stdout with exit 0 is the right CLI design is a genuine question, and it is section 8.4's rule ('failure is a value') with its own phase. It is not something to flip inside this PR. Refs #3881 * chore(#3881): backfill changeset PR number for the project-dir fix Refs #3881 * test(#3881): keep the frontmatter mutation shard inside its time budget The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard cap. Root cause is NOT row-level spawn overhead (contrast the #2790/ core-utils precedent): the three shard test files' own logic runs in ~413ms total (356+30+27ms) with all 392 assertions passing. Instead, src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to the vendored YAML parser, proportionally growing the mutant count Stryker generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner bills the full 'node --test <3 files>' invocation once per mutant, and node:test's default per-file process isolation forks a child process for each of the three files on every one of those invocations — pure fork overhead multiplied by a much larger mutant population. Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares isolation: 'none', and .github/workflows/mutation.yml passes --test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e. unchanged behavior — for the other 8 shards, which were not individually audited for cross-file state leakage under shared-process execution). Measured locally via node:test's run() API on the exact 3-file set: isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same 392 passing assertions. The true CI-shard number can only be confirmed on the GitHub Actions run (Stryker cannot run locally, and 'node --test' is hard-blocked in this environment). Refs #3881 * test(#3881): register the vendored-parser tests in the frontmatter mutation shard stryker.config.mjs's own rule ("Keep this list in sync with the tests arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew src/frontmatter.cts from ~825 to 1496 lines but its new tests (tests/feat-3881-yaml-parser-consequences.test.cjs, tests/frontmatter-golden-parity.test.cjs, tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in tests/frontmatter.test.cjs) were never added to the frontmatter shard's tests array, so Stryker's mutants in the new vendored-js-yaml adapter had nothing constraining them. PR #3888 measured 55.8% against the 65 floor (748 killed / 593 survived / 17 timeout) and the shard was separately cancelled at 15m04s against the 15-minute per-shard cap. Registers all four files (each earns its slot on evidence of a unique constraining assertion, documented inline), gives the shard a measured/projected 180-minute budget via a new per-module timeoutMinutes field threaded through mutation.yml's job-level timeout-minutes the same way isolation is threaded, and removes the prior isolation:'none' override (re-measured at this file-set size, its savings are within run-to-run noise, not worth the unaudited cross-file-state-leakage risk). Refs #3881 * feat(#3881): derive the mutation test list and ratchet the score floor Refs #3881 * test(#3881): ratchet five stale mutation floors and close the frontmatter gap Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1): config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78, context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs RATCHET_BASELINE in the same diff per the ratchet's own contract. Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter), repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via extractFrontmatter), and the null-byte sentinel round-trip surviving at region offset 1. Did not lower minScore. Refs #3881 * test(#3881): decouple the ratchet test from real module floors The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour. Refs #3881 * refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency Refs #3881 * fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives. Closes #2571 Refs #3881 --------- Co-authored-by: sim <sim@local> |