Commit Graph

242 Commits

Author SHA1 Message Date
Tom Boucher
9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
90771ddf02 enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable

ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.

Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.

The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.

Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.

Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.

Closes #2904

* fix(#2904): bound and sanitize untrusted registry `interactions` strings

Review findings from the pre-PR passes.

Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.

`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.

Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.

Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.

* chore(#2904): backfill changeset PR number (#2912)
2026-07-31 08:11:57 -04:00
Tom Boucher
4bd6fb066b chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam

The no-adhoc-markdown-parsing rule matched only a negated class whose sole
member was a pipe ([^|]), so the stricter and more common [^|\n] spelling
evaded it entirely -- src/state-document.cts hand-rolled exactly that shape
and linted clean. Widen the fingerprint to any negated class excluding a
pipe, which is the ADR-2143 section 7 prohibition as written.

With the rule fixed, state-document.cts goes red. Replace tableRowPattern
with locateFieldRow: a line scan using the markdown-table seam's
splitTableRow for cell semantics, returning the value cell's byte range, and
splice that range instead of running a whole-document content.replace. An
edit now physically cannot cross a row boundary (section 4).

Behavior is frozen -- stateReplaceField has 79 dependents across 5 command
processes. Characterization tests lock all 14 table-branch rows plus CRLF,
extract round-trip and the withFallback caller shape; a fast-check property
asserts every non-target line stays byte-identical.

Refs #2880, epic #2143

* fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint

Isolated adversarial review found four defects in the first commit.

1. locateFieldRow split lines on \n only. JS treats a lone \r as a line
   terminator, so the regex it replaced matched rows separated by bare CR.
   "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now
   CR, LF and CRLF are all terminators, byte offsets unchanged.

2. The field name was normalised with trim().toLowerCase(). The old regex
   embedded it verbatim, so its whitespace had to be absorbed by the row's
   own padding -- and because the group is a literal-character match rather
   than a whitespace class, a tab-padded cell does not accept a
   space-padded name. Replaced with an offset-aligned search reproducing
   the original backtracking exactly.

3. The widened fingerprint regex had two unbounded [^\]]* around an
   optional and ran quadratically over every regex source in every linted
   file: 256000 chars took 23 seconds. Replaced with a single-pass scanner
   that never rescans; the same input is now ~1ms.

4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|].
   Narrowed to a class excluding the pipe plus only \n, \r or \t.

Differential fuzz against origin/next: 20000 cases, 0 mismatches.

Refs #2880, epic #2143

* test(#2880): drop wall-clock assertion from the ReDoS regression guard

local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md
bans timing assertions outright as flaky. The 256000-char input stays as
the regression guard for the quadratic scan; correctness of the verdict is
what is asserted. If the quadratic path returns, the test stops completing
and surfaces as a suite timeout rather than a silent pass.

Also adds the changeset fragment for #2880.

Refs #2880

* fix(#2880): spec-correct case folding, property tests, naming

Code-review findings.

The field-name comparison used toLowerCase(). The regex it replaced used
/i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a
non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII
K where the old code returned null. Replaced with spec-correct
Canonicalize, including the multi-character uppercase case (eszett -> SS),
which a naive uppercase comparison also gets wrong.

Added the fast-check property tests CLAUDE.md requires for parsers: one
for the negated-class scanner, one for the field-name fold semantics, each
against an independent reference implementation. Both reference impls
failed on first run against real bugs, so neither property is vacuous.

Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a
duplicated comment to a cross-reference.

Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the
harness proven to discriminate the KELVIN case.

Refs #2880

* chore(#2880): backfill changeset PR number (#2889)

* docs(#2890): correct the local ESLint plugin path in CONTEXT.md

CONTEXT.md named the local AST-rule plugin directory as
scripts/eslint-rules/, which does not exist. The real location is
eslint-rules/ at the repo root -- what eslint.config.mjs actually
imports -- and CONTEXT.md's own later entry already says so
explicitly, so the file disagreed with itself.

Found by a line-by-line audit of all 1036 lines against the live
graph; this was the only confirmed inaccuracy.

Closes #2890

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:55:03 -04:00
Tom Boucher
7372d99a26 enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales

The reviewer lane roster was hand-enumerated across five documentation
surfaces and three workflow files that had drifted apart: --kimi-code was
missing from all four translated COMMANDS.md mirrors, --coderabbit from
every workflow forwarding list, and --antigravity from FEATURES.md.

Adds checkReviewerDocsParity, a second pure gate deliberately separate from
checkReviewerLaneParity so a stale doc cannot make the runtime checker look
red. Workflows now derive their flag lists from a new review-lane flags
query instead of hand-enumerating them, which also retires the unanchored
grep that matched --agy inside --antigravity.

Documents the previously absent reviewer body and hostBehaviors field in
the capability manifest reference.

Closes #2800
Closes #2781
Closes #2272

* fix(#2800): key the docs parity table arm on first-cell position

Review found the flag arm was file-scoped, so the forwarding row that lists
every flag in its third cell satisfied it on its own. Deleting a lane's own
reviewer-table row -- the #2781 regression this gate exists to prevent --
therefore passed undetected.

Arm 4 keys on the FIRST table cell, which separates a lane row from the
forwarding row structurally and in every locale. Regression test included.

* fix(#2800): shape-filter the flags subcommand output

All three consumers read review-lane flags through an unquoted command
substitution so the output word-splits into loop items. Phase 2 admits
third-party overlay lanes, so an overlay flag containing whitespace would
inject a second loop item and one containing a glob would expand against
the cwd. Emit only well-formed flags so neither reaches the shell.

* fix(#2800): remove the regex length ceiling and count only prose mentions

Review found two real defects in the docs parity gate.

The never-throws contract was false: building a RegExp from a declared flag
or section title throws SyntaxError past ~100k chars, and Phase 2 admits
overlay lanes whose declared strings are untrusted in length. Every one of
these matches is literal, so String.includes replaces the regex outright,
which also deletes escapeLiteral and the llama.cpp escaping it existed for.

Arm 1 was context-blind: a flag mentioned only inside a fenced example or a
commented-out row counted as documented. Both are stripped before matching.

Also advertises all 13 lane flags in the argument-hint and corrects a stale
eleven-lane count in the slug grammar note.

* test(#2800): repoint the convergence suite off deleted workflow text

The derived flag loop deleted the literal per-flag grep lines four tests
matched on. Two of those failed loudly. The behavioral and property tests
failed SILENTLY instead: their end marker no longer resolved, so the parse
block extracted empty and both passed vacuously, and the property test's
gsd_run stub had a no-op default that hid it.

All now share one extractor and execute the real deployed block through a
gsd_run shim backed by the actual binary. The whitelist assertions become an
anti-parity check: re-adding a hand-written flag list must fail.

Also repairs two vacuous cases in the docs parity suite. The unreadable-doc
test called its own mock rather than the reader, and the integration test
bounded nothing, so a doc losing its marker would have been silently skipped
and still passed green.

* fix(#2800): run the derived flag loop after the launcher preamble

The remote matrix caught a real runtime bug, not a test artifact. In
autonomous.md and plan-review-convergence.md the launcher preamble that
defines gsd_run lives in a separate, LATER bash fence than the derived loop.
Each fence is its own shell, so gsd_run was undefined where the loop ran:
the command substitution yielded nothing and zero reviewer flags would have
been forwarded. Worse than the drift this epic fixes, and silent.

The whole CONVERGENCE_ARGS construction moves as one unit, because the
--max-cycles append sits between the loop and the preamble and would
otherwise have run against an uninitialized variable and then been dropped
by the relocated initializer.

Also documents all 13 lane flags in help/modes/full.md, which the repo gates
bidirectionally against each command's argument-hint.

* test(#2800): repoint the two converge suites off deleted flag literals

Both asserted workflow.includes('--codex') against the hand-enumerated list
the derived loop removed. They now assert the derivation itself, keep --all
and --text (convergence controls, still literal), and add an anti-parity
guard so re-adding a hardcoded list fails.

The lost pass-through proof is replaced with a real one: every flag the
tests used to hardcode is asserted present in the actual roster emitted by
the binary, which is the property the old assertion was protecting.

* test(#2800): acknowledge the workflow byte growth from the derived flag loop

* chore(#2800): backfill changeset pr number to 2882

* fix(#2800): strip HTML comments to a fixed point in the parity gate

CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the
single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick
construction smuggles a commented-out row past the gate and it counts as
documented. Not an injection risk here since nothing is rendered, but it is
the exact false pass this helper exists to prevent.

Strips to a fixed point, then treats any surviving opener as unterminated so
the multi-line branch closes it on a later line. Terminates because every
pass strictly shortens the string.

* test(#2800): pin the comment-smuggling regression with a real reproducer

The obvious fixture for this class does not reproduce it: <!--<!---->-->
leaves a dangling --> rather than a live <!--, and is caught either way, so
it would have passed with and without the fix. The join-trick construction
(<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely
regresses on the single-pass strip and is what the test now uses.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:14:13 -04:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00
Tom Boucher
8b44a0da43 chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract

Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as
the declared contract for all 11 cross-AI reviewer lanes, and the
DEFECT.GENERATIVE-FIX parity assertion the roster has never had.

The lane contract lived in three unrelated surfaces — the roster, ~640
lines of hand-authored per-CLI bash in invoke_reviewers, and the
write_reviews section headings — so cross-cutting fixes landed per-leg
(#2494 and #2605 were the same empty-output defect filed twice).

- src/review-lane-descriptor.cts: frozen table declaring per lane the
  slug, flags, probe, invoke shape, timeout floor, empty-output policy,
  REVIEWS.md section, evidence class, required binaries, prompt-budget
  key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so
  Phase 2 harvests the shape with no translation layer. It declares;
  it does not execute — invoke_reviewers iterates in Phase 5b.
- checkReviewerLaneParity: bidirectional parity across descriptor,
  roster, invoke_reviewers legs and write_reviews sections. Forward-only
  would miss the failure it exists to catch (#2718 added a leg, #2781
  was the drift). ADR-1517 instance headings are exempt per D8.
- Legs carry an explicit <!-- reviewer-lane: slug --> marker; five
  non-lane bold labels share the bold-then-fence shape a heuristic
  matcher would key on.
- ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an
  error in both the core module and the workflow prose that mirrors it.
  A code-only change would be unobservable — the module has no
  production caller; the workflow narrates the policy. Discovery paths
  (--all, review.default_reviewers) stay lenient.
- Fixes the qwen leg, the last one discarding stderr to /dev/null.

Two ADR-2782 D2 vocabulary widenings were forced by surveying the
shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and
outputChannel 'file-arg' (Codex writes via -o and discards stdout,
#1698). Both are additive and closed; Phase 2 owns the validator.

Closes #2690

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): make the parity checker total and pin the lane slug grammar

Findings from the orthogonal review passes.

Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2
"verbatim" while diverging in three undisclosed ways, which is the
translation layer Phase 2 was supposed to be spared:
- `transport` moves from `invoke.transport` to the LANE level, a sibling
  of `probe`/`invoke`, exactly as D1's manifest example places it. The
  nested form read better as a TS discriminated union; the union is now
  discriminated at the lane level instead, which costs nothing.
- The header and the CONTEXT.md glossary now enumerate all FOUR
  widenings (adding `outputArg` and `flags[]`), not two.

Standards axis — CLAUDE.md requires a fast-check property test for a
parser, and `checkReviewerLaneParity` parses markdown for markers and
headings. Adding one found two real defects that the hand-written
matrix missed:
- NOT TOTAL: a malformed descriptor entry threw on `lane.flags`
  iteration, contradicting the module's own "never throws" claim. Every
  field is now narrowed from `unknown` at the trust boundary and
  reported as MALFORMED_LANE / INVALID_SLUG. This matters because
  Phase 2 feeds this function third-party overlay data, and a parity
  gate that crashes is indistinguishable from one never run.
- SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a
  slug outside that class was unmatchable — its marker could be present
  and correct and the scan would still report LEG_MARKER_MISSING
  forever. LANE_SLUG_RE now pins the grammar and a violating slug is
  reported INVALID_SLUG. A loud named violation beats a silent miss.

Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371):
seeding from the module's own matchers could only produce documents
those matchers already recognize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): register the new bin/lib module in the ESLint ignore list

The remote runner caught this; lint:ci did not, because the invariant
lives in the test suite rather than the lint chain:

  tests/repo-invariants.test.cjs
  "each bin/lib/*.cjs is linted xor ignored according to migration state"
  -> tsc-generated bin/lib modules not yet added to ESLint ignore list:
     review-lane-descriptor.cjs

Adding a src/*.cts module ripples to six surfaces (.gitignore, the
ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md
glossary, the capability/inventory manifests, and any size baseline).
The other five were covered; this was the miss.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced

Building the Phase 1 descriptor table against all eleven shipped legs is
the first time every lane's contract was written in one place, and it
surfaced four cases the ADR's original survey did not cover. Amending
the design lock rather than diverging from it, so Phase 2 (#2795)
implements the manifest validator against the amended vocabulary instead
of rediscovering the gaps.

All four are additive widenings of closed enums; no decision reverses:

- D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it
  reviews the working-tree diff.
- D2 outputChannel gains `file-arg` — the ADR called a file-writing lane
  a shape a real CLI *could* take; codex already is one, writing via
  -o/--output-last-message and discarding stdout (#1698).
- D2 gains `outputArg`, required iff file-arg — knowing the review lands
  in a file is useless without the argument naming it.
- D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes —
  antigravity is selected by both --antigravity and --agy, which a
  single-valued field cannot express.

This is the same evidence path that produced the openai-http transport:
the vocabulary widens on a lane that exists, under review, never on
speculation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2794): backfill changeset pr number to 2820

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:31:32 -04:00
Tom Boucher
9624167eec fix(#2810): accept the documented effortSurface axis on EoS registry entries (#2813)
* fix(#2810): accept the documented effortSurface axis on EoS registry entries

The EoS registry schema required an exact eight-key `interactions.axes`
object, while `docs/registries/README.md` and `CONTEXT.md` both documented
nine keys including `effortSurface`. An entry that faithfully mirrored its
upstream descriptor's `effortSurface` key was rejected outright.

`effortSurface` reached the runtime-descriptor vocabulary through ADR-1239
amendment #2481 (`HOST_INTEGRATION_AXES`), but the registry's hand-maintained
copy of that vocabulary never picked it up. The runtime-descriptor surface is
guarded by tests/host-integration-validator-parity.test.cjs; the registry copy
had no equivalent guard, which is what let the two drift.

Accept `effortSurface` as an OPTIONAL ninth axis validated against the
canonical ['argv','none'] rather than a required one: registry entries mirror
their upstream registry/eos-entry.json byte-for-byte, so requiring it would
retroactively invalidate every entry published before the amendment.

Adds tests/registry-axes-parity.test.cjs, which asserts that every key shared
between the registry vocabulary and HOST_INTEGRATION_AXES has an identical
enum array, plus limit-1/limit/limit+1 boundary coverage on the axes key set.

Closes #2810

* test(#2810): fail when a canonical axis is added but never mirrored

The enum-equality assertion compares only keys the registry and
HOST_INTEGRATION_AXES already share, so it is blind to the exact drift that
produced #2810: a new canonical axis appears and the registry copy is never
told. Verified by simulation — mutating an enum is caught, adding a new
canonical key is not.

Assert instead that every HOST_INTEGRATION_AXES key is either modeled by the
registry or named in an explicit NOT_MODELLED allowlist (subagentToolkit and
isolation, both dispatch sub-fields the registry collapses into its free-form
dispatch summary). Adding a canonical axis now fails until someone decides
which bucket it belongs in. The allowlist is itself guarded against going
stale.

Refs #2810

* fix(#2810): harden the axis value lookup with the CodeQL barrier pattern

Both orthogonal reviews flagged the same line: `AXES[key] !== undefined`
is not an own-property test, and the bracket reads are shaped like a
prototype-pollution sink even though the unknown-key gate above provably
makes them unreachable.

Switch the presence test to `Object.hasOwn` and add the repo's inline
literal guards (`capability-state.cts:146-155`, "Prototype-pollution guard
(inline literal, CodeQL barrier)"), which CodeQL can follow where it cannot
follow the `.includes()` filter that actually does the work.

Behavior is unchanged — re-verified all five axes key-count shapes plus a
genuine own `__proto__` property built through JSON.parse (the shape a
third-party registry PR would submit): it is rejected as an unknown key and
Object.prototype is untouched.

Refs #2810

* chore(#2810): backfill changeset PR number
2026-07-29 07:00:35 -04:00
Tom Boucher
1e3c995e6f fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it

Every input to `diffEmitted` is base-relative -- `baseline` vs `current`,
`changedPaths` from `git diff base...HEAD` -- except the ack set, which
was read absolutely, from the working tree only. A differential machine
consulting a non-differential input.

So `staleAcks` asks exactly one question, "did a delta consume you?", and
that cannot distinguish an ack that never explained anything (an
authoring mistake) from one whose ripple is now absorbed into the base
(the ack's SUCCESS condition). After merge an ack is in the second state
but reports as the first.

The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request
events only, so a push to `next` falls through to origin/next -- the very
commit under test. Both sides build identical content, no deltas remain,
and every live ack is reported stale. PR #2768 acked a deliberate 40866
-> 42020 byte growth, was green on its own lane, and reddened `next` the
moment it merged. It also reds every PR branching off the poisoned base,
and since publish-emitted-baseline is gated on the test job, it blocked
baseline publication too.

Give the ack the base side it was missing. `diffEmitted` now takes
`baseAck` -- the same document at the base ref, via `readAckFileAtRef`.
An entry already present there is SPENT: it may no longer consume a delta
and is never reported stale, only surfaced as `spentAcks` for tidying. An
entry new or reworded in this diff stays live, and if nothing consumes it
that genuinely fails, with blame on the author who just wrote it.

This closes a hazard the IMPLEMENTATION named but could not prevent -- a
leftover ack silently pre-clearing the next ripple on its path. (ADR-2719
§3 asserted only that TOUCHING the file is the alarm; its residual-risk
list never covered pre-clearing, and §3 now carries an amendment.)
Verified against the two-PR laundering sequence -- land an innocuous ack,
then change the artifact -- which passed silently before and now fails on
both the hash pass and the size ratchet.

Three things the design has to get right, each of which was wrong first:

  - A read failure on the base document THROWS; only absence-at-the-ref
    returns null. Returning null on error LOOKS armed (every entry stays
    live) but a live entry's defining power is that it CONSUMES a delta,
    so null is armed on the staleness axis and DISARMED on consumption --
    silently the whole pre-#2789 gate. `git show` cannot tell absence
    from fault, so absence is established with `ls-tree`.
  - Re-arming a spent ack costs actual PROSE. Internal whitespace and the
    zero-width family collapse, and `runtime` is not compared: a doubled
    space, an invisible character, or a decorative field would otherwise
    re-arm an ack whose justification still describes the previous
    ripple, showing a reviewer nothing.
  - `baseAck` is REQUIRED once an ack declares entries -- omission is an
    error, not a silent "inherit nothing" -- so a dropped argument fails
    loudly instead of quietly restoring this bug with the suite green.

Because a corrupt document ON THE BASE is expensive (the loud base-side
failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs
blocks one from landing. It is standalone rather than importing parseAck
-- scripts/ ships in the npm package and tests/ does not -- so a parity
test runs both surfaces over one corpus and fails on divergence; it
caught one immediately, a `null` document, now classed as policy rather
than schema. Deadlock is separately foreclosed: a tree carrying no ack
never reads the base, so the PR that DELETES a corrupt file still lands.

`readAckFileAtRef` takes an injected git runner so all four branches are
tested deterministically; it never executes in the remote runner, where
the real-tree test skips for want of a base ref. It also refuses an
option-shaped ref, since execFileSync's array form stops shell
metacharacters but not git's own option parsing.

Rejected: skipping the differential when base == HEAD. It treats the
symptom, costs real coverage on the push-to-next lane, and does nothing
about the downstream PRs the same flaw was reddening.

Deletes the now-spent tests/emitted-drift-ack.json, and updates the
CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a
LIVE entry is, and a spent one is inert.

Closes #2789

* chore(#2789): backfill changeset PR number
2026-07-28 21:28:10 -04:00
Tom Boucher
e276cc7f00 enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate

check-glossary-refs asserts that every backticked tests/ token in
CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719
section 3) is absent on a healthy next BY DESIGN — it appears only
inside a PR that needs it, which is what makes touching it the alarm.

It passed before only by accident of backtick pairing: CONTEXT.md's
RULESET entries are themselves backtick-wrapped and contain backticks,
so the token happened to fall outside a code span. Any edit that
shifted the parity exposed it. A gate that passes by luck is not
passing.

The exemption is exact, not a prefix hole: a sibling missing tests/
path still fails, and a test locks that.

* feat(#2778): make the size-ratchet failure name its own remedy

The growth branch stated a requirement and withheld the means of
satisfying it: no ack file named, no schema, no key format, and no
do-not-regenerate line — so the likeliest guess was to hunt for a
baseline that #2724 deleted. Observed live on #2543.

All remediation now comes from one frozen REMEDIATION export whose
example document is rendered from ACK_VERSION, so the taught schema
cannot drift from the schema parseAck accepts. A round-trip test feeds
the printed document back through parseAck.

The report is now built as a typed IR (buildReport) that formatReport
renders, so tests assert on structure rather than prose, per
CONTRIBUTING.md's raw-text-matching rule.

Two defects found and fixed inline while building:
- diffEmitted's validation early-return omitted newFileCapExceeded
  while formatReport reads its length, so the branch that reports a
  failed git diff threw a TypeError instead of naming the problem.
- Printing one complete ack document per failing branch made each read
  as the whole file, so pasting the second over the first silently lost
  an acknowledgment. One document now covers the whole report.

Closes #2778

* chore(#2778): backfill changeset pr number to 2780
2026-07-28 18:26:55 -04:00
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Tom Boucher
9138271b5f test(#2723): differential emitted-attribution check, dual-run beside the golden (#2737)
* test(#2723): differential emitted-attribution check, dual-run beside the golden

Phase 3 of #2719. The conservation law itself, running BESIDE
golden-install-parity.test.cjs -- both green, fixtures untouched.

Every emitted path whose hash moved between next HEAD and PR HEAD must be
attributable, through the Phase 2 table, to a path the PR actually changed.
Unattributable deltas fail with the paths NAMED. The only way through is a
committed acknowledgment, never a flag -- a contributor facing a red gate
sets a flag, which is what UPDATE_GOLDEN=1 is today.

The central decision is that the law is a PURE function (no fs, git,
installer, or clock), with I/O confined to a separate resolver. The naive
one-big-integration-test shape would need ~38 installer spawns per assertion,
so #2723's four failing-first criteria would not in practice have been
written -- which is exactly how a phase ships promised-but-not-built. Pure,
they are millisecond table tests, and the Stryker gate can actually bite.

Buckets are conserved: every moved path lands in exactly one of
attributed | unattributable | acked, property-tested at 400 runs. A path the
provenance table cannot resolve surfaces as an error, never a silent skip.

Asymmetries that are deliberate, each with a test:
- an ADDED emitted key is a ripple too, not just a modified one
- synthesized paths are exempt; code-derived ones are NOT (Phase 2 refused to
  mark them exempt precisely because exempt means permanently blind)
- shrinkage needs no ack; growth does. Gating shrinkage would punish exactly
  what the size ratchet wants
- a STALE ack is a hard failure -- an ack outliving its ripple pre-clears the
  next one on that path
- a failed `git diff` is an explicit error, never an empty changedPaths set;
  reading it as "nothing changed" would make everything unattributable and
  produce a failure storm that reads like a real finding
- prefix sources are SEGMENT-aware, so `agents/` does not attribute
  `agentsfoo/x.md`

Baseline is cached, not committed, keyed on the next sha. A stale key is
refused rather than used: absence fails loudly and gets fixed, whereas
staleness produces a confident wrong answer. An explicitly pointed-at
GSD_EMITTED_BASELINE that is stale is a hard stop; a stale cache falls
through to the in-job build. No baseline-unavailable path returns -- ADR-2719
section 6 names that trap, since in node:test a bare return is a PASS.

Both the conservation property and the staleness gate were mutation-verified
(injecting a swallowed key fails 9 tests; disabling the staleness comparison
fails 5).

Fixtures, generators, the merge driver and the ADR status are untouched --
those are Phase 4 (#2724).

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): compute stale acks once, after the size pass

Self-review defect found while the reviewers were running. `staleAcks` was
computed twice: once between the hash pass and the size pass, then again
after. Only the second value was returned, so the first was dead code -- and
the dead one was placed where it would have been WRONG.

An acknowledgment can be consumed by either a hash move or a size growth.
Computing staleness before the size pass reports a legitimate growth ack as
stale, which is a false failure that pushes a contributor to delete the very
ack that is doing its job.

Now computed once, after both passes, with a regression test. Verified by
mutation: restoring the early computation fails 2 tests.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): wire the attribution check to the real tree, not just synthetic input

An isolated reviewer caught that the first cut was INTERFACE-ONLY: nothing
read the ack file from disk, nothing shelled git, nothing built real
manifests. Every test was true of hand-built inputs and none of the repo, so
the acceptance criterion "both this check and golden-install-parity green on
the same tree" was trivially true rather than meaningfully true. That is the
promised-but-not-built failure this epic keeps finding in its predecessors,
recurring one phase later for the wiring itself. Taken, not argued.

Adds tests/helpers/emitted-runtime.cjs -- the only module that touches git,
disk, or the installer -- and an integration test that runs the same pure law
against reality:

- CURRENT side: 19 real installer spawns via runMinimalInstall +
  buildParityManifest, the same machinery the golden harness uses.
- BASELINE side: `git show origin/next:<fixture>`. That is next's RECORDED
  emitted state and it costs nothing. Deliberately NOT the working-tree
  fixtures, which are whatever this PR's author regenerated -- comparing
  against those would be vacuous. Phase 4 deletes the fixtures and swaps in
  resolveBaseline's cache path, already implemented and tested.
- changed paths from real `git diff --name-only origin/next...HEAD`, with the
  git subprocess bounded at 30s per CLAUDE.md's unbounded-subprocess rule.
- the real tests/emitted-drift-ack.json (absent is legal; present-but-empty
  or unparseable throws rather than being read as absent).

Verified it can actually fail: an uncommitted edit to a shipped workflow
moves emitted output but never appears in the committed diff, and the check
names all 18 affected emitted paths with the message format ADR-2719 §1
specifies. Restores clean.

Also from review:
- readAckFile now has a real test exercising the SUT across absent / valid /
  empty / unparseable / unreadable. The previous test asserted fs behaviour
  rather than SUT behaviour, because no SUT ack-reading path existed yet.
- formatReport's sampleLimit gains true limit-1/limit/limit+1 coverage at
  19/20/21. A test was previously NAMED "(limit+1)" while testing no numeric
  limit at all, which is worse than no coverage because it reads as covered.

Windows uses an explicit t.skip (install output is platform-specific there,
mirroring the golden harness) -- never a bare return, which node:test scores
as a PASS.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): cover the claude-local manifest family in the real-tree check

Isolated adversarial review, MAJOR. The real-tree wiring enumerated
Object.keys(RUNTIME_META) -- 18 entries -- while the emitted manifest set has
19 families. The 19th is claude-local: claude is the reference host and the
only runtime with a distinct LOCAL "legacy flat-commands" layout
(commands/gsd-*.md + agents/gsd-*.md at project scope), which
golden-install-parity.test.cjs guards with a hand-coded test outside its
RUNTIME_META loop (#2086).

The family was dropped from BOTH sides, so the test's own self-check
(current.length === baseline.length) passed vacuously at 18 === 18. A PR
changing Claude's local-scope output would have failed the golden while this
check reported ok -- and that disagreement is precisely what the dual-run
window is designed to surface as a provenance-table hole. A wiring omission
masquerading as one is the worst available failure here, because it would
have been read as evidence about Phase 2 rather than a bug in Phase 3.

Fixed by deriving MANIFEST_FAMILIES explicitly (18 global + claude-local at
local scope) instead of inferring the set from RUNTIME_META.

The self-check is also repaired: it now asserts both sides against the
INDEPENDENT EXPECTED_MANIFEST_COUNT from the Phase 2 table, and asserts
claude-local specifically. Comparing the two sides to each other can never
catch a family missing from both -- the assertion has to come from outside.
Verified by mutation: removing claude-local again fails the test.

Also from the same review:
- sourceSatisfiedBy returns the matched source string, so an empty-string
  source would return '' and the caller's `if (hit)` would silently discard a
  real match. Unreachable today (every rule source is a non-empty template)
  but a footgun for the next rule author; now `!== null`.
- the purity fixture used a single-element changedPaths array, so an in-place
  sort would have been invisible. Now three elements in unsorted order.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2723): resolve the base ref tolerantly instead of hard-requiring origin/next

The first matrix run failed on both linux lanes:

  differential attribution over the real tree
  cannot resolve origin/next (Command failed: git rev-parse origin/next)

Not a flake, and not an environment excuse -- a real defect in this diff. The
gsd-test runner shallow-clones and merges base+head, so no origin/* remote-
tracking refs exist in the container. My own fail-loud path fired correctly;
what was wrong was hard-depending on that ref existing. GitHub Actions has the
same shape by default, which is exactly why changeset-required.yml carries an
explicit `git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"`.

Now resolved through an ordered candidate list -- GSD_EMITTED_BASE (explicit
lane override), then origin/$GITHUB_BASE_REF and $GITHUB_BASE_REF, then
origin/next and next -- de-duplicated, each verified with
`rev-parse --verify <ref>^{commit}`.

When NO candidate resolves the test takes an explicit t.skip() naming every
ref it tried and stating that the gate did not run here. That is the
ADR-2719 section 6 distinction: t.skip is REPORTED as skipped, whereas a bare
return is scored as a PASS. Hard-failing was the other option and is wrong --
it would make the suite permanently red wherever a base ref cannot exist by
construction, which is a statement about the checkout, not a propagation
finding.

The candidate ordering is pinned by a unit test rather than left implicit,
since the ordering IS the fix.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 23:22:37 -04:00
Tom Boucher
1f6822ccba test(#2722): emitted-artifact provenance table with a totality guard (#2735)
* test(#2722): emitted-artifact provenance table with a totality guard

Adds the declarative emitted-path -> source-path table that ADR-2719 §2
specifies, plus the totality guard that keeps it honest. Phase 2 of #2719.

Every emitted path across all 19 committed golden-parity manifests (8,524
paths) must match exactly one rule. Zero matches, two matches, and a rule
matching nothing are all hard failures, so a new emitted family fails the
build loudly instead of passing through unattributed.

The measured surface is larger than #2722 estimated from claude.json alone
(26 top-level families across 19 runtimes, not 13), which is itself what the
totality guard exists to surface. It resolves to 19 rules.

Building the table caught three false attributions that were total but
resolved to repo files that do not exist -- Copilot's `<name>.agent.md`
rename, Kimi's code-literal `agents/gsd.{yaml,md}` root agent, and Copilot's
`hooks/gsd-session.json` registration. The "every attributed source exists"
test is therefore a first-class gate, not a nicety.

Notable correctness decisions:
- Emitted shapes are hard-coded; deriving them from the installer would make
  the guard tautological (it would follow any installer change silently).
  Only source paths read a first-party descriptor, and only where the
  descriptor is the sole declaration (hostBehaviors.nativePlugin.source).
- Emitted skills attribute to commands/gsd/*.md, NOT the repo skills/ dir --
  that directory is generated from commands/gsd by gen-plugin-skills.cjs, so
  attributing to it would be false attribution that still passes totality.
- Attribution is keyed on (rel, runtime): plugins/gsd-core.js has different
  sources for opencode and kilo.
- Rule order carries no semantics (property-tested), since exactly-one
  matching is enforced rather than first-match-wins.

Nothing here reads a git diff, builds a live manifest, or touches a fixture;
the differential check, drift-ack file and size ratchet are #2723.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* docs(#2722): record the delivered provenance table in the CONTEXT.md glossary

The `### Emitted Artifact Provenance` entry landed in #2721 describing the
table as future work. Phase 2 delivers it, so the glossary now records what
actually exists and the invariants #2723 must preserve:

- where the table lives, its rule count, and that it is total over all 8,524
  emitted paths across the 19 manifests
- dead-rule detection, so table rot is loud in both directions
- the corrected surface measurement (26 families, not the 13 estimated from
  claude.json alone)
- the two invariants #2723 inherits: shapes hard-coded (deriving them would
  make the guard tautological), and attribution keyed on (rel, runtime)
- the skills/ false-attribution trap, and that totality does NOT catch a
  wrong-source rule — the source-existence assertion is what does

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2722): close five review findings on the provenance table

Two orthogonal review passes plus an isolated adversarial reviewer returned
findings at minor..major. All fixed; no blockers were raised.

Standards axis (CONTEXT.md:456, RULESET.TESTS.guard-toplevel-readFileSync):
- module-level loadManifests() threw at require time before any test()
  registered, turning a missing fixture dir into an opaque crash instead of
  one named failure. Now a memoized lazy accessor.
- matchRules and assertTotality each carried their own copy of the matching
  loop; assertTotality now calls matchRules. That is the #2266 divergence
  class, and two copies could let the guard and the attributor disagree.
- named the corpus stride constant; dropped an inline require.

Spec axis:
- the CONTEXT.md glossary carried a "26 families" figure that is not
  reproducible from the code and that no test pinned -- a hand-maintained
  number in permanent canon, i.e. exactly the silent drift this epic exists
  to end. All volatile counts are now removed from the glossary, with the
  reason stated inline: the guard recomputes them every run, so they belong
  in a failure message, not in prose. No test was added to pin the count,
  because that would rebuild the brittle committed number we are deleting.

Isolated adversarial review:
- `.+` tail captures let a `..` segment reach a constructed source path that
  resolves outside the repo. Not live-exploitable (fixtures are committed and
  the only consumer is an existsSync probe) but Phase 3 feeds these strings
  into a diff-consuming check, so assertSafeRelPath now fails closed once, in
  matchRules, rather than per-rule.
- attributeEmittedPath's ambiguous-match branch was never exercised; only
  assertTotality's parallel path was. Now tested directly.
- sampleLimit's truncation branch had no limit-1/limit/limit+1 coverage.
- the fast-check property could not fail for the reason it was named for.

That last one took two attempts and is the one worth reading. The property
hand-rolled its shuffled side from the per-rule matchOne primitive, which is
order-independent by construction, so it held for reasons unrelated to the
shipped matchRules. Routing it through the real matchRules was still not
enough: on an unambiguous table, first-match-wins and collect-all return
identical results for every path (measured: 0 of 190 corpus paths differ).
Order can only matter where more than one rule matches, so the property now
also asserts that an intentionally ambiguous table reports BOTH hits as a set
under every permutation. Verified by mutation -- injecting a `break` into
matchRules makes it fail, and restoring makes it pass.

Enabling all of the above: matchRules and attributeEmittedPath now take an
injectable rules table, so tests can drive the real code path instead of
re-implementing it by hand.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:51:51 -04:00
Tom Boucher
80778e2674 fix(#1881): report an unreadable ROADMAP instead of reading it as absent (#2729)
* test(#1882): stage one file per commit in the base-ref ancestry fixture

CI failed on ubuntu-24 inside this test's setup loop, before any code under test
ran: at commit 32 of 60 the index referenced a blob whose object write had not
landed -- "invalid object ... for 'base-31.txt' / Error building trees".

The loop staged with `git add .`, which re-stages every file already in the tree.
Across 60 iterations that rehashes O(n squared) blobs -- roughly 1,800 stagings
and 60 full index rewrites to add 60 one-line files -- and that churn is what the
object store failed under. Each commit only ever adds a single new file, so
staging that one path is equivalent and removes the redundant work entirely.
Verified the loop still builds the intended history: 61 commits, git fsck clean.

The fixture already carries a note from an earlier fix in this epic recording
that it passed on ubuntu-22 and windows-24 and failed on ubuntu-24 for the same
commit. That was a different stage -- fetch versus diff -- but the same lane and
the same brittleness, so this is the second time this fixture's cost has surfaced
as a red build rather than as a test failure.

Not caused by this PR's change, which touches two configuration lists and cannot
reach a scratch git repository in tmpdir. Fixed here rather than deferred,
because the run surfaced it.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1881): prove an unreadable ROADMAP is indistinguishable from an absent one

Failing-first. Encodes the issue's runtime repro: an unreadable ROADMAP.md makes
getRoadmapPhaseInternal return the same null it returns for "phase not found",
and getMilestoneInfo return the same {v1.0, milestone} it returns for a project
with no roadmap at all -- so a permission or I/O fault reads as a brand-new
project.

Half these cases exist to hold the opposite line. getMilestoneInfo has no
existsSync guard, so platformReadSync's null-for-ENOENT is converted to a
synthetic Error carrying no errno, and that lands in the SAME catch as a real
EACCES. Reporting unconditionally there would flag every project without a
ROADMAP.md -- every brand-new project -- as corrupt. The absent case, the
errno-less error, a non-string errno, unparseable content and a genuinely missing
phase are all pinned silent.

One case guards a decision rather than behaviour: an unreadable STATE.md alone
must stay silent, because the inner catch that swallows it is deliberate and
documented under the #2245 audit as an optional enhancement falling back to
ROADMAP-only heuristics.

Two more pin the invariant ADR-1411 names explicitly -- neither function may
throw, because src/state.cts removed its own defensive try/catch on the strength
of that guarantee.

Assertions are on the frozen reason enum and the emission counter, never on
diagnostic prose. Faults are injected by overriding the platformReadSync seam and
restoring in t.after(), never chmod 0o000, which root bypasses.

Adds the ROADMAP_UNREADABLE reason to the shared vocabulary as scaffolding; no
call site emits it yet, which is what makes these tests red.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1881): report an unreadable ROADMAP instead of reading it as absent

getRoadmapPhaseInternal returned null for a read failure exactly as it does for
"phase not found", and getMilestoneInfo returned {v1.0, milestone} exactly as it
does for a project with no roadmap -- so a permission or I/O fault presented as a
brand-new project and workflows synthesised a blank phase or skipped requirement
extraction with no signal.

Both return values are preserved exactly, per ADR-1411's amendment: continuity is
correct, the silence was the defect. Each catch now reports through the shared
unusable-input seam that shipped with #1882 rather than a second copy of the same
mechanism.

The discriminator is the errno, and it is load-bearing in the silent direction.
getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT
is converted into a synthetic Error with no code that lands in the same catch as
a real EACCES. Reporting unconditionally there would flag every project without a
ROADMAP.md -- every brand-new project -- as corrupt. A genuine read fault always
carries an errno; absence never does. The parse is regex over text and cannot
throw, so nothing else reaches these catches.

Neither function gains a throw. ADR-1411 names this explicitly: src/state.cts
removed its defensive try/catch around getMilestoneInfo under the #2245 audit
because it never throws, and two tests pin that. The inner STATE.md catch stays
untouched and silent -- its fallback to ROADMAP-only heuristics is a deliberate,
documented optional-enhancement path, not a fault.

Where the fix belongs was the design question. platformReadSync does not leak: it
keeps absent and unusable as two channels, exactly as an abstraction should. Both
callers re-collapsed that distinction, so the fix is caller-side and the
projection seam -- with roughly ninety other dependents -- is untouched.

Closes #1881

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1881): admit the roadmap reason to the locked vocabulary

The seam documents adding a reason as three coordinated changes -- the enum
entry, the emitting call site, and the test that locks Object.keys(...).sort().
This PR made the first two and the lock caught the third, which is the whole
point of pinning the key set rather than asserting each value exists.

The roadmap suite no longer re-locks the full set. Two complete locks would mean
two files to update every time a later phase adds a reason, and #1883 and #1884
are both going to. The canonical lock stays in the seam's own suite; the roadmap
suite asserts only the value it introduces.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1881): resolve the roadmap path inside the try, not outside it

Naming the file in the diagnostic required the resolved path in the catch, and
the obvious way to get it was to hoist `path.join(planningDir(cwd), 'ROADMAP.md')`
above the try. planningDir throws a plain Error for an invalid GSD_WORKSTREAM or
GSD_PROJECT segment -- one containing a slash, backslash or `..` -- so hoisting it
let that throw escape uncaught.

That broke the exact invariant ADR-1411 names as this file's hazard: src/state.cts
removed its defensive try/catch around getMilestoneInfo under the #2245 audit
because that function never throws. Of its callers only archivePhaseDirectories
wraps it; cmdInitExecutePhase, cmdInitNewMilestone, cmdInitMilestoneOp,
cmdInitManager, cmdInitProgress, cmdProgressRender and cmdStats all call it bare,
so a workstream name with a slash in it crashed the CLI outright instead of
degrading.

The previous commit asserted "neither function gains a throw -- two tests pin
that". That was false. Both tests inject faults through platformReadSync only and
never through planningDir, so neither could have exercised the path that broke.
The guarantee was claimed, not demonstrated.

The path is now declared before the try and resolved inside it, so the catch can
still name the file when there is one, and a path that never resolved reports
nothing and returns the sentinel unchanged. The two test names are narrowed to
what they actually prove -- that a failing READ does not throw -- and a new case
injects the planningDir failure directly, which is what would have caught this.

getRoadmapPhaseInternal carried the same hazard, resolving the path outside its
try since before this branch. It is fixed the same way rather than left: ADR-227
is explicit that throwing breaks pipeline continuity, this read path already
degrades to null for every other failure, and a PR whose purpose is hardening
this invariant is the wrong place to leave the sibling crashing.

Behaviour otherwise unchanged and re-verified: healthy lookups, EACCES reporting
on both functions, absent-roadmap silence, and the errno discriminator all
unaffected.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1881): backfill changeset pr number to 2729

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 20:17:09 -04:00
Tom Boucher
a613caaeef enhance(#2721): regenerating merge driver, regen:derived, and a name for the emitted-artifact family (#2730)
* test(#2721): failing-first suite for the gsd-regen driver and CONTEXT.md parity

Tests precede the implementation per the TDD gate. The driver module does not
exist yet, so tests/git-merge-regen-driver.test.cjs fails at require time; the
contributor-standards parity assertions fail against next as it stands today,
where the standards doc names two CONTEXT.md headings that have never existed.

Refs #2721

* feat(#2721): add the gsd-regen merge driver and regen:derived

The golden parity manifests and the two size baselines are pure functions of
the source tree, so their only correct merge is "recompute" -- something git's
ours/theirs interface cannot express. 140 of 143 conflicted-file instances
across the open PR queue are these files.

The driver deliberately does NOT regenerate. Four probes established that at
merge-driver time neither the working tree nor the index reflects the merge:
both hold the ours side, a file added by theirs does not exist yet, and
MERGE_HEAD is unwritten. Git also invokes the driver once per conflicted path
(20 here). A regenerating driver would therefore read the ours-side tree and
emit a plausible-but-wrong hash manifest -- worse than a conflict, because a
conflict is visible. So it accepts %A, runs zero subprocesses, records the
resolved paths, and prints one notice pointing at npm run regen:derived.
Staleness stays caught where it already was, by golden-install-parity in CI.

Every failure path degrades toward today's behaviour (a normal conflict).
install-tree is deliberately excluded per ADR-2719 section 7.

Also folded in, per the no-defer rule: workflow-size.cjs claimed .md files have
no eol=lf in .gitattributes; git check-attr shows eol: lf, set by .gitattributes
line 2 since #1088.

Refs #2721

* docs(#2721): document regen:derived and the gsd-regen merge driver

Adds the how-to a contributor actually reaches for when the generated parity
manifests or size baselines conflict, in both places they would look: the
merge-conflict path in CONTRIBUTING.md and the full guide in TESTING-SUITES.md,
including what the driver deliberately does not do (it does not clear GitHub's
CONFLICTING label, and it does not regenerate mid-merge).

Also scopes the new contributor-standards parity assertion to the doc's own
CONTEXT.md section. Its first run flagged `## Decision`, `## Consequences` and
`## Standards followed`, which the doc attributes to an ADR body and a PR body
rather than to CONTEXT.md -- a doc-wide extractor would have demanded CONTEXT.md
grow headings that do not belong to it.

Refs #2721

* fix(#2721): stop passing %P to the merge driver — shell injection

The isolated adversarial review found, and I independently reproduced, local
arbitrary command execution.

Git does not invoke a merge driver with an argv array. It substitutes %O %A %B
%L %P textually into the configured string and runs the whole thing through a
shell, and $(...) executes inside POSIX double quotes -- so quoting the
placeholder does not neutralise it. %O/%A/%B are git-generated temp names and
%L is an integer, but %P is the file's own path, chosen freely by any
contributor. A branch renaming a covered fixture to
evil$(touch PWNED_SENTINEL).json executed that command on the machine of every
maintainer who merged it, and the merge still reported success.

Fix removes the input rather than filtering it: %P is no longer registered, so
the driver receives no attacker-controlled argument at all. The marker records
a count instead of path names. A metacharacter filter would have been a guess
about shell grammar; passing nothing is a property. Re-ran the identical
exploit against the fixed command: nothing executed, conflict still resolved.

Two regressions guard it -- a platform-independent assertion that the
registered command carries no %P, and a real merge driven by the actual
planInstall output with a $(...) filename.

Also from review: CLI dispatch had no coverage at all (CONTRIBUTING's
"CLI and command routing" matrix), which is why runInstall/runStatus now take
{repoRoot} -- hardcoding REPO_ROOT was what made them untestable. Renamed
planResolution to resolveAndRecord since the plan* prefix promised purity it
did not have. Reconciled the eleven-vs-twelve generator count across
CONTEXT.md, CONTRIBUTING.md and the changeset.

Refs #2721

* test(#2721): scope safe.directory for the check-attr helper

The 66f4d85a run failed 11 assertions, all in the .gitattributes scoping block,
with "fatal: detected dubious ownership in repository at '/work'". The test
container checks the repo out at a path its user does not own, so git refuses
check-attr outright. Everything else passed (27,185).

`check-attr` is a pure read of .gitattributes -- no hooks, no filters -- so the
exemption is scoped to that one invocation. It is deliberately NOT applied to
the driver's own production `git config` calls, which run in the user's own
clone and should keep the protection.

Refs #2721

* test(#2721): delete the stale assertion that the driver command carries %P

The plex2 run on bdfd0856 left exactly two failures, both this test: it still
asserted the pre-fix command string, i.e. the vulnerable behaviour. Deleted
rather than relaxed, per RULESET.TESTS.delete-bad-tests -- its useful half is
already covered, in both directions, by
registeredDriverCommandNeverPassesThePlaceholderForTheFilePath.

Refs #2721

* test(#2721): drive the end-to-end merges from the real planInstall output

The e2e helper hand-rolled its own driver registration, and still carried %P.
That meant the five real-git tests were not exercising the production command
string at all -- planInstall could drift and they would keep passing. They now
register exactly what a contributor gets from npm run setup:merge-driver.

Refs #2721

* chore(#2721): backfill changeset pr number to 2730
2026-07-27 19:55:37 -04:00
Tom Boucher
e48eb44003 fix(#1856): give the executor-worktree refusal a handoff instead of a dead end (#2727)
* test(#1856): failing-first contract for the orchestrator cwd-drift guard handoff

The #48 guard correctly refuses to execute waves from an agent worktree, but the
refusal is a dead end: that worktree can hold committed fixes AND uncommitted
work, and "re-run from the orchestrator's own worktree" silently means
abandoning them. The reporter was left choosing between continuing from a
blocked worktree and losing the work.

The guard is shell embedded in execute-phase.md, so these tests extract the
block by a stable marker and EXECUTE it against real git fixtures — the shipped
text is the runtime contract. Covers the stranded-commit and dirty-tree report,
the integration commands, both agent- namespaces, commit-count boundaries 0/1/2,
and the constraints the guard's own comment records: it must NOT fire on an
ordinary branch, on 'agentic-refactor', or on a legitimate feature worktree under
.claude/worktrees/, and must degrade cleanly with no resolvable base or a
detached HEAD.

RED expected: the marker does not exist, so extraction fails and every case errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): give the executor-worktree refusal a handoff instead of a dead end

The #48 cwd-drift guard correctly refuses to execute waves from an agent
worktree — its comment records why ("this is how a wrong-base merge nearly
shipped ~1000 files"), and that refusal is untouched here. The defect is that it
was a dead end.

At the moment it fires, the worktree can hold committed product work, uncommitted
product and planning changes, and the live gap-planning context. Telling the user
to "re-run from the orchestrator's own worktree" silently means abandoning all of
it, because the orchestrator worktree cannot see commits that live only on the
agent branch. The reporter was left choosing between continuing from a blocked
worktree and losing five commits plus uncommitted work.

The refusal now reports what is actually stranded — the commit count and log
against the resolved base, and the uncommitted files — followed by the concrete
integration sequence (commit here, switch to an orchestrator-safe checkout,
merge or cherry-pick, re-run) and a verify command. Nothing is claimed that is
not there: a clean worktree with no commits ahead prints the plain refusal with
no empty sections.

Every added command is diagnostic and `|| true`-guarded, so a failure degrades to
the original refusal rather than crashing before the message prints. Verified: an
unresolvable base still refuses cleanly.

Deliberately NOT done: auto-merging or auto-cherry-picking the agent branch. That
is precisely the operation #48 exists to stop the orchestrator performing from a
drifted cwd, at the moment it has least confidence about which tree is which.
Reporting beats acting here.

The guard block carries a `gsd:guard=orchestrator-cwd-drift` marker so the new
contract test can extract and EXECUTE the shipped shell against real git
fixtures rather than asserting on its characters.

Fixes #1856

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): report true counts, add the changeset, and document the seam

Review findings, all from the isolated adversarial pass:

- The dirty-file list was capped at 20 with no indication, so a worktree with 27
  uncommitted files reported 20 — under-informing the user about exactly what is
  stranded, which is the entire point of this change. Both lists now count BEFORE
  truncating and print "… and N more". Verified with 25 commits / 27 dirty files.
- The has-commits condition was written out twice and could drift on a future
  edit. Collapsed to a single _WT_HAS_COMMITS flag.
- The changeset fragment existed but was untracked, so it was in neither commit
  on this branch and the PR gate would have failed against real history.
- CONTEXT.md:122 documents this exact seam ("the orchestrator runs a cwd-drift
  guard at execute_waves entry…") and was not extended. Now records the handoff
  report, that the refusal condition and exit code are unchanged, and that every
  added command is diagnostic and || true-guarded.

Verified NOT a defect, correcting the review's premise: the guard block does break
when its line endings are CRLF, but .gitattributes:2 is `* text=auto eol=lf`,
which OVERRIDES core.autocrlf and forces LF on checkout on every platform
including Windows — so the shipped file is LF there too, and the installer copies
it through Node without translating endings. The reproduction (mine and the
reviewer's) required injecting CRLF by hand. It is also not fixable from inside
the script: a \r breaks the shell parse at the block's first line, before any
#1856 code runs. Neither introduced nor amplified by this change.

Also noted and left as-is by design: the review flagged that #1856's "offer an
explicit recovery option" could be read as requiring an interactive/automated
integration rather than printed instructions. Deliberate — see the commit that
added the block: auto-merging is the exact operation #48 exists to prevent the
orchestrator performing from a drifted cwd. Called out in the PR body for a
maintainer decision rather than silently chosen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#1856): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 18:32:00 -04:00
Tom Boucher
9a76ca6783 fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter

extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.

Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.

The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.

Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (3eb1cede2), making it
the only non-text file under src. file(1) reported it as data and text tools
silently skipped it, defeating the audit rule that says to search the authored
source; tsc passed because the bytes sat inside a comment, so no gate caught it.
It is live on next.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): pin unterminated-frontmatter detection and its negative space

Covers the discriminator on both sides. The positive rows are the issue's own
repro (LF and CRLF) plus the key-count boundary 0/1/2 around the ">= 1 parsed
key" threshold. The negative rows are the documents that reach the same branch
and must stay silent -- above all a Markdown thematic break at byte 0, which is
how this class of check has previously shipped a false positive on valid
Markdown.

Deduplication is tested on both halves of the composite key: a repeat of the
same (path, cause) is suppressed, a genuine second failure in a different file
is not, and a Windows and POSIX spelling of one path resolve to a single key.
The reset seam is asserted to actually clear -- #2674 is the precedent where a
reset that silently failed to clear made every later dedup assertion a vacuous
pass, and the cases only passed because each happened to pick an unused key, so
every case here uses a path unique to itself.

Assertions are on typed surfaces throughout -- the frozen reason enum and the
dedup-set size -- never on diagnostic prose. The one CLI-level case asserts a
differential between two runs (whether stderr is empty) rather than matching a
message, and is the wired user-reachable surface for this fix. Stream failure is
injected by overriding process.stderr.write and restoring it, never chmod 0o000,
which root bypasses.

Two properties guard the ~50 call sites of the changed function: the new
optional path argument is inert with respect to the parsed value, and LF/CRLF
spellings of a document still parse identically.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): raise the truncation threshold and repair the dedup key

Isolated adversarial review found the one-key discriminator false-positives on
ordinary Markdown: a thematic break above a single labelled line -- `Note:`,
`Author:`, `TODO:`, `See:` -- parses as exactly one key and was reported as
corruption, which is the precise failure the design claimed to prevent and the
changeset promised was fixed. The threshold is now two keys. A file truncated
after exactly one key becomes a false negative; that is the same
precision-over-recall direction already taken at zero keys, and every GSD
artefact this guards carries two or more frontmatter keys.

Three dedup-key defects, each of which could silently swallow a real diagnostic:

- Backslash normalization is removed. A backslash is a legal filename character
  on Linux and macOS, so folding it to a forward slash made two genuinely
  different files share one key. Two spellings of one Windows path may now
  report twice; two distinct files can never silence each other. Lost signal is
  the worse failure.
- The key namespaces are tagged so a file literally named like the unnamed
  digest fallback can no longer collide with a path-less caller whose content
  hashes to that digest -- computable for any predictable content, no brute
  force needed.
- The source identity is computed once rather than hashed twice per emission.

Corrects the previous commit. The two NUL bytes in src/config-loader.cts were
NOT in a JSDoc comment as that message claimed; they were deliberate separators
in the live dedup key, and stripping them degraded it to bare concatenation.
They are restored as escape sequences -- byte-identical runtime string, and the
file is text again so grep can see it. The diagnostic script that misled me
indexed a character-offset string with a byte offset.

Also threads sourcePath through the STATE.md and PLAN.md readers so the two
artefacts epic #1879 is actually about name their file rather than reporting
under a content digest.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): correct fixtures and assertions left behind by the review fixes

The previous commit changed two behaviours deliberately and the suite still
encoded the old ones, so gsd-test came back red with six failures across both
lanes -- all of them mine.

Fixtures carrying a single frontmatter key no longer clear the two-key
truncation threshold, so the CLI differential and the two path-less dedup cases
were asserting a diagnostic that is now correctly withheld. They now carry two
keys, which is what a real interrupted write of a GSD artefact looks like.

The Windows/POSIX case asserted that two spellings of one path collapse to a
single key -- the exact folding that was removed because it also collapsed
genuinely distinct POSIX files whose names contain a backslash. Inverted to
assert they now report separately, with the reasoning recorded inline so the
trade is not silently reversed later: mild duplicate noise on one Windows path
is acceptable, a swallowed diagnostic is not.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): name the file at every read site, and report each file once

The diagnostic reached only the four frontmatter CLI verbs, so ~47 of 53 call
sites reported a truncated file under an anonymous content digest instead of
naming it. Since naming the file is the whole point -- it is what an operator
can act on -- that was a gap in the deliverable, not a scoping choice. 43 of 53
sites now pass the resolved path.

Closing it surfaced a defect the original design missed. A single truncated
STATE.md is parsed twice in a normal run: once by the read wrapper, which holds
the path, and again by a pure core downstream, which is handed only the string
and cannot know it. Those two parses keyed separately, so one file produced two
diagnostics -- and wiring more sites made the collision more likely, not less.
Every emission now registers both identities the input could be known by and
checks both before writing, so whichever caller arrives first speaks and the
other is suppressed. Distinct files with distinct content still report
separately, which is the property ADR-1411 actually requires; two files whose
truncated content is byte-identical collapse to one report, which stays the
documented limit.

Ten call sites deliberately keep no path. Two are frontmatter's own round-trip
checks during set and merge, where passing a path would report on every write.
The other eight are the state-transition pure cores, which ADR-1769 defines as
(content, intent, deps) -> newContent with injected I/O; threading a path
through them would contradict that recorded decision, so it is surfaced rather
than taken unilaterally. With the widened key they no longer double-report, and
in the normal flow the named parse runs first, so the file is still named.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): inject the STATE.md path into the transition cores

The six state-transition cores parsed STATE.md frontmatter without knowing
which file it came from, so a truncated STATE.md reached the operator as an
anonymous content digest on exactly the artefact epic #1879 is named for.

ADR-1769 section 3 shapes these as (content, intent, deps) -> newContent with
injected deps, and deps is the seam for precisely this: something the core
cannot derive without doing I/O. It already carries roadmapProvider and a
phase-inventory provider on that basis, each documented as injected rather than
imported so the core stays pure and testable without disk access. A resolved
path is data, not I/O, so an optional sourcePath member extends the established
pattern rather than contradicting it, and every existing stub keeps compiling
because the member is optional.

updateCore and reconcileCurrentPosition take no deps and are left alone. With
the widened dedup key they cannot double-report, and in the normal flow the read
wrapper has already named the file by the time they run.

Also regenerates gsd-core/bin/lib/state-transition.cjs. That artifact is tracked
rather than gitignored, unlike most of its siblings, so leaving it stale would
have shipped a runtime without this change to anyone reading the repo without
building. tsc had skipped the re-emit because its incremental build info still
recorded an emit that had since been reverted, so the stale output survived a
clean build; clearing tsconfig.build.tsbuildinfo forced it. The
compiled-artifact-sync gate is what surfaced the drift and now reports all nine
tracked artifacts matching their source.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): stop the widened dedup key from hiding a second file

The previous commit widened the dedup guard so one file parsed twice -- once by
a read wrapper holding the path, once by a pure core holding only the string --
reported once instead of twice. It did that by checking BOTH keys before
emitting, which silently traded one defect for a worse one: two DIFFERENT files
whose truncated content happened to be byte-identical now collided on the shared
content digest, and the second file's diagnostic was swallowed. That is the
over-coarse keying ADR-1411 explicitly forbids, reintroduced while fixing
something else.

The guard now checks only the key matching what the caller actually knows -- a
named read checks its path key, a path-less read checks its digest key -- while
still recording every key the input could later be identified by. The redundant
path-less re-parse of an already-named file stays silent, and two distinct files
always both report.

Verified across all six orderings: same file named-then-anonymous reports once;
two different files with identical content report twice; two different files
with different content report twice; the same path twice reports once; two
path-less parses of identical content report once; two path-less parses of
different content report twice.

The suite caught this -- twenty failures, all in the unusable-input tests that
reuse one truncated fixture across different paths. The local probe written
alongside the broken change did not, because it compared two files with
different content and could therefore only confirm the expected behaviour.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): count diagnostics emitted, not identities interned

The suite measured the size of the dedup set as a stand-in for "how many
diagnostics were emitted". That held only while one emission recorded exactly
one key. Once an emission began recording every identity the input could later
be matched by -- a path key and a content key for the same file -- the set grew
by two per write and twenty assertions read 2 where they expected 1.

The production behaviour was correct throughout; the proxy was not. Set size
counts identities, which is an implementation detail of the guard. The
behavioural claim these tests exist to make is how many diagnostics an operator
actually saw, so the module now exposes that directly as an emission counter and
the suite asserts on it. The set-size accessor stays for assertions genuinely
about key shape.

The local probe written alongside the change did not catch this because it
counted process.stderr.write calls -- the right thing -- while the suite counted
set growth. Verification now asserts both and requires them to agree, so a
future divergence between the counter and real writes fails immediately rather
than being discovered a bench run later.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): retire two assertions that outlived the behaviour they described

Both tests encoded assumptions the dedup fix invalidated, and both were caught
by the suite rather than by the probe written alongside the change.

The forged-path case asserted that a file named like the anonymous digest
fallback must not suppress a later path-less report. That premise is gone: an
emission now records every identity its input could be matched by, so ANY named
report of some content silences the anonymous re-parse of that same content --
which is the same-file guard working as intended, and has nothing to do with the
crafted name. The property still worth defending is that a crafted filename can
never silence a real file reported under its own path, so that is what the test
now asserts, with the deliberate suppression documented beside it.

The reset-seam case ended by reading the size of the dedup set and expecting 1.
Set size counts interned identities, not diagnostics written, and one emission
now interns two. It asserts the emission counter for the event and keeps a
weaker set-size check for the interning.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): close the review findings on the discriminator, dry-run and counter

Three orthogonal review passes ran against the final diff. Their findings:

A labelled preamble under a leading rule was still misreported. Raising the key
threshold to two only moved the boundary, because two colon-labelled lines are
as common in ordinary prose as one -- a document opening with a rule over an
Author and a Reviewed-by line, then prose, was called corrupt. Key count alone
cannot separate the two. What does is what follows: a write interrupted part way
through a frontmatter block ends mid-block, so every line of the region is still
frontmatter-shaped, whereas a document merely opening with a rule goes on to
prose. Both conditions are now required, and each closes a false-positive class
the other leaves open. Nested list values and indented continuations stay
frontmatter-shaped, so legitimate truncations are unaffected.

`state rebuild --dry-run` reported a truncated STATE.md anonymously. The write
path is named only because readModifyWriteStateMd parses with the path first;
the dry-run branch reads the file directly and never did. Dry-run is the
read-only mode an operator reaches for first when they suspect corruption, so it
is the one that most needed to name the file. reconcileCurrentPosition takes the
path as an optional argument now and rebuildCore passes it down. That function
was previously left alone on the grounds that a read wrapper always names the
file first -- this is the flow that disproves it.

The emission counter counted write attempts rather than writes, so on a broken
stderr it claimed a diagnostic had reached the operator when nothing had. It is
incremented only after a write that completed, and the broken-stderr test now
asserts the count as well as the return value.

Two documentation defects. The module described a guarantee it does not keep:
one file yields one diagnostic only when the named read comes first. The reverse
ordering emits twice, and that is deliberate -- a path-less caller cannot
identify its file, so suppressing the later named report would also suppress a
genuine second failure in a different file whenever two files share identical
truncated bytes, which ADR-1411 ranks the worse failure. The comment now states
the asymmetric guarantee and a test pins it. Separately, the CONTEXT.md glossary
entry still described backslash normalization that a later commit removed, and
asserted the opposite of what the tests pin; no lint checks prose against code,
so nothing caught it.

Also converts three body-level try/finally blocks to t.after(), per
CONTRIBUTING.md's rule that try/finally belongs only in helpers with no test
context -- the file's own emissionsDuring helper already did this correctly.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#1882): tell the operator what the truncated-frontmatter warning means

A user who has just seen the new warning is acting, not studying, so this lands
in the How-To quadrant beside the other "if you see X" branches in
debug-a-failed-execution, not in reference or explanation. It gives them what
the warning means for this run, three steps to restore the file, and the fact
that the warning changes no return value or exit code.

It also states the case that matters more than the warning itself: silence does
not prove the file is intact. GSD says nothing when the partial block carries
fewer than two fields or reads as prose, because a Markdown document opening
with a horizontal rule is indistinguishable from one of those. A reader chasing
missing metadata needs to know not to treat quiet as clean. Why that threshold
exists is explanation and deliberately stays out of a how-to.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1882): backfill changeset pr number to 2712

* test(#1882): constrain each branch of the frontmatter-shape check

CI's mutation gate came in at 61.56 against a threshold of 62, and the surviving
mutants were concentrated in isFrontmatterShaped -- the function added last, in
response to review, and the only one never given tests of its own. It was
exercised solely through extractFrontmatter, which covers the composite decision
but leaves each branch of the predicate unconstrained: drop the blank-line
filter, or any one of the three shape alternatives, and every existing assertion
still passed.

Four cases now pin the halves independently. A blank line inside an interrupted
block must not disqualify it, which constrains the filter and its comparison. An
unindented list item and an indented folded-scalar continuation each exercise one
shape alternative that no other case reaches on its own -- the folded line is
neither a key nor a list item, so it is the only input that distinguishes the
indented branch. And two keys followed by prose must stay silent, which is the
negative half: it fails if the predicate is ever mutated to accept everything,
and it is the case that proves key count alone was never sufficient.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): register the unusable-input suite with the frontmatter mutation shard

The mutation gate reported an identical 61.56 across two runs whose only
difference was four added tests. That is the tell: the tests were never
executed. The frontmatter shard runs a fixed file list in stryker.config.mjs and
scripts/mutation-matrix.cjs, and tests/unusable-input.test.cjs was in neither, so
the entire suite covering the new unterminated-fence branch was invisible to the
gate while passing perfectly well in the normal run.

So the score was not measuring weak tests, it was measuring absent ones: #1882
added mutants to frontmatter.cjs and no test in the shard covered them. Both
lists gain the file; the config already notes they must stay in sync.

This is a registration ripple a new test file carries when it covers a
mutation-tracked module, alongside the .gitignore, eslint, inventory, glossary
and size-baseline ripples a new module carries. Nothing warned about it, which
is why two runs were spent before the identical score gave it away.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:50:12 -04:00
Tom Boucher
09477f925e fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading

The Workflow backend emitted every agent() call with no model, so
model_overrides / model_policy / model_profile were silently inert on that path
while the inline path honored them (ADR-1411). Neither existing suite contained
the string 'model' at all.

The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor')
rather than hardcoding either, so it asserts backend parity rather than a fixed
string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when
nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids
reaching the code generator, the #2285 composed seam, CLI config-defaulting, and
a fast-check round-trip property.

RED expected: no model key is emitted anywhere, and --executor-model does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2686): thread the resolved executor model into the Workflow backend

The Workflow backend emitted every agent() call with no model at all, so
model_overrides / model_policy / model_profile_overrides / model_profile were
silently inert on that path while the inline path honored all of them. The model
was not dropped at the last step — it was absent from the whole seam:
agentOptions() took no model, EmitInput had no field to carry one, and
ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could
not forward one. The generated script asserted the parity it broke.

VERIFY-FIRST, which #2686 flags as the question that decides the fix: the
Workflow tool's agent() DOES accept a per-call model. Its documented signature is
  agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? })
so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled
out. ADR-1143:24's option enumeration omitting `model` is an incomplete
enumeration, not a decision to exclude it.

- agentOptions(p, executorModel) emits `model` only when it is a non-empty string
  that is not "inherit" (#2517: an empty model 404s on runtimes without native
  tier aliases). A non-string is a malformed config: omit, never throw.
- executorModel threaded through EmitInput and ResolveWaveDispatchInput.
- The CLI resolves gsd-executor from project config by DEFAULT rather than
  requiring a flag, reading the same source the inline path reads. An
  orchestrator that never learns about a new flag would otherwise silently keep
  the old bug. --executor-model exists only to pin/override.
- ADR-1411 provenance: the generated header now states which model was applied,
  or that none resolved and why. A fallback must be a visible value.

Compatibility: when nothing resolves, the emitted options object is byte-identical
to before, so every existing caller and assertion is unaffected.

Behavior change (Hyrum's Law): opted-in users move from session inheritance to the
catalog-resolved executor model. Adding a `model` key also changes agent() opts,
which invalidates the cached prefix of any in-flight resumeFromRunId run — a
one-time re-execution. Both disclosed in the changeset.

Fixes #2686

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2686): reject script-breaking model ids and share the emit predicate

The isolated adversarial review found a BLOCKER in my own provenance comment,
proven by execution (the emitted script exited 42 from an injected statement).

U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment
in EVERY engine — the ES2019 change legalized them inside string LITERALS only.
So quoteString (JSON.stringify) is sufficient for the `model: "..."` object
literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in
a model id closed the comment and made the rest of the line live top-level code.
The value is reachable from `.planning/config.json` (model_overrides /
model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any
non-claude runtime — attacker-influenceable in a cloned repo.

`emitWorkflowScript` now rejects a string executorModel carrying any character in
UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to
phaseDir/runId, which is proof the codebase knew this hazard. Rejection is
ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an
emit failure to the inline backend WITH that reason, so the degradation is
visible. A non-string stays on the existing defensive path (omit, never throw) —
that is malformed config, not an injection attempt.

Also from the reviews:

- The predicate deciding "is this model emittable" was duplicated between the
  emission and the comment asserting it. Extracted to emittableModel() so a
  generated comment can never claim something the generator did not do — the
  exact failure class #2686 was filed for.
- That predicate now trims and lower-cases before comparing, closing a real
  #2517-class gap: " " and "INHERIT" were previously emitted verbatim.
- The adversarial test was pass-always against this very vulnerability — it
  asserted only that JSON.stringify appeared. Replaced with the real contract
  (rejection) plus an execution-level check that no LineTerminator survives into
  the comment. A raw U+2028 had also been committed into that test's fixture
  array where a tab was intended; both are now explicit \u escapes.
- optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated
  model contained — silently not testing what it claimed. Now brace- and
  string-aware.

Stale-test corrections in tests/fix-2285-*: three assertions froze the exact
options literal `{ agentType: "gsd-executor" }`. The object legitimately gained
an optional additive `model` key, so they now assert the invariant they exist to
protect (agentType present, isolation absent) rather than a frozen literal. The
CLI-vs-pure equality test pins --executor-model on both sides; otherwise it
compared a config-resolved CLI run against a pure call given no model.

CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the
new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* test(#2686): fix the options extractor and model the rejection path

Two defects in my own test helper, caught by the full matrix:

- optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The
  emitted shape is agent("brief", { ... }), so that never matched and the helper
  returned an empty array, making every assertion over it vacuously true. It now
  anchors on agent( and takes the first balanced, string-aware {...} after it.

- The fast-check property predated the security fix and asserted ok:true for any
  generated string. Strings carrying an unscriptable character are now rejected,
  so the property models the real three-way contract: unscriptable -> ok:false;
  trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the
  trimmed value.

Verified locally against the built module: omit values clean, both plans carry
the model on the parity path, property passes 500 runs at seed 42. Test file
re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are
explicit \u escapes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* test(#2686): scope no-control-regex on the mirrored unscriptable-char class

The class is the point of the assertion — those bytes are exactly what must be
rejected — so the rule is disabled at that line rather than the class weakened.
UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence
the mirror.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#2686): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 16:24:21 -04:00
Tom Boucher
1008aabd31 fix(#2615): document the effortSurface axis in the host-integration matrix (#2698)
* fix(#2615): document the effortSurface axis in the host-integration matrix

#2481 added `effortSurface` as the ninth negotiated `hostIntegration` axis and
wrote documentation-sourced values into 18 descriptors, but never touched
`docs/reference/host-integration-capability-matrix.md`. The matrix that ADR-1239
designates the cited source of truth had zero occurrences of the axis: no entry in
the axes legend, and no row in any of the per-runtime tables. `src/host-integration.cts`
states "every value is documented or explicitly 'undocumented'" — for this axis
that was false for every runtime.

Adds the legend entry (the `argv` / `none` / `undocumented` vocabulary, plus why
there is deliberately no config-file member) and an `effortSurface` row to all 19
per-runtime tables. Every citation is carried over from #2481's own commit message,
where the values were sourced:

- claude   argv -- `claude --help` documents `--effort <level>`
- opencode argv -- `opencode run --help` documents `--variant`
- codex    argv -- `model_reasoning_effort` is a config.toml key, not a dedicated
                   flag, so the generic `-c key=value` override is the only argv
                   route (still argv)
- 15 hosts undocumented -- their docs state no reasoning setting

kimi-code is the nineteenth section (added by #2603 after #2481) and is the one
runtime with no declared value. Its row and a Documentation-gaps entry record why
rather than inventing one: Kimi Code documents `/effort` (alias `/thinking`), but
only as an INTERACTIVE slash command — `-m, --model` is the only model-adjacent
argv. Neither vocabulary member is accurate (`none` would deny a mechanism the host
has, `argv` would claim one it does not expose), so closing that gap needs a
vocabulary decision, which is a negotiation change and not a documentation one. The
absent value already degrades closed exactly as the sentinel does.

The regression test derives its runtime list from the registry rather than
hardcoding it, so a runtime added later fails until its matrix row exists — the
ratchet whose absence let #2481 add an axis with nothing catching the missing docs.

Closes #2615

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* docs(#2615): honest citations for the undocumented rows; fix the four stale 8-axis lists

Two findings from the orthogonal review of the first commit.

1. The 15 `undocumented` rows shared byte-identical text — "searched the runtime's
   official docs (see Sources consulted above)" — which is weaker than this file's
   own convention ("no authoritative doc — searched: <url>") and, worse, implies a
   per-host targeted search that did not happen: each section's Sources-consulted
   list was gathered for OTHER axes and contains no CLI-reference or
   reasoning-effort source. The rows now say plainly what the finding is — an
   ABSENCE established by #2481's cross-host survey — and cite that survey rather
   than implying a URL was checked per host.

2. Four normative docs still described "the eight negotiated axes" and omitted
   effortSurface entirely. The worst of them is
   docs/how-to/add-or-update-a-host-integration.md — the maintainer's own guide for
   onboarding a host, whose Step 2 axis table would have a maintainer reproduce
   exactly the gap #2615 exists to close. Also fixed:
   docs/reference/host-integration-interface.md (which calls itself the normative
   reference and had no effortSurface row at all),
   docs/how-to/author-a-host-plugin.md, docs/registries/README.md ("**exactly** the
   eight … axes keys"), and CONTEXT.md's matching EoS-registry sentence.

Deliberately NOT changed, because they are historical records rather than current
contract: docs/whats-new-1.7.0.md and docs/FEATURES.md's 1.7.0 entry (effortSurface
shipped in 1.8.0 via #2481 — rewriting them would falsify the release history),
ADR-1239's pre-amendment body (already superseded by its own
"Amendment (2026-07-21): effortSurface axis (#2481)"), and ADR-1016's "original
eight axes", which refers to a different axis set entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2615): backfill changeset PR number (#2698)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:02:20 -04:00
Tom Boucher
a5633bb32f enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures

* feat(#2671): brand raw vs calibrated token types

* refactor(#2671): hoist type-compile into a before() hook

Two review responses:

- The fixture compile ran in the describe() body, so it executed at
  collection time even when the block was filtered out, and a failed
  precondition collapsed eight independent assertions into one opaque
  describe-level failure. A before() hook is this repo's documented
  idiom and preserves per-test granularity.

- parseTokensFlag now records WHY it returns an unbranded number: it
  validates the magnitude of --tokens, but the basis is decided by
  --calibrated, so branding here would be wrong for half its callers.
  The assertion belongs to cmdEstimateCheck, its only caller.

* test(#2671): pin each brand diagnostic to its OFFENDING marker

Adversarial review demonstrated that asserting only exactly-one-diagnostic-
at-code-N is not airtight. Repairing a fixture's brand violation while
injecting an unrelated error of the same code (a string passed as the
budget argument) still yielded exactly one TS2345, so the fixture would
have reported green while no longer testing its regression at all.

Each bad-* fixture now routes its violating value through a const named
OFFENDING, and the test asserts the diagnostic's start offset falls inside
that node — located through the AST, so it survives reformatting and never
pattern-matches source text. Replaying the proof-of-concept against the new
assertion rejects it: the diagnostic lands on the budget literal, not the
marker.

Also corrects a doc comment that claimed the program type-checks all of
src/; it covers phase-estimation.cts and its transitive dependencies.

* chore(#2671): backfill changeset PR number (#2676)
2026-07-26 21:42:50 -04:00
Tom Boucher
c3958018dd docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern

ADR-1411 reasons only about a resolution miss. It is silent on input that
is present but not usable, which is how five engine read paths (#1879) could
fold an unusable input into the value meaning 'genuinely absent' without
contradicting an Accepted ADR.

Read together, ADR-1411 and ADR-227 converge and do not license throwing as
the cluster's answer: ADR-227 requires malformed input to be coerced rather
than propagated and carves out only genuinely-fatal fields, while ADR-1411
already permits a fallback provided it is 'a visible value, not a silent
substitution'. The defect in these five sites is therefore not that they fall
back but that they fall back invisibly.

Records the pattern that follows: every current return value is preserved, and
the cause is made visible in-band where the result already carries a
provenance envelope, or out-of-band via a deduplicated stderr diagnostic where
it returns a bare value it cannot extend. Throwing stays confined to ADR-227's
genuinely-fatal carve-out, decided per call. Also names the per-applier caller
audit and the lint-resolution-provenance registry gap.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2674): prove the warning-state reset misses the unknown-key dedup set

The two existing cases in this suite only pass because each picks a key
name no other case reuses, so neither can observe whether the reset the
beforeEach calls actually runs.

Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after
_resetRuntimeWarningCacheForTests(). It is not - the helper clears only
_warnedConfigKeys despite documenting itself as resetting per-process
warning state.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2674): reset the unknown-key dedup set with the runtime warning cache

_resetRuntimeWarningCacheForTests documents itself as resetting per-process
warning state but cleared only _warnedConfigKeys, leaving
_warnedUnknownConfigKeys populated across cases. The suite that exists to
test that set - 'loadConfig - unknown-key warning dedup' - calls the helper
in beforeEach expecting exactly this, so the reset was a silent no-op for
it; both cases passed only because each picked a key name the other never
reused. Any later case reusing a key would have had its warning suppressed
by leaked state.

Found while amending ADR-1411, which names this dedup guard as the pattern
five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the
fix would have propagated the footgun to each of them. Folded in here per
CLAUDE.md's no-defer rule rather than filed.

RED verified on 3c4895841 (test only, no fix): linux-node22 reported
'FAIL tests/config-loader.test.cjs - the documented per-process
warning-state reset must clear the unknown-key dedup set too'.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): document src/ in the changeset-lint trigger list

CONTRIBUTING.md presented the Changeset Required trigger list as bin/,
gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which
scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the
TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is
the most-edited user-facing path in the repo and the omission sends any
contributor who touches it into a CI failure the doc says cannot happen.

Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so
running it bare locally reports success without evaluating the branch. This
PR hit exactly that: a local run said ok_fragment_present and CI failed
fail_missing_fragment on the same diff.

Found while opening this PR; folded in per the no-defer rule.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): restore the round-2 review corrections to the amendment

These edits were made in response to the second isolated review pass but
never staged: later commits used targeted `git add <file>` for the test and
the source fix, so the two markdown files stayed dirty and shipped nothing.
The branch carried the round-1 text, including the ADR-227 misquote the
reviewer raised as a blocker.

Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in
was never implemented, so citing it as the precedent was wrong), the dedup
key, #1882 folded into the out-of-band mechanism instead of a fourth
mechanism-less category, the narrowed caller-audit rationale, and the
test-methodology clause.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 20:25:39 -04:00
Tom Boucher
bd570618d4 feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop

* fix(#2632): calibrate against the raw projection so the loop converges

* test(#2632): add closed-loop convergence guard and codify the feedback-loop rule

* fix(#2632): pair calibration samples per plan; atomic write; amend adr

* chore(#2632): backfill changeset pr to 2672

* fix(#2632): retry renameSync on transient windows errnos and clean up the temp
2026-07-26 16:28:56 -04:00
Tom Boucher
46ba02acde feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs

* fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions

* fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests

* chore(#2630): backfill changeset pr to 2661
2026-07-26 01:42:47 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
4a66d62d10 feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver

Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.

worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.

resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.

Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: rebuild tracked state-transition.cjs to match #2400 source

The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit 2bcfaa2e2) added the progress.total_plans frontmatter sync to source but the tracked bin/lib/state-transition.cjs was never rebuilt, so the fix was not shipping to consumers of the compiled artifact. The mandatory build:lib step for Phase 2 surfaced the drift; recompiling makes the already-merged, already-changelogged #2400 fix effective. Artifact-only resync (no source/test change); drift class tracked by #2591.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 21:23:06 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
BeeHiggs
bf9fe4630d feat(#2249): bracket phase-id core grammar — parse/render/toDir round-trip pair (epic #612 PR-1) (#2258)
* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards

PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md).
Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single
canonical owner — as a pure, additive extension. The 17 locked exports and
PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical,
so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs)
stays green.

New pure round-trippable model (ADR Decision 4):
- PhaseId { project, milestone, phase, subphase?, plan? }.
- parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token
  `GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens
  (`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new
  parser — normalizePhaseName and every legacy reader keep accepting those
  tokens unchanged (conservative default; no existing path gains a throw).
- renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug`
  with a slug guard that sanitizes path-traversal input.
- getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone
  from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning
  the `vN.0` form (parity with READING-A). The optional parameter keeps the helper
  pure (no config read) and byte-compatible — every existing single-arg caller
  resolves to the unchanged READING-A body (ADR Decision 6).
- extractPhaseToken(dirName, convention?): bracket dir branch GATED on
  convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is
  string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal
  family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no
  string-only discriminator is complete — an ungated auto-detect silently
  reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit
  convention signal keeps every existing convention-less call site byte-identical
  (pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs).
- comparator: no new code — comparePhaseNum already orders the dot-decimal
  `PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a
  PR-2 resolution concern (bracketQualifiedKey), not core grammar.
- SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999}
  non-milestone guard; the bracket-prefix reading is gated the same way (an
  ungated read called `P0.0-foundation` a sentinel), legacy leading-int form
  unchanged.
- BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately
  more permissive than parsePhaseId — a read-tolerance source for PR-2, not the
  emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt
  owner so PR-2 builds every bracket read regex from the canonical source and
  check:phase-id-drift stays green stack-wide.

The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*`
grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the
config permits parses. parsePhaseId has no live callers in PR-1, so this grammar
choice is forward-facing for PR-2 with zero PR-1 behavior impact.

Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips,
full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases),
extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens,
sentinel + slug guards, bare-token rejection, exported-source behavioral
assertions, and two generative fast-check properties: render∘parse identity over
well-formed displays, and the toDir/disk↔display bijection. Plus a #2043
numeric-tail characterization (single- AND multi-digit rows) in
tests/phase-id.test.cjs pinning the convention-less reading byte-identical.

The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish)
and rebuilt by CI, so it is intentionally not committed.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3)

PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4).

B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers,
over-padded numbers, multi-space separators, stray whitespace), so
render(parse(x)) === x did not hold for every well-formed x as ADR-612
Decision 4 requires. Both branches now enforce canonicality by construction:
parse permissively, rebuild the canonical string via the same emit path
(renderPhaseId for display, a hand-rebuilt token for dir/token), and throw
"parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's
entry is removed — the match anchors now reject leading/trailing whitespace
outright, folding into the existing "not a bracket phase id" rejection.

M1 (major): toDir only ever guarded the slug; project/milestone/phase/
subphase were interpolated unsanitized, so a hand-built PhaseId (a
structural, not nominal, type) could smuggle a path-traversal segment onto
disk. Every field is now validated against the exact shape parsePhaseId
itself would produce before use.

M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling
trailing hyphen in the emitted dir name. toDir now throws in that case.

M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable
from the dir-branch's plan tail, so it silently broke the disk<->identity
bijection on read-back. toDir now rejects all-digit slugs.

Nits: toDir now rejects a non-string slug instead of coercing it to the
literal token 'undefined'/'null'; sentinel boundary tests added for
milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999},
not an inclusive range — these were already correct, now locked by test).

Test-first: every new assertion (concrete examples + fast-check mutation
property for B1; concrete cases for M1-M3 and the nits) was written and
confirmed red before the implementation changes, per repo TDD convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2249): reformat changeset body to house convention (review Mi2)

The fragment added in ab26190a was a plain paragraph — no bold headline,
no trailing issue reference. Reformat to the repo's
`**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g.
.changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md).

Uses (#2249), the issue every commit on this branch references, not the
PR number already carried in frontmatter (`pr: 2258`) — the changelog
serializer appends `(#{pr})` unconditionally, so a body also ending in
`(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered
bullet directly via parseFragment + serializeChangelog: it now reads
`... (#2249) (#2258)`, matching the dominant convention across the other
fragments (frontmatter pr = merged PR, body reference = originating issue).

Also moved the docs-exempt marker back before the paragraph -> after it
(matching the file's original order): the marker sits on its own line and
is stripped before the body is used, but placing it first left a leading
blank line in front of the bold headline once reformatted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
two property-generator coverage gaps the reviewer flagged; no source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged).

Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no
property fed a 3+-digit milestone/phase/subphase/plan through parse/render/
toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen
numArb to 1–999 so the round-trip and disk↔display bijection properties both
span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no
leading zero, so canonicality still holds). Add a concrete regression pinning
the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and
toDirs to 'GSD.100-05-feature' without truncation.

Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered
milestone/phase pad + whitespace mutations but never a subphase pad. Add
unpad-subphase / overpad-subphase to the mutation set and a generated
`includeSub` boolean that decides whether the canonical carries a `.SS`
(forced in for the subphase mutations so there is always a `.SS` to mutate);
non-subphase mutations keep their original no-subphase coverage.

Non-vacuity verified against the compiled lib by temporarily probing each
widened/new property and confirming it fails: round-trip counterexample
["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit
tokens are genuinely generated and reach the body; a no-op unpad-subphase
mutation trips the mutated===canonical guard (counterexample
["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached
with a subphase present. Probes reverted; numRuns unchanged.

Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail;
`npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major)

BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that
re-derived the grammar as an unbounded `\d+` literal instead of consuming
PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket
path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug
whose first word is a year) over-collected the token as `01-14-2026` instead of
`01-14`.

Interpolating the cap verbatim at every position was rejected on evidence: the
bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent.
The exactly-2 cap at the others would under-collect ids toDir itself emits —
`PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit
sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and
`[GSD.100] 05` is a pinned regression. Those positions are delimiter-
disambiguated (a required field separator; a dot a slug can never contain),
not heuristically recognized, so they have no year collision to defend against.
Upstream draws the same line for the same reason: core-utils/phase cap the
paired PLAN component while the leading phase component stays unbounded.

So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and
each position takes the width its delimiter affords: leading unbounded, dash-1
and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner
seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of
the token grammar.

Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE
instead of re-spelling it as a literal, so the emit-side gate and the read-side
token source are one rule — the same single-owner discipline this fix is about.
Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant).

Refs #2249

* test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major)

The comment block alone cannot hold the divergence: src/phase-id.cts is exempt
from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not
catch the bracket token source drifting from the seam. Per the Generative Fix
Divergence rule, the divergence is pinned behaviorally instead.

Surface 6 joins the existing #2232 parity gate rather than starting a rival one:
the review named the bracket token source "a sixth continuation-recognition
site", and continuation-grammar-parity.test.cjs is already the invariant-named
home where the five #2043 sites agree with the owner on a shared width corpus.
Surface 6 asserts the same contract at the bracket run's slug-adjacent position
(`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so
the bracket path now fails the same gate the other five do.

A second block pins the DELIBERATE half — the wider canonical width at the
delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out
of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would
look like a cleanup rather than a regression.

The generative property ties the READ side to the EMIT side metamorphically: for
every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that
id's numeric run — no more, no less. It needed a new arbitrary: the existing
slugArb generates one [a-z0-9] word and so can never produce the number-leading
slug the collision requires.

Probe-falsified, both directions (probes reverted):
- reverting the source to the old unbounded `\d+` fails 8: the parity gate
  reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the
  review's scenario verbatim — and the property shrinks to
  ["A",1,1,undefined,"100-a"].
- interpolating the seam at EVERY position (the rejected verbatim option) leaves
  the repro and parity green but fails the divergence gate `'02' !== '02-105'`
  and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the
  evidence that a verbatim cap under-collects ids toDir emits.
Width 2 stays green under both probes — the corpus agrees with the owner exactly
where the old and new rules coincide, so the gate discriminates rather than
merely mirroring the regex.

Refs #2249

* docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major)

* test(#2249): pin deterministic grammar boundary cases (re-review m1)

PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
the m1 proof gap — the grammar's bounds were exercised only incidentally
through the fast-check domain (1-999, [a-z0-9] slugs). No source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged).

Adds a deterministic boundary block (7 describe groups, +22 tests) pinning
the compiled lib's CURRENT behavior — a proof gap, not a behavior gap:

- m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse
  (display + dir) -> render/toDir round-trip byte-equality. The plan
  position is identity-symmetric (parse/render accept 99/100/101) but toDir
  drops it (filename-surface dimension only).
- m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs
  99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps
  the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of
  the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry.
- m1.3 leading-zero 007 -> not-canonical rejection at every position/form.
- m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing
  slug (never stored, never mis-read as a plan) and rejects a line
  terminator; toDir's allow-list sanitizer collapses each to a safe
  [a-z0-9-] token or rejects sanitize-to-empty.
- m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test);
  an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE;
  an abs-path string is not a bracket id; an abs-path dir slug is dropped to
  a clean tuple.
- m1.6 whitespace-only -> not-a-bracket-phase-id.
- m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral):
  garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes.

No accept-not-reject case is a src bug: parse never STORES an abusive slug
(dropped from the identity tuple) and toDir independently re-sanitizes on
emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by
parse is the documented positional design (toDir drops the plan; the
read-token caps it) — divergence pinned, not papered over.

Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both
its parse-side and emit-side), ran -> 8 distinct named failures, reverted ->
66/66 green. Confirms every new group executes and can fail.

Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar +
continuation-grammar-parity + collision-characterization + phase-id family
175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green
except one pre-existing, unrelated env failure (npm-integrity-gate: a live
npm-audit advisory in the production dep tree — reproduces with this change
stashed; no package.json/lock change here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 12:49:53 -04:00
Tom Boucher
bf8f320083 feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI)

PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting
GSD's kimi support into two distinct products per the user's directive:
- kimi       (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python)
- kimi-code  (new):      Moonshot's Node Kimi Code CLI (~/.kimi-code,
                         runtime: node, KIMI_CODE_HOME env)

Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json
descriptors, not hardcoded branches in install.js. The new descriptor uses
the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml
hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs).

Critical Kimi Code constraint reflected in the descriptor:
  hostIntegration.dispatch.namedDispatch: false
  hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan']
  hostBehaviors.namedSubagentsSupported: false
Kimi Code's official docs confirm only 3 built-in subagents with NO custom-
subagent registration (the [subagent] table only has timeout_ms). The
kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in
kimi-code's artifactLayout.

Schema adjustments:
- subagentToolkit set to 'undocumented' (the existing escape hatch); the
  schema enum (full/read-only) lacks a 'limited'/'built-in-only' value.
  A follow-up PR can extend the schema enum to add 'built-in-only' as a
  first-class axis value reflecting Kimi Code's documented model.

Registration:
- capabilities/kimi-code/capability.json (new descriptor, modeled on codex)
- bin/install.js: allRuntimes array + --all list + --kimi-code flag
- gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases
  (kimi-code, kimicode, kimi_code)
- src/runtime-name-policy.cts: FALLBACK_ALIASES map
- gsd-core/bin/lib/capability-registry.cjs: regenerated via
  scripts/gen-capability-registry.cjs --write

Tests:
- tests/multi-runtime-select.test.cjs updated for the new runtime count (18)
  + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19.

Out of scope for PR 1 (follow-up PRs in the sequence):
- Install-time decision logic (kimi vs kimi-code detection / prompt)
- agent-install-check semantics for kimi-code (verify Agent Skills presence)
- cmdAgentSkills fallback returning subagent prompt content
- Workflow template mapping (named agents → built-in coder/explore/plan)
- Migration guidance for users currently on 'kimi' who are actually on Kimi Code
- Schema enum extension for subagentToolkit: 'built-in-only'

Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS).

* fix(#2454): complete drift-guard registrations for kimi-code runtime

The drift guards caught every surface that pins runtime enumeration. Each
update is mechanical, driven by the guard's named failure mode:

- src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code
- src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the
  isKimiCode predicate generator
- bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream
  entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19
- gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry
  (null/null/null — same as kimi, no model tier defaults until configured)
- docs/reference/capability-matrix.md: regenerated via
  scripts/gen-capability-matrix.cjs --write (kimi-code row added)
- tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP:
  kimi-code → '.kimi-code'
- tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden
  (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated)

The capability-registry is already regenerated from the prior commit.

* test(#2454): update drift-guard tests for kimi-code runtime registration

Multiple drift guards pin runtime enumeration counts and option numbering.
Each update is mechanical, driven by the guard's named failure mode:

- tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17);
  'all 16 flags' → 'all 17 flags' in test names + messages.
- tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering
  cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15,
  trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice
  test for kimi-code (option 11). Prompt test updated for new numbering.
- tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains
  kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs);
  EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per
  docs, same as Python kimi/opencode).
- tests/global-config-home-fragment.test.cjs: table-count test renamed
  13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier).

* fix(#2454): empty artifactLayout for kimi-code (PR 1 scope)

The skills kind requires a converter (existing converters are per-runtime
like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only
registers the descriptor; the actual Agent Skills converter (and a new
'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside
the install-time decision logic. Empty artifactLayout.global is valid and
means 'nothing to install yet via the layout seam'.

Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs
(localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as
option 11 in install.js's buildRuntimePromptText (renumbered downstream
options 11..17 → 12..18, All 18 → 19).

* fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode)

The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved)
for the new kimi-code runtime id. Property names with hyphens are awkward for
consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new
runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing
the conventional PascalCase flag name. The 16 prior single-word runtime ids
are unaffected (the regex finds no hyphens).

* fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures

- tests/runtime-flags.test.cjs drift guard: use proper kebab-case
  conversion (isKimiCode → kimi-code, not 'kimicode') so the registry
  comparison doesn't false-positive on hyphenated runtime ids.
- tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae
  single-choice tests for the renumbered options (kilo 11→12, opencode
  12→13, pi 13→14, qwen 14→15, trae 15→16).
- tests/install.test.cjs: Kilo integration option 11→12, prompt test
  regex updated.
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
  The kimi-code install produces the standard GSD install layout (skills,
  contexts, references, etc.) — 436 paths, same shape as other runtimes
  that have no custom converter yet.

* fix(#2454): add kimi-code install contract + global config home fragment

- src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code
  → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path
  instead of falling through to the default '.claude'.
- tests/installer-migration-install.integration.test.cjs
  RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for
  PR 1; PR 2 will specialize once the Agent Skills converter lands).
- tests/multi-runtime-select.test.cjs: fix space-separated-choices test
  for the renumbered kilo option (11 → 12).
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: regenerated after rebasing onto current next (new
  planner-reversibility.md from #2471 etc. now included).

* test(#2454): skip kimi-code install contract until PR 2 ships install layout

The end-to-end install test (tests/installer-migration-install.integration
.test.cjs) asserts every allRuntimes entry installs a runtime-specific
artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the
capability descriptor + flags + labels, but the install LAYOUT (Agent
Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands
in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and
self-removing — PR 2 removes the entry alongside adding the install
surface, restoring the contract loop to full coverage.

* fix(#2454): restore compact model-catalog.json format (M1 review)

Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations'
commit used python json.dump(indent=2) which inflated the file from 165→607
lines (every nested entry got expanded) and lost the trailing newline. The
semantic change was just a 3-line kimi-code entry. Restored the original
hybrid format (top-level indent=2 + inner entries' one-line style) and
added kimi-code in matching form.

Regenerated golden install parity + install tree fixtures since the
model-catalog.json hash changed.

* fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code)

CI lint-tests job failed on the glossary drift guard
(scripts/check-glossary-refs.cjs --check):
  ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but
    bin/install.js's allRuntimes array has 18.
  ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js
    (missing from CONTEXT.md's list: kimi-code).

Missed in the prior commits because gsd-test does not run the glossary
check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims
to 18 values + kimi-code in the member list.

* chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511)

* docs(changeset): Phase 1 kimi-code runtime Added (#2511)

* test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands

* docs(changeset): backfill PR #2519 for Phase 1 (#2511)
2026-07-22 00:27:35 -04:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
c5e0371775 feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging

Red phase for issue #1951 (reversibility tagging: classify decisions by
undo cost, gate one-way doors behind a checkpoint:decision).

Tests assert, per the issue's acceptance criteria:
- discuss-phase CONTEXT.md template records a **Reversibility:** field with
  a rationale on captured decisions, and states it is optional
- gsd-planner @-references planner-reversibility.md and stays under the
  49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs)
- a one-way rating inserts a checkpoint:decision before the dependent task;
  reversible inserts none; costly is flagged but never blocks
- the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard)
  and inserting a checkpoint implies autonomous: false
- docs/reference/plan-md.md documents <reversibility> as optional with all
  three ratings
- --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected
  into the planner prompt, and is advertised in the command argument-hint
  and help full mode (argument-hint parity)
- the override suppresses the gate but still persists the rating
- cmdVerifyPlanStructure accepts every rating and the absent case
  (additive-validator guarantee, behavioral via runGsdTools)
- parity: thinking-models-planning.md #4 adopts the canonical three-level
  taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone
- no content loss from the planner extraction made to fit under the cap

Prose-contract assertions are Red until the implementation lands. The
behavioral validator assertions pass immediately — regression guards
proving the validator already accepts unknown optional tags.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1951): reversibility tagging — gate one-way-door decisions

Classify planning decisions by what undoing them would cost, and give a
one-way door a human beat before the agent walks through it (issue #1951,
The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way
door framing).

Acceptance criteria met:
- discuss-phase records an optional reversibility rating with a rationale
  on <decisions> entries in the phase CONTEXT.md template. Unrated
  decisions are treated as reversible, so existing phases are unaffected.
- a one-way rating makes gsd-planner insert a checkpoint:decision before
  the task that implements the decision, reusing the existing checkpoint
  mechanism -- no new checkpoint machinery.
- reversible ratings trigger no checkpoint; costly ratings are flagged in
  the plan but never block.
- the rating persists on the task as the optional <reversibility rating=>
  element. cmdVerifyPlanStructure accepts every rating and the absent
  case; the structural validator does not reject unknown optional tags.
- --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses
  checkpoint insertion for intentionally-unattended runs while still
  recording ratings -- the override changes what stops the run, not what
  the plan remembers.

Single taxonomy, not two: references/thinking-models-planning.md #4
already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is
loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the
canonical three-level vocabulary and now points at planner-reversibility.md
as the taxonomy owner, with a parity test that fails if the surfaces
diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE).

agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the
checkpoint DO/DON'T guidance was relocated verbatim into
planner-antipatterns.md -- already @-referenced from the same section for
the same topic, so the planner still loads it and nothing was dropped. A
test guards the relocation against content loss.

Files: gsd-core/references/planner-reversibility.md (NEW, canonical
taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase
workflow/command/help (flag wiring + parity), plan-md.md schema,
discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest,
size baselines, install goldens, plugin skills regen, changeset.

Closes #1951

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1951): address orthogonal review findings

Two isolated reviewers (correctness + security), neither of which authored
the change. Every finding fixed:

Security — the rationale is untrusted input (ADR-1577). It originates in
conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each
hop an LLM reading the previous hop's output, with no validation on the
path. planner-reversibility.md and the discuss-phase template now state
it is data and never instructions, and name the </reversibility>
early-termination hazard explicitly -- a rationale that closes its own
element injects sibling structure the executor reads as real tasks.
Four tests guard it.

Correctness 1 — nothing machine-enforced the feature's own promise: a task
rated one-way with no preceding checkpoint:decision validated as fully
clean, so a planner error silently reopened the gap this feature exists to
close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A
warning, not an error: <reversibility> stays additive and the plan stays
valid. Four tests cover ungated (warns), gated (silent), still-valid, and
reversible/costly never flagged.

Correctness 2 — pass-always test. The --no-reversibility-gates parse test
substring-matched the whole workflow file, and plan-phase.md prose mentions
both tokens in one sentence, so it passed with the bash conditional
deleted: it was testing the documentation, not the parser. Now scoped to
the fenced bash blocks and matched as one physical line, with a negative
control confirming prose alone cannot satisfy it.

Correctness 3 — costly had no itemized emission rule, only one-way did, so
two agents could diverge on whether to tag costly at all.

Correctness 4 — template convention break: the example ratings were bare
while every sibling field uses [...] to signal substitution, inviting an
LLM to copy one-way/costly forward as boilerplate. Now bracketed.

Correctness 5 — latent false-green: .includes('reversible') also matches
inside irreversible/irreversibility, which appear in anti-pattern
prose, so a surface that dropped the real taxonomy entry would still pass.
Now word-boundary matched.

ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md
1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on
next). The ceiling may only rise for privileged host machinery, and
reversibility gating is optional-feature logic, so the wiring was slimmed
to its minimum and the explanatory prose moved to the reference files the
planner already loads. plan-phase.md is now 94400 bytes -- 119 under the
ceiling and 70 bytes SMALLER than on next, so the host loop shrank while
gaining the feature, which is what phase 6 ratchets toward. The tracer
contract (tests/tracer-bullet.test.cjs) is unchanged.

Lint — fixed an unnecessary non-null assertion in verify.cts and a
CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST-
PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture must carry the common task elements

The gated-one-way fixture built a checkpoint:decision task from the
abbreviated skeleton in gsd-planner.md, which shows only the
checkpoint-specific elements (<decision>/<context>/<resume-signal>).
cmdVerifyPlanStructure requires <name> and <action> on EVERY task
regardless of type, so the fixture failed validation for reasons that had
nothing to do with reversibility:

  errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"]

Caught by gsd-test on 14d14a39 (2 failures, both this fixture).

The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task
carries <name>/<files>/<action>/<verify> like any other. Fixture corrected
to match. Verified behaviorally against the real gsd-tools CLI across all
four cases: gated one-way (valid, silent), ungated one-way (valid, warns),
costly (valid, silent), absent (valid, silent).

Not a product defect: the validator's every-task contract is intentional
and pre-existing, and docs/reference/plan-md.md scopes its required-element
list to type=auto/tracer only because those are the elements a planner must
author, not because checkpoints are exempt from <name>.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1951): backfill changeset pr number to 2471

* fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision

Both CI failures were real defects in code this PR added, not false
positives.

CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 —
the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`,
which escapes the hyphen but not backslash, so the escape was incomplete.
It was also unnecessary: `-` carries no special meaning outside a character
class. Replaced with a complete metacharacter escape (backslash included).
Word-boundary behavior verified unchanged across all three ratings — notably
that "irreversible" prose still does not satisfy a "reversible" match, which
is the false-green this helper exists to prevent.

Prompt injection scan — the checkpoint fixture used the human-verification
child element inside <verify>. That tag name is a fake-instruction-boundary
pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed
files, so copying the shape from tests/verify.test.cjs (unflagged only
because it is not in this diff) tripped the gate. Switched to the documented
plain-prose <verify> form.

The first attempt at that fix failed the same gate a second time: the
comment explaining the collision quoted the offending tag literally. The
comment now names it in prose instead — the scanner does not care whether a
match is code or commentary, which is the whole point of the
DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md.

Verified locally before push: scan reports 0 findings across 57 changed
files, eslint clean, and both fixtures still validate as designed (gated
one-way silent, ungated one-way warns, neither errors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): record measured cost and halve gsd-tools spawns

The Windows shard 1/3 job timeout was traced to the sharding layer, not to
this PR's assertions — see #2472. Two contributing factors were this file's
own, and are fixed here.

1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs,
   so scripts/run-tests.cjs weighted it at the table's median fallback
   (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x
   under-weight. Recorded the measured value from the green gsd-test run
   (max across the node22/node24 lanes, per gen-test-timings.cjs's
   convention). Only this one entry: a full regen churns 634 entries of
   run-to-run drift, and the table is explicitly advisory and un-gated, so
   a 637-line diff does not belong in a feature PR.

2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost.
   Spawns cut from 9 to 6 with no coverage lost:
   - the ungated-one-way warning and its stays-valid assertion now share
     one plan instead of building the same plan twice;
   - the reversible/costly never-flagged-as-ungated test was strictly
     subsumed by the additive suite, which already runs those two ratings
     ungated and asserts no /reversibilit/ warning at all — and the gate
     warning's text contains both "reversibility" and "one-way", so the
     broader assertion catches it. It only re-spawned gsd-tools twice to
     prove the same thing.

Both are symptom fixes. The shard imbalance itself (19/11/10 minutes
against a 20-minute cap, from a cost-blind round-robin partition that also
reshuffles downstream files whenever one is inserted) is tracked in #2472.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture adopts the #2444 type-branched contract

Surfaced by rebasing onto next, which gained #2444 (branch plan-structure
validation on task type=checkpoint:*) while this PR was in review.

cmdVerifyPlanStructure no longer applies one required-element set to every
task. A checkpoint:decision now requires <name> + <resume-signal> +
<decision> + <options>, and is exempt from the <action>/<verify>/<done>/
<files> set that auto and tracer tasks carry. The gated-one-way fixture
predated that split and failed on the new requirement:

  errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"]

Fixture rewritten to mirror the checkpoint:decision contract exactly — real
<options> with two <option> children — rather than padding it with fields
checkpoints no longer need. That also drops the plain-prose <verify> the
earlier revision carried purely to dodge the prompt-injection scan; a
checkpoint task has no <verify> requirement at all, so the workaround is
moot.

Verified against the real gsd-tools CLI across all four cases: gated one-way
(valid, silent), ungated one-way (valid, warns), costly (valid, silent),
absent (valid, silent).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 10:44:55 -04:00
Tom Boucher
d16a66479a feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship

Adds a new  capability (#1950) that operationalizes GSD's
no-defer discipline as a tracked, enforced artifact:
accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths
across phases, and /gsd-ship blocks while any entry is open.

Implementation:
- src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR +
  I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed
  + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed
  error assertions. Windows-safe atomic rename with retry on transient
  EPERM/EBUSY/EACCES.
- gsd-tools.cjs: new  subcommand (status | append | waive | fixed),
  wired via routeWindows + HOST_COMMAND_ROUTERS.windows.
- capabilities/broken-windows/capability.json: one ship:pre gate with
  artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0.
  activationKey windows.enabled (default true) + sibling windows.enforce
  (default true, separate so tracking can precede enforcement).
- gsd-core/workflows/ship.md: capId==broken-windows branch in preflight,
  sibling to security — reads gsd_run windows status --raw, fails closed
  on open_count > 0 or unreadable ledger.
- agents/gsd-executor.md: extends the existing ## Known Stubs instruction
  to also append to WINDOWS.md via gsd_run windows append (best-effort,
  never blocks execution).
- agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify
  items in WINDOWS.md.
- gsd-core/workflows/progress.md: surfaces open + waived counts.
- docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md:
  document the gate, waiver mechanism, and new module.
- tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check
  roundtrip property; fail-closed on malformed ledger; security boundary on
  path traversal in --file.

Backward-compatible: a project with no .planning/WINDOWS.md reports
open_count: 0 and ships cleanly. Disable enforcement per-project with
gsd config-set windows.enforce false (tracking continues, gate stays open).

* chore(#1950): ratchet size baselines, defer verifier integration

- Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632
  (broken-windows preflight branch + open-windows surface).
- Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also
  appends to WINDOWS.md). gsd-verifier.md unchanged.
- LARGE_CAP (49152) preempted the planned verifier integration
  (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the
  documented 'real headroom'). Verifier integration deferred to a follow-up
  PR that extracts the VERIFICATION.md template (lines 739-859) to
  gsd-core/references/ — a pre-existing cap-tightness defect this PR
  exposed but does not expand scope to fix. Verifier integration is not in
  the issue's acceptance criteria (executor writes is; unmet-truths
  recording was an enhancement, not a gate).

* fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens

Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44
cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all
caps off → empty hooks'):

- capability manifest: rename windows.enabled+windows.enforce (default
  true) → single federated key workflow.windows_enforce (default FALSE,
  opt-in). Matches security's workflow.security_enforce convention and
  makes the adr857 all-caps-off test pass without modification (the test's
  buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps
  the gate out of the registry's default ship:pre resolution so existing
  loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId
  'security') stay valid; users opt in via
  gsd config-set workflow.windows_enforce true.
- drop activationKey (security doesn't have one either; workflow.* key
  doubles as the activation toggle).
- regenerate docs/reference/capability-matrix.md to include broken-windows
  (capability-matrix-sync test).
- regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) —
  installer now emits the new capability + lib file.
- update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md,
  agents/gsd-executor.md to use the new key name and /gsd:colon slash
  syntax (slash-command-namespace test).
- restore accidentally-regressed /gsd:capture in progress.md.

Tracking-only by default; enforcement is opt-in. Acceptance criterion
'/gsd-ship fails while any ledger entry is open' is met when
workflow.windows_enforce=true (test fixture enables it).

* test(#1950): update ship:pre structural invariants for 2-gate registry

- loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre
  (security + broken-windows), regardless of activation. Activation tests
  above still pin security-only or empty behavior via fixtures; these
  structural tests pin the REGISTRY shape, which has 2 gates as of #1950.
- workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce
  rename added 17 bytes).

* fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker

Adversarial isolated review (Step 6.3) found 2 HIGH findings that block
the PR and 3 mediums. All addressed:

H1 (HIGH): description containing the markdown 3-backtick fence would
terminate the ledger's JSON code block early inside JSON.stringify output
(JSON doesn't escape backticks), corrupting the file and bricking the
next parse. Fix: use a 4-backtick fence (json ... ) which
JSON.stringify cannot produce on its own, AND validate that no entry
text field contains a 4-backtick run (reject at append time with new
WINDOWS_INVALID_TEXT reason code). Locked by a regression test.

H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger',
silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate
would then pass on an unreadable ledger — the precise vector the
workflow doc claims is impossible. Fix: only ENOENT returns null;
every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the
gate blocks and the operator sees a real diagnostic. Locked by a
regression test that chmod 000s a ledger with open_count=1 and
asserts the result is never a false-green 0.

M1: writeLedgerAtomic left an orphaned .tmp file on rename failure.
Wrapped renameWithRetry in try/catch with best-effort unlink.

M2: validateLine silently coerced 'abc' → NaN → null, hiding type
drift. Removed the line === 0 special case (was undocumented) and
made the error message match the strict check. Now any non-positive-
integer line value throws, including strings.

M3: tests/broken-windows.test.cjs (with its fast-check property test)
was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate
src/broken-windows.cts but no test would catch the mutations,
producing false surviving-mutant scores. Added to the list.

L1 (dead throw e after error()), L7 (line boundary tests, H1/H2
regression tests, 4-backtick CLI test) also addressed.

* docs(#1950): inline concurrency + busy-wait notes (review L2+L3)

* fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test

gsd-test v4 caught two issues:
- goldens I regenerated earlier (commit 526682084) predated the L1
  routeWindows catch-block cleanup (commit dd844d565). Regenerated
  via 'npm run gen:golden' against current HEAD so the install
  parity hash for gsd-tools.cjs matches.
- 'append --line boundary' test expected --line 0 to succeed with
  null entry.line, but the M2 fix correctly rejects 0 (lines are
  1-indexed; 0 is not a valid source line). Updated the boundary
  test to assert --line 0 fails alongside -1 and 'abc'.

* chore(#1950): regen goldens after rebase onto next

* chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase)

* chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md

* fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization)

CodeQL flagged the markdown-table cell escaper:
  String(s ?? '').replace(/\|/g, '\\|')
— it escapes pipe but not backslash first. A description containing '\|'
would render as '\\|' which markdown parses as 'literal backslash' +
'cell separator', splitting the column.

Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|).
Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped
pipe), which markdown renders as a single '\|' inside the cell. The JSON
code block (the parse source-of-truth) was already correctly escaped via
JSON.stringify; only the display-only table was affected.

Locked by a regression test that:
1. Verifies the JSON block reparses with the description intact.
2. Walks the rendered table row counting unescaped pipes — must be
   exactly 11 (the row separators for 10 cells), proving no in-cell
   pipe added a split.
2026-07-19 20:24:21 -04:00
Behruz Nassre Esfahani
d04e287fa9 fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases

detectApiIntegration fired on any integration verb co-occurring anywhere on a
line with any API noun, treated / as a word boundary (so a first-party Next.js
src/app/api/... route path matched the noun "api"), and read any capitalized
word before API/SDK/REST/GraphQL as a service name behind a fixed stopword
denylist (so threat-model prose like "Resolver-only API" fired). Because the
verify:pre seal gate is BLOCKING, a phase touching no external API could not
reach UAT without fabricating a coverage matrix.

The compound rule now requires the verb and noun to share one clause (sentence
punctuation and table-cell walls end a clause) within a bounded word gap.
Non-prose spans are excluded before matching: fenced code (already), inline
code spans (new stripInlineCode in the markdown-sectionizer seam), and
path-shaped tokens. The <Service> API surface rule requires proper-noun
position — a clause-initial capitalized word is ordinary English and needs
dependency evidence (URL / package reference) on the same line — and rejects
compound modifiers ("Resolver-only", lowercase after the hyphen).

A phase that integrates no external API now has a first-class, reasoned way to
say so: a COVERAGE.md containing "No external API integration: <reason>"
satisfies the gate (declaration + rows is contradictory and blocks). The
true-positive path is pinned by regression tests: every default-vocabulary
positive still fires, including the widest word-gap pairing and the
surface-rule-only shape.

Fixes #2365

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2365): tighten api-coverage detector per Codex review (round 2)

Applies the Codex review findings on the initial #2365 fix:

- S-1: a COVERAGE.md "no external API integration" declaration is the human
  override for a fallible detector, so it must PASS even when detection still
  fires — but the contradiction is now SURFACED in the gate output (overridden
  signal count + terms) instead of passing silently.
- S-2: verb/noun pairing is now a term-group nearest-pair merge walk over
  precomputed word ordinals (computeWordStarts / minWordGap), not a match×match
  cross product — a hostile line repeating one pair thousands of times stays
  linear instead of going quadratic.
- FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`)
  are kept as noun/dependency evidence rather than being fully masked, so a
  genuine dependency reference inside code ticks still corroborates.
- C-1: the <Service> API surface rule now scans every candidate in every
  clause; a rejected first candidate no longer shadows a later genuine service.
- Cross-clause binding: a verb may bind a noun in the immediately following
  clause only when its own clause names a service object, within a tight gap —
  admits "Integrate Stripe, exposing its endpoints …" without re-admitting the
  unrelated-clauses false-positive class.
- Internal-descriptor negative evidence ("internal Payments API",
  "the internal endpoint") never pairs; URL/scheme matching generalized beyond
  http(s).

All 5 acceptance criteria still hold: the three reported false positives are
clean and "integrate the Stripe API" still fires. Built .cjs committed
alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) +
lint:regression-names clean; affected suites 256/256 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3)

Codex's second-round review found the round-2 tightening had over-corrected into
FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal
gate silently let through (the catastrophic class, since a missed API surface is
worse than a dismissable false positive). Retuned the detector to be explicitly
fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external
API integration" declaration dismiss the residual false positives.

Fail-open false negatives fixed (all now detect):
- F1 clause-initial `<Service> API` with a plain follower ("Stripe API for
  payment processing") — dropped the follower-allowlist / corroboration gate on
  clause-initial surfaces; a service that is not a stopword, descriptor, or
  compound modifier is a real name from any clause position.
- F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an
  alphabetic final label now contributes its API nouns; a first-party route
  path (no dotted host) still does not.
- F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the
  compound path no longer filters nouns on "internal"/"first-party" (Codex: the
  qualifier can describe the vendor's own API, not the consuming project's).
- F4 long single integration clause — removed the word-gap cap entirely: it
  could not separate a 21-word genuine clause from an 18-word internal one, so
  the clause boundary is now the whole relationship test.
- F5 lowercase cross-clause service — cross-clause binding no longer requires a
  capitalized "service object".

New false positives fixed (all now clean):
- F6 a URL token that swallowed a trailing clause comma, merging two clauses —
  trailing clause punctuation is kept literal so the split survives.
- F7 a capitalized internal component authorizing cross-clause binding — the new
  gate requires a dependent elaboration, not a new coordinate clause opened by a
  conjunction ("…, then document…").
- F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol
  and locality descriptors are rejected in the `<Service>` position.

- Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick
  runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms),
  semantics preserved (148 sectionizer tests unchanged).

Net simplification: the fail-closed model removed the round-2 minWordGap /
groupByTerm / follower / corroboration machinery (350 insertions vs 445
deletions across the touched files). Under fail-closed, three round-2 negative
tests now correctly detect (integration verb + "internal"-qualified noun, and
the distant-same-clause case); none were trek-e acceptance FPs.

Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown-
parsing) + lint:regression-names clean; all 8 review findings reproduced as
regression tests, both directions. Built .cjs committed alongside the .cts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4)

Codex's round-3 adversarial review found the fail-closed retune had introduced
new holes in both directions. Resolved:

Fail-open false negatives (now detect):
- External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an
  integration surface and contributes an endpoint noun even when the host names
  no vocabulary word. A bare domain link with no path ("https://example.com")
  stays a non-signal, so "Integrate … from example.com, document …" is still
  clean.
- Locality qualification ("internal", "private") no longer leaks across a
  sentence or clause boundary: only plain spaces may separate the descriptor
  from the service, so "The cache is private. Stripe API …" now detects.
- Cross-clause binding: the fragile head-word cap (which could not tell a
  genuine "Connect … to Stripe payments, exposing its endpoints" from an
  unrelated "Integrate … from URL, document …" — both 4 words after the verb)
  is replaced by a participial-continuation rule: a verb binds a noun in the
  next clause only when that clause begins with an "-ing" elaboration. This
  fixes the 4-word-head false negative AND the false positive below at once.

False positives (now clean):
- Cross-clause no longer binds a finite continuation regardless of separator:
  "Wire the settings form. Document endpoint props." / "…; document …" /
  "…, document …" are separate actions, not elaborations.

Perf (quadratic → linear):
- The trailing-punctuation peel is a backward char scan instead of an
  unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms).
- SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile
  "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms).

Consumer fail-open (blocking gate):
- readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is
  unreadable. On a read error the gate BLOCKS ("could not read the phase
  scope …") instead of silently certifying no-integration from partial scope —
  an unreadable plan could be the one describing the integration.

Documented fail-closed tradeoffs, now pinned with tests so they are not
"fixed" back into a fail-open: a clause-initial capitalized common word before
"API" ("Payment API", "Search API") reads as a service name; a long clause
pairs a verb with a distant noun; and a CommonMark inline code span that wraps
a newline is matched within-line only. Codex judged these acceptable because
the COVERAGE.md declaration is a cheap override.

One documented limitation remains out of scope: "Integrate Stripe, and
authenticate requests with its API" (a coordinate finite clause whose noun
refers back by pronoun) needs coreference resolution, beyond a lexical detector.

Verified: 379/379 affected + command-router tests pass (+14 new regression
tests covering every round-3 finding, both directions); tsc + eslint
(no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5)

Round-4 review confirmed the detector's two most complex features generate
findings in both directions no matter how they are tuned, because they need a
vendor dictionary + coreference the issue rules out in principle. Per the
operator's "ship the robust core" decision, both are removed and their gaps are
documented rather than chased further:

- Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused
  a fail-open on finite continuations ("Integrate Stripe; use its OAuth
  endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing
  endpoint terminology…" — wrongly fired). Detection is now same-clause only.
- URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint
  fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo
  link) and recreated routine UI-phase false positives. An external URL is
  evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1").

Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future
maintainer does not re-add the heuristics that caused the false positives above:
a service named only in a clause separate from its API noun, and a bare external
host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md
declaration and rare in real phase prose ("integrate the X API").

Also fixed from the round-4 review:
- Qualification now survives markdown emphasis ("The **internal** Payments API"
  stays clean) while still not crossing a sentence/clause boundary.
- readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the
  phase directory or reading the roadmap fallback — not only per-plan-file
  failures; a missing directory/section remains a legitimate no-op. The
  declaration-override path surfaces scope_read_error so an incomplete-scope
  override stays visible.
- SERVICE_SURFACE_API_RE length-bound comment no longer overclaims.

Net: the detector is same-clause verb+noun + `<Service> API` surface, with
path/code/inline masking and a fail-closed posture. All five acceptance criteria
hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) +
lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review)

The round-5 sanity review confirmed the detector simplification is sound (all
acceptance positives fire, all required negatives clean) and flagged one real
blocker plus a nit:

- Blocker: readPhaseScope's roadmap fallback could still silently pass an
  UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which
  returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent",
  no exception reached isRealReadFailure, and the blocking gate certified empty
  scope. Fixed at the source: read the roadmap directly and honor the function's
  OWN documented contract — null only on ENOENT (genuinely absent), otherwise
  throw. Both existing callers already wrap it in try/catch expecting that throw,
  and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by
  a new e2e test (unreadable roadmap fallback → block).

- Nit: the detectApiIntegration JSDoc still described the removed cross-clause
  participial binding and "every external hostname counts" — corrected to the
  actual same-clause-only behavior and the names-a-vocab-word URL rule.

Verified: full unit suite green; tsc + eslint + lint:regression-names clean.
Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): backfill changeset PR number (#2397)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): sync generated capability-registry + recapture install goldens

CI surfaced two generated-artifact staleness issues (all failing test shards +
lint-tests traced to these, not to a logic defect):

- gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the
  ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API
  integration" declaration section) but did not regenerate the registry, which
  embeds an inline copy of that fragment. Regenerated via
  `gen-capability-registry.cjs --write` — the diff is exactly the fragment text
  sync. Fixes `lint:generated-sync` and the "committed registry is in sync" +
  "registry integration" tests.

- The 18 golden-install-parity fixtures were stale by exactly one hash line each
  — `gsd-core/references/api-coverage.md`, which this PR edits and which is a
  hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that
  single hash per runtime and nothing else. Fixes the `golden parity — *` tests.

No source or behavior change — generated artifacts only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): flip representative-corpus manifest to assert the fixed behavior

The #2371 representative corpus (merged into next after this branch was cut) is a
known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test
fails loudly the moment #2365 is fixed, at which point — per its own contract in
representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the
assertion checks expectedDetected instead.

This is that moment. Removed currentBuggyOutput from the three detector fixtures
(nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now
asserts detected:false, which the fail-closed same-clause detector satisfies.
Notes updated to describe the fix rather than the bug. The #2366 matrix corpus
is left untouched — that tripwire belongs to its own PR (#2374).

Corpus test: 7/7 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): skip chmod-000 fail-closed e2e tests on Windows

The three fail-closed gate tests induce an unreadable plan / directory / roadmap
with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync
still succeeds, so the gate never reaches the read-error path and the assertion
fails on the windows-latest CI leg. The fail-closed LOGIC is platform-
independent (readError → block) and is fully exercised on the macOS/Linux legs;
only the method of inducing EACCES is POSIX-specific. Guard the three tests to
skip on win32 as well as root, mirroring golden-install-parity's win32 skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds

Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged;
this closes the test/process-hygiene findings.

Major:
- CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix
  relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate).
- Replaced the banned wall-clock assertion in the "hostile repeated-term line"
  test (Clock Seams rule — no elapsed-time asserts) with a deterministic
  signal-count assertion, which also directly verifies the term-dedup that keeps
  pairing linear (one signal for a 10k-pair line, not thousands).
- Rewrote the three fail-closed read-failure tests: instead of chmod 0o000
  (a no-op under root / on Windows, the pattern the repo's IO-failure convention
  avoids) they now exercise the newly-exported `readPhaseScope` in-process and
  inject the failure by monkeypatching fs.readFileSync/readdirSync to throw,
  restoring in finally. Deterministic and platform-independent (no skip needed),
  and they add the ENOENT-is-absence case that the chmod tests couldn't express.

Minor:
- Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40}
  service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN
  (200) on the declaration reason.
- Added a fast-check property that fuzzes the tokenizer / clause splitter /
  masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial
  tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector
  is total (never throws), shape-stable, holds detected <=> signals, and is
  deterministic.

readPhaseScope is exported for the in-process tests. Verified: 125 detector +
19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) +
lint-regression-test-names + lint-test-file-count clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 14:59:25 -04:00
Tom Boucher
1720aacf0c feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element

Red phase for issue #1949 (Design by Contract: <precondition> element
asserted before task execution). Tests assert:

- docs/reference/plan-md.md documents the new <precondition> element
- agents/gsd-planner.md @-references planner-preconditions.md and stays
  under the 49152-char cap (progressive-disclosure requirement)
- gsd-core/references/planner-preconditions.md exists and documents the
  three emission cases mandated by the issue (user_setup / prior-phase
  artifact / env-var) and the contract triad mapping
- agents/gsd-executor.md asserts <precondition> before task execution
  and routes unmet preconditions through existing checkpoint machinery
- cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both
  with and without <precondition> — the additive-validation guarantee
- Parity assertion: plan-md.md and planner-preconditions.md agree on the
  canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard)

Most prose-contract assertions are Red until the implementation lands.
The behavioral validator assertions pass immediately (regression guards
proving the validator already accepts unknown optional tags).

* feat(#1949): <precondition> task element — Design by Contract

Add an optional <precondition> element to <task> in PLAN.md (issue #1949,
The Pragmatic Programmer Topic 23). The front-of-task side of the plan
contract — preconditions (before) ↔ postconditions (<verify>/<done>/
<acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the
whole plan). Together with the tracer-bullet proposal (#1945), this closes
both ends of the 'outrunning your headlights' failure mode for an
autonomous AI executor.

Acceptance criteria met:
- <precondition> is an optional element on <task>; plans that omit it
  validate unchanged (cmdVerifyPlanStructure checks for presence of
  required tags, does not reject unknown optional tags).
- gsd-executor evaluates the precondition before any other task work.
  Unmet halts execution with a checkpoint:human-verify and no partial
  commit; met or absent produces no visible change to execution flow.
  Unmet is never auto-approved under AUTO_CFG=true — a missing
  prerequisite is a fact the executor cannot establish on its own.
- gsd-planner emits <precondition> in exactly the three cases the issue
  mandates: user_setup consumption, prior-phase artifact dependency, and
  env-var/runtime-config dependency.
- Tests cover met, unmet, and absent preconditions plus the additive-
  validator guarantee.

Files:
- gsd-core/references/planner-preconditions.md (NEW): full emission
  rules, the three cases with worked examples, format guidance,
  anti-patterns, the contract triad mapping, and the executor assertion
  contract. Progressive disclosure.
- agents/gsd-planner.md: slim <precondition> note in Task Anatomy with
  @-reference to the new file. To stay under the 49152-char agent-file
  cap (27-char headroom before this change), the inline
  <comment_text_discipline> and <region_scoped_negative_gate> summaries
  are compressed to one-line pointers — their full rules already live in
  planner-antipatterns.md, so no content is lost.
- agents/gsd-executor.md: new step 0 'Precondition check' in the
  execute_tasks loop, before the type dispatch, routing unmet through
  checkpoint_return_format.
- docs/reference/plan-md.md: new Preconditions section in the schema
  reference, with the canonical example and the three emission cases.
- CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet.
- docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new
  references/planner-preconditions.md (regen via gen-inventory-manifest).
- tests/precondition-element.test.cjs: failing-first tests covering
  schema docs, planner emission contract, executor assertion contract,
  reference-file presence + the three cases, behavioral additive-
  validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX-
  DIVERGENCE guard).
- .changeset/quick-hawks-bark.md: Added fragment.

Companion to #1945 (tracer bullets).

* chore(#1949): regen agent-size baseline + install-tree goldens

Documented baseline regenerations required by the feat(#1949) prose changes
(RULESET.AGENT_SIZE_BUDGET + golden-install-parity):

- npm run size:baseline — locks in the new gsd-executor.md size (+1050
  bytes: the precondition-check step 0 block). gsd-planner.md is net
  smaller (-142 bytes: compressed two inline summary blocks whose full
  rules already lived in planner-antipatterns.md to make room for the
  slim <precondition> pointer). No hard-cap breach.
- npm run gen:golden — pick up the new references/planner-preconditions.md
  + the two changed agent files across all 18 runtime install trees.

Both regens are CI-mandated after intentional agent/reference changes;
see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in
tests/golden-install-parity.test.cjs.

* fix(#1949): bound <precondition> checks to read-only (security review)

Apply the security-review finding (LOW, isolated /security-review subagent):
the executor's 'run the cheapest check' phrasing for a plan-author-controlled
prose line was broader than ideal — a hostile plan author could craft a
<precondition> whose 'cheapest check' is side-effecting (curl to an attacker
host under the guise of verification, rm -rf before checking, secret emission).

The risk is inherited from GSD's existing plan-trust model (<verify>, <action>,
<done> already direct the executor to run arbitrary shell), so <precondition>
does not materially expand it. But the new prose actively directs execution
('run the check') rather than passively consuming the element, so the bound
is worth making explicit.

Tightened across all four surfaces that describe the check shape:
- agents/gsd-executor.md step 0: 'Verify with read-only checks only — file
  existence, env var presence (no value output), idempotent GET /health-style
  pings. Do NOT run commands with side effects (writes, network POSTs, secret
  emission) as the check; if a side-effecting check seems required, halt and
  surface via checkpoint instead.'
- gsd-core/references/planner-preconditions.md Format section: same bound,
  plus the halt-and-surface escape hatch.
- docs/reference/plan-md.md Preconditions section: mirrored.
- CONTEXT.md Precondition glossary entry: mirrored.

Regenerated agent-size baseline (executor grew 46186 -> 46440; still under
the 49152 cap) and install-tree goldens.

* chore(#1949): backfill changeset pr number 2422

Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7:
backfill the placeholder pr:0 with the real PR number immediately after
gh pr create returns. Avoids the fail_invalid_fragment gate.

* fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456)

CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires
every // allow-test-rule: exemption on a NEW test file to carry an issue
reference (#NNN or URL). My earlier push omitted it.

Local 'npm run lint' (eslint) does NOT run this check — only 'npm run
lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs
lint:ci; a local pass is not the gate.' I should have run lint:ci before
pushing; correcting now.

Pattern matches the companion feature's test file:
tests/tracer-bullet.test.cjs:1  // allow-test-rule: source-text-is-the-product [#1945]
2026-07-19 07:52:36 -04:00
Tom Boucher
f2c077df38 chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate (#2391)
* chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate

Apply the audit-and-enforce concept from the ADR index (#2356) to CONTEXT.md:
correct stale facts, and add a CI gate so the machine-verifiable claims can't
silently re-rot.

CONTEXT.md was entirely hand-maintained with nothing checking its claims against
the shipped tree, so it had rotted. An audit against live code (Memtrace +
filesystem + gh), each finding adversarially re-verified, drove 38 factual
corrections + 1 surfaced by the new gate:

- Dead references: Package Identity named @opengsd/get-shit-done-redux (package
  is @opengsd/gsd-core); Shell Command Projection named run-git/run-npm/run-tool
  (real exports execGit/execNpm/execTool); a partial docs/adr/1606 ref; retired
  sdk/ framing.
- Superseded facts: allRuntimes 15 -> 17 (pi #2102, zcode); "seven nested-loader
  runtimes" -> five (claude reverted flat #924, antigravity flat); stacked-PR
  examples rebasing onto main -> next; QUOTA_SENTINELS precedence corrected to
  match src/agent-command-router.cts.
- Drifted CONTRIBUTING.md line citations refreshed.

Per CONTRIBUTING.md:179, only stale FACTS were corrected -- no maintainer intent,
lesson, or opinion was rewritten, and the append-only session log is untouched
except one dated in-place superseding note. The three tests that assert on
CONTEXT.md content (phase6-capstone-conformance, tracer-bullet,
external-job-waiting) keep all their anchors.

New scripts/check-glossary-refs.cjs (--check, wired into lint:generated-sync):
- Check A: every backticked file reference under a TRACKED_PREFIXES allowlist
  resolves on disk. Generated gsd-core/bin/lib/*.cjs (77 refs, gitignored),
  ~/-paths, .planning/, and bare filenames are deliberately skipped so a clean
  CI checkout never false-fails.
- Check B: the allRuntimes count + member set in the glossary prose match
  bin/install.js's allRuntimes literal (drifts on every runtime addition).
tests/check-glossary-refs.test.cjs covers both, including the false-positive
guard that a missing bin/lib/*.cjs ref does NOT trip the gate.

Closes #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2387): confine glossary-gate file refs to ROOT (no `..` traversal)

Pre-PR security review finding (low): extractTrackedRefs fed tokens straight to
fs.existsSync(path.join(ROOT, token)), and PATH_TOKEN_RE admits `.` in a segment,
so a CONTEXT.md token like `src/../../../etc/passwd` passed the `src/` prefix
check and normalized to an out-of-tree absolute path — turning the doc lint into
a filesystem-existence oracle on the CI host (existsSync only; CONTEXT.md is a
trusted committed file, hence low severity, but a defense-in-depth gap).

Add isWithinRoot() confinement in extractTrackedRefs: a token is dropped unless
path.resolve(ROOT, token) stays within ROOT. A CONTEXT.md reference is always a
plain in-repo path, so a `..` escape is never legitimate. Regression test asserts
a `..`-bearing token is skipped and never named in output.

Refs #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2387): drop legacy `get-shit-done` name from a CONTEXT.md defect entry

CI lint-legacy-dir-name failed: the line-928 upstream-issue re-point I applied
wrote the historical provenance as "gsd-build/get-shit-done#3545", and
scripts/lint-legacy-dir-name.cjs forbids the legacy `get-shit-done` name. Reword
to "moved from #3545 in the predecessor repo" — same provenance, no legacy name.

Caught by `npm run lint:ci` (the CI lint chain), which I had not run locally —
lint:generated-sync + eslint do not include lint-legacy-dir-name.

Refs #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 19:24:28 -04:00
Tom Boucher
15b3cc8690 docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a
two-layer dispatch (registry families + leaf-verb table filling the prepared
_dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions
ADR-959 leaves open: full dissolution, family/leaf classification rule, shared
parseFamilyArgs, and the capability-arm extraction shape. Phased under epic
#2345 (P1-P4). Behavior-preserving; each cutover proven by the
audit-command-cutover equivalence template.

- docs/adr/2346-command-dispatch-completion.md (new)
- docs/adr/959-*.md: Status Proposed -> Accepted + amendment section
- docs/adr/README.md: index rows for 959 + 2346
- docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub
- CONTEXT.md: seed glossary entry

Closes #2346 (docs-only; no production code).
2026-07-17 07:19:17 -04:00
Tom Boucher
315d94f6d4 feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate

Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode.

- gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top.
- gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer.
- --no-tracer flag wired through plan-phase workflow/command/help/skill.
- CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled.
- tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1945): backfill changeset PR number to 2294

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:41:36 -04:00
Tom Boucher
c4237df8e6 docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a
conceptual Embeddable Orchestration System (EoS) explanation
(docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md
with a v1.7.0 feature section, and wire both new docs into the docs index
(docs/README.md) and the root README.

Covers the release's marquee changes: the ADR-1239 Host-Integration Interface /
EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS
discoverability registries, the gsd-mcp-server companion, model-catalog advances
(GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format,
plus a themed summary of the 100 fixes and 4 security hardenings.

Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay
now documents the #2009 fail-open behavior for a load-failed gate-declaring
capability (previously described as fail-closed).

American house style; no parity-gated reference docs hand-edited.

Refs #2276, #1678

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:49:44 -04:00
Tom Boucher
8b70db343b fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207

completePhaseCore was writing the overloaded bare 'Milestone complete' on the
last phase — the same string space the milestone-close verb owns for terminal
state. Per ADR-2207, phase-completion now writes the existing intermediate
value 'All phases complete' (already used in gsd2-import.cts). Milestone
termination ('<version> milestone complete' / 'Awaiting next milestone')
remains solely with milestoneCompleteCore.

Status lifecycle: Ready to plan → All phases complete → <version> milestone
complete → Awaiting next milestone.

Changes:
- src/state-transition.cts: completePhaseCore status value
- src/phase.cts: #2028 guard comment
- tests/state-transition.test.cjs: assertion + test name
- tests/phase.test.cjs: 8 assertion updates (positive + negative)
- tests/state.test.cjs: normalizeStateStatus test case + reset regex
- tests/workstream.test.cjs: fixture status to terminal value
- gsd-core/workflows/progress.md: Route D label
- gsd-core/workflows/transition.md: Route B label
- CONTEXT.md: Status lifecycle glossary entry (ADR-2207)
- .changeset/brave-geese-jump.md

* test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline

Workflow file edits (progress.md, transition.md) changed install payload
hashes and pushed past the committed workflow-size baseline. Regenerated
all 17 golden-install-parity fixtures + claude-local via the standalone gen
script (which now also covers the local-scope claude layout). Updated
workflow-size-baseline.json and agent-size-baseline.json via size:baseline.

* fix(#2204): correct claude-local golden hashes + document gen-script limitation

The gen-script's claude-local generation produces macOS-specific hashes
incompatible with Linux CI (local-scope install embeds platform-varying
node-runner paths). Reverted to manual update using Linux FAILURES.md
+actual hashes for the 2 changed workflow files. Added explanatory
comment in the gen script.

* test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary

Addresses orthogonal code-review findings (Medium #1 + #2):
- Add isCompletedInventory test cases for ADR-2207 status lifecycle
  (terminal 'milestone complete' → true; intermediate 'All phases
  complete' → false; archived → true; active statuses → false)
- Clarify CONTEXT.md glossary: note that isCompletedInventory
  intentionally excludes the intermediate value

* docs: backfill changeset PR number (#2259)

* docs(#2204): add Status lifecycle table to state-md reference (ADR-2207)
2026-07-14 14:47:03 -04:00
Tom Boucher
2cbf186420 chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3

Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112,
#2118) were already fixed tactically on next; this introduces the reusable
structural contracts and rewires the primary #2140 site onto them.

- src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the
  per-surface write-set (`WriteOutcome {surface, applied, requirement?}`,
  `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports
  `Result` from here (single source; distinct from command-routing-hub's Result).
- requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT,
  per-surface write-set; `write_set_complete` is true only if every surface of
  every requirement applied — structurally forbidding the #2140 OR-into-one-flag
  masking, including across a multi-ID batch (adversarial-review regression).
  Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed).
- deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial
  null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow;
  RoadmapProgress return contract unchanged.
- commit --files (#2112) and milestone complete --dry-run (#2118) left as-is
  (single-surface commit / pre-mutation preview — not genuine multi-surface writes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2244): backfill changeset PR number (#2251)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:39:58 -04:00
Tom Boucher
efd04716da chore(#2143): withSection bounded-mutation seam + phase.cts migration — Phase 2 (#2250)
* chore(#2143): bounded-mutation seam (withSection/withPhaseSection) + phase.cts migration — Phase 2

Phase 2 of epic #2143 (ADR-2143 §4): add a bounded-mutation primitive so a
per-phase ROADMAP edit is structurally confined to that phase's own section,
and migrate the phase-scoped mutation sites in `phase.cts` onto it.

- `src/markdown-sectionizer.cts`: `withSection(content, target, edit, opts?)` —
  resolves a section via `collectSection` and applies `edit` to ONLY that
  section's body, re-serialising via `replaceSection`. The edit callback sees
  only the section body, so any regex it runs is physically confined.
- `src/roadmap-parser.cts`: `withPhaseSection(content, phaseId, edit)` —
  resolves a phase's `### Phase N` detail-section heading via the #2121
  phase-id source and delegates to `withSection`. Heading match is anchored to
  the heading start (a sibling phase whose title mentions the number is not
  hijacked) and bounds at the next ATX heading of any level (`levelBounded:false`).
- `src/phase.cts`: `mutateMilestonePhase`'s plan-count and per-plan-checkbox
  writes now route through `withPhaseSection` — structurally retiring the
  #2130 / #2067 / #2080 boundary-crossing class for these sites. The phase-LIST
  checkbox is intentionally left milestone-slice-scoped (it lives outside any
  `### Phase N` detail section). Cross-phase renumbering is untouched.
- Property test (fast-check): editing phase k leaves every sibling section
  byte-identical; regression tests for title-collision + mixed heading depth.

Behaviour-preserving (verified by old-vs-new differential runs on real
fixtures). Extend-never-mutate (ADR-2143 §2). Registration: CONTEXT.md +
docs/INVENTORY.md export lists.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2243): backfill changeset PR number (#2250)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:39:22 -04:00
Tom Boucher
d49ac81306 chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1

Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a
canonical seam and migrate the pilot reader.

- Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed
  {columns, rows} addressed by column NAME; ragged rows are typed parse
  errors, not silent), a single-source TABLE_SCHEMAS registry
  (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with
  variants under one id), matchTableSchema, and findTableBySchema. Result<T>
  is scoped to this seam (distinct from the dispatch Result).
- Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the
  position-anchored regex to name-based resolution via the seam — fixes #2137
  (the 5-column milestone-grouped Progress table previously returned all-null).
- Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route
  fast.md's log_to_state through it, retiring the inline `awk NF-2` column
  arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped
  (| and newlines) and the STATE.md read-modify-write is atomic under
  readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230).
- Writer/reader/template parity test guards TABLE_SCHEMAS against drift
  (ADR-2143 §3 Generative-Fix-Divergence).

Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md +
INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md.

Behaviour-preserving for the canonical 4-column Progress table; the named
bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2242): backfill changeset PR number (#2248)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): escape backslash before pipe in markdown-table cell escaping

CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not
the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow
unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl.
literal backslashes) round-trip exactly. Added backslash round-trip tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan

Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself
"pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via
a new seam helper findTableWithColumns (first table whose header is a superset of
Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads
cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact
TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while
staying seam-based and preserving its `## Progress` scoping (#2012/#1445).
Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale
state.test.cjs assertion that predated the Phase-1 migration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:36:14 -04:00
Tom Boucher
60e3c4988a feat(#2182): scaffold community capability + EoS registry (tests + stubbed core)
Adds the discoverability-registry surface for issue #2182: JSON-sourced
capability/eos catalogs, a pure schema/vocab module (registry-schema.cjs)
with the ADR-857 loop points + ADR-1239 axes, thin validate/gen CLIs,
the registry-entry PR template, README spec, and CONTEXT.md glossary terms.

The three pure functions (isValidGsdRange/validateEntries/renderMarkdown)
are stubbed here so the comprehensive test suite fails first (red), per the
feature-implementation red-first directive; the next commit implements them.

Refs #2182

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 14:13:02 -04:00
Tom Boucher
474ca08e06 docs(#2171): record the statusline data-source scope boundary (#2178)
Records the statusline scope boundary decided during triage of #2160-2164:
the statusline sources only local, read-only data (refine-existing + new-local),
never credentials or external/network APIs. #2164 (account-usage segment) is
out of scope on this boundary; #2163 (git) is in-scope but on the feature track;
#2160/2161/2162 are approved enhancements.

- docs/adr/2164-statusline-scope-boundary.md (new ADR, Accepted)
- docs/adr/README.md (index row)
- CONTEXT.md (### Statusline glossary/seam entry)
- .out-of-scope/statusline-account-usage.md (#2164 rejection record)

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 12:52:21 -04:00
Tom Boucher
5695522d5f feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads:
getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix
(→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite),
getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead
inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity
destructures (antigravity is already on the descriptor-agents path). subagentToolkit
flipped undocumented→full (Context7: antigravity.google/docs/cli/features);
namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical
golden parity for all 16 runtimes.

UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions
merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's
settings.json — non-destructive, idempotent, symmetric uninstall. Added to
VALID_PERMISSION_WRITERS + the FinishPermissionWriter union.
UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json
registering the gsd-core companion MCP server (Gemini-successor mcpServers schema,
best-effort — raw schema unpublished). Both writers dispatch from finishInstall.
settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable,
no absolute paths) is golden-tracked → only antigravity.json changes.

Tests: declarative-reference-antigravity extended (source-grep guard across 4
modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) +
antigravity-upgrades (permission-writer + mcp_config live-install, idempotency,
user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md +
connect-gsd-mcp-server docs updated; changeset (Changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 03:07:01 -04:00
Tom Boucher
ab04916682 feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven
hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle,
reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES.
Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by-
name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain.

UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker-
delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml
in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts).
GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7-
confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/
SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards
removed). config.toml holds absolute install paths so it's golden-excluded via
an exact relative-path (.kimi/config.toml), not a basename (which would blind
Codex's config.toml). New hooksSurface value added to the closed enum in
capability-validator + runtime-config-adapter-registry.
UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's
Agent tool takes run_in_background; root agent already gets the Agent tool), so
negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented'
per AC (coder/explore/plan have distinct tool policies).
MCP transport explicitly deferred (no installer-driven MCP for any runtime).

Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others +
claude-local byte-identical (kilo/zcode keep their own exclusions). Tests:
kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep
guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency
+ backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated;
changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 00:53:34 -04:00
Dave
31a500b970 docs(#1867): replace stale plan-phase.md:921 line-pointer with section-name reference (review #6) 2026-07-10 14:40:03 -04:00
Dave
c1756d0cd5 chore(#1867): register ui-consideration probe + regen install cascade (SHIP-01)
Ship-safe registration + regenerated snapshots for the #1867 UI-consideration
probe (Phase 3, SHIP-01):

- CONTEXT.md: PROBE.ui.{verification,axis,seam} predicates + ui-consideration
  -probe added to PROBE.family (machine-canon for the 3rd adapter, MIXED axis).
- agents/gsd-ui-{researcher,checker}.md: one @-include of
  references/ui-consideration-probe.md each (both under the 24576 agent cap).
- docs/INVENTORY.md + INVENTORY-MANIFEST.json: register the reference doc and
  the compiled ui-consideration-probe.cjs (inventory-manifest-sync green).
- tests/fixtures/golden-install-parity/*.json (16 runtimes): recaptured against
  a clean full build — folds in the deferred Phase-1 (ref doc, plan-phase lift)
  and Phase-2 (ui-phase step, UI-SPEC section) install-surface changes.
- tests/agent-size-baseline.json: ratcheted the two grown UI agents.
- .changeset/vivid-orcas-chatter.md: type Added (pr updated at PR-open).

Inventory/golden/size gates green; lint:ci + lint:docs + lint:changeset green.
The plan-phase.md PRE_PHASE6 ceiling stays RED pending #1852 (unchanged).

Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
2026-07-10 14:40:03 -04:00