Commit Graph

753 Commits

Author SHA1 Message Date
Tom Boucher
c5e0371775 feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging

Red phase for issue #1951 (reversibility tagging: classify decisions by
undo cost, gate one-way doors behind a checkpoint:decision).

Tests assert, per the issue's acceptance criteria:
- discuss-phase CONTEXT.md template records a **Reversibility:** field with
  a rationale on captured decisions, and states it is optional
- gsd-planner @-references planner-reversibility.md and stays under the
  49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs)
- a one-way rating inserts a checkpoint:decision before the dependent task;
  reversible inserts none; costly is flagged but never blocks
- the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard)
  and inserting a checkpoint implies autonomous: false
- docs/reference/plan-md.md documents <reversibility> as optional with all
  three ratings
- --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected
  into the planner prompt, and is advertised in the command argument-hint
  and help full mode (argument-hint parity)
- the override suppresses the gate but still persists the rating
- cmdVerifyPlanStructure accepts every rating and the absent case
  (additive-validator guarantee, behavioral via runGsdTools)
- parity: thinking-models-planning.md #4 adopts the canonical three-level
  taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone
- no content loss from the planner extraction made to fit under the cap

Prose-contract assertions are Red until the implementation lands. The
behavioral validator assertions pass immediately — regression guards
proving the validator already accepts unknown optional tags.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1951): reversibility tagging — gate one-way-door decisions

Classify planning decisions by what undoing them would cost, and give a
one-way door a human beat before the agent walks through it (issue #1951,
The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way
door framing).

Acceptance criteria met:
- discuss-phase records an optional reversibility rating with a rationale
  on <decisions> entries in the phase CONTEXT.md template. Unrated
  decisions are treated as reversible, so existing phases are unaffected.
- a one-way rating makes gsd-planner insert a checkpoint:decision before
  the task that implements the decision, reusing the existing checkpoint
  mechanism -- no new checkpoint machinery.
- reversible ratings trigger no checkpoint; costly ratings are flagged in
  the plan but never block.
- the rating persists on the task as the optional <reversibility rating=>
  element. cmdVerifyPlanStructure accepts every rating and the absent
  case; the structural validator does not reject unknown optional tags.
- --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses
  checkpoint insertion for intentionally-unattended runs while still
  recording ratings -- the override changes what stops the run, not what
  the plan remembers.

Single taxonomy, not two: references/thinking-models-planning.md #4
already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is
loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the
canonical three-level vocabulary and now points at planner-reversibility.md
as the taxonomy owner, with a parity test that fails if the surfaces
diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE).

agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the
checkpoint DO/DON'T guidance was relocated verbatim into
planner-antipatterns.md -- already @-referenced from the same section for
the same topic, so the planner still loads it and nothing was dropped. A
test guards the relocation against content loss.

Files: gsd-core/references/planner-reversibility.md (NEW, canonical
taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase
workflow/command/help (flag wiring + parity), plan-md.md schema,
discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest,
size baselines, install goldens, plugin skills regen, changeset.

Closes #1951

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1951): address orthogonal review findings

Two isolated reviewers (correctness + security), neither of which authored
the change. Every finding fixed:

Security — the rationale is untrusted input (ADR-1577). It originates in
conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each
hop an LLM reading the previous hop's output, with no validation on the
path. planner-reversibility.md and the discuss-phase template now state
it is data and never instructions, and name the </reversibility>
early-termination hazard explicitly -- a rationale that closes its own
element injects sibling structure the executor reads as real tasks.
Four tests guard it.

Correctness 1 — nothing machine-enforced the feature's own promise: a task
rated one-way with no preceding checkpoint:decision validated as fully
clean, so a planner error silently reopened the gap this feature exists to
close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A
warning, not an error: <reversibility> stays additive and the plan stays
valid. Four tests cover ungated (warns), gated (silent), still-valid, and
reversible/costly never flagged.

Correctness 2 — pass-always test. The --no-reversibility-gates parse test
substring-matched the whole workflow file, and plan-phase.md prose mentions
both tokens in one sentence, so it passed with the bash conditional
deleted: it was testing the documentation, not the parser. Now scoped to
the fenced bash blocks and matched as one physical line, with a negative
control confirming prose alone cannot satisfy it.

Correctness 3 — costly had no itemized emission rule, only one-way did, so
two agents could diverge on whether to tag costly at all.

Correctness 4 — template convention break: the example ratings were bare
while every sibling field uses [...] to signal substitution, inviting an
LLM to copy one-way/costly forward as boilerplate. Now bracketed.

Correctness 5 — latent false-green: .includes('reversible') also matches
inside irreversible/irreversibility, which appear in anti-pattern
prose, so a surface that dropped the real taxonomy entry would still pass.
Now word-boundary matched.

ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md
1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on
next). The ceiling may only rise for privileged host machinery, and
reversibility gating is optional-feature logic, so the wiring was slimmed
to its minimum and the explanatory prose moved to the reference files the
planner already loads. plan-phase.md is now 94400 bytes -- 119 under the
ceiling and 70 bytes SMALLER than on next, so the host loop shrank while
gaining the feature, which is what phase 6 ratchets toward. The tracer
contract (tests/tracer-bullet.test.cjs) is unchanged.

Lint — fixed an unnecessary non-null assertion in verify.cts and a
CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST-
PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture must carry the common task elements

The gated-one-way fixture built a checkpoint:decision task from the
abbreviated skeleton in gsd-planner.md, which shows only the
checkpoint-specific elements (<decision>/<context>/<resume-signal>).
cmdVerifyPlanStructure requires <name> and <action> on EVERY task
regardless of type, so the fixture failed validation for reasons that had
nothing to do with reversibility:

  errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"]

Caught by gsd-test on 14d14a39 (2 failures, both this fixture).

The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task
carries <name>/<files>/<action>/<verify> like any other. Fixture corrected
to match. Verified behaviorally against the real gsd-tools CLI across all
four cases: gated one-way (valid, silent), ungated one-way (valid, warns),
costly (valid, silent), absent (valid, silent).

Not a product defect: the validator's every-task contract is intentional
and pre-existing, and docs/reference/plan-md.md scopes its required-element
list to type=auto/tracer only because those are the elements a planner must
author, not because checkpoints are exempt from <name>.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1951): backfill changeset pr number to 2471

* fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision

Both CI failures were real defects in code this PR added, not false
positives.

CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 —
the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`,
which escapes the hyphen but not backslash, so the escape was incomplete.
It was also unnecessary: `-` carries no special meaning outside a character
class. Replaced with a complete metacharacter escape (backslash included).
Word-boundary behavior verified unchanged across all three ratings — notably
that "irreversible" prose still does not satisfy a "reversible" match, which
is the false-green this helper exists to prevent.

Prompt injection scan — the checkpoint fixture used the human-verification
child element inside <verify>. That tag name is a fake-instruction-boundary
pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed
files, so copying the shape from tests/verify.test.cjs (unflagged only
because it is not in this diff) tripped the gate. Switched to the documented
plain-prose <verify> form.

The first attempt at that fix failed the same gate a second time: the
comment explaining the collision quoted the offending tag literally. The
comment now names it in prose instead — the scanner does not care whether a
match is code or commentary, which is the whole point of the
DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md.

Verified locally before push: scan reports 0 findings across 57 changed
files, eslint clean, and both fixtures still validate as designed (gated
one-way silent, ungated one-way warns, neither errors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): record measured cost and halve gsd-tools spawns

The Windows shard 1/3 job timeout was traced to the sharding layer, not to
this PR's assertions — see #2472. Two contributing factors were this file's
own, and are fixed here.

1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs,
   so scripts/run-tests.cjs weighted it at the table's median fallback
   (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x
   under-weight. Recorded the measured value from the green gsd-test run
   (max across the node22/node24 lanes, per gen-test-timings.cjs's
   convention). Only this one entry: a full regen churns 634 entries of
   run-to-run drift, and the table is explicitly advisory and un-gated, so
   a 637-line diff does not belong in a feature PR.

2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost.
   Spawns cut from 9 to 6 with no coverage lost:
   - the ungated-one-way warning and its stays-valid assertion now share
     one plan instead of building the same plan twice;
   - the reversible/costly never-flagged-as-ungated test was strictly
     subsumed by the additive suite, which already runs those two ratings
     ungated and asserts no /reversibilit/ warning at all — and the gate
     warning's text contains both "reversibility" and "one-way", so the
     broader assertion catches it. It only re-spawned gsd-tools twice to
     prove the same thing.

Both are symptom fixes. The shard imbalance itself (19/11/10 minutes
against a 20-minute cap, from a cost-blind round-robin partition that also
reshuffles downstream files whenever one is inserted) is tracked in #2472.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture adopts the #2444 type-branched contract

Surfaced by rebasing onto next, which gained #2444 (branch plan-structure
validation on task type=checkpoint:*) while this PR was in review.

cmdVerifyPlanStructure no longer applies one required-element set to every
task. A checkpoint:decision now requires <name> + <resume-signal> +
<decision> + <options>, and is exempt from the <action>/<verify>/<done>/
<files> set that auto and tracer tasks carry. The gated-one-way fixture
predated that split and failed on the new requirement:

  errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"]

Fixture rewritten to mirror the checkpoint:decision contract exactly — real
<options> with two <option> children — rather than padding it with fields
checkpoints no longer need. That also drops the plain-prose <verify> the
earlier revision carried purely to dodge the prompt-injection scan; a
checkpoint task has no <verify> requirement at all, so the workaround is
moot.

Verified against the real gsd-tools CLI across all four cases: gated one-way
(valid, silent), ungated one-way (valid, warns), costly (valid, silent),
absent (valid, silent).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 10:44:55 -04:00
Tom Boucher
909a3b180b fix(#2470): install pi's extension as gsd.js so pi actually discovers it (#2478)
* test(#2470): failing-first — pi extension must satisfy pi's auto-discovery filter

pi auto-discovers extensions/ entries through isExtensionFile(), which accepts
only .ts and .js. GSD installs its extension as gsd.cjs, so pi silently skips
it: no /gsd command, no error, no log line.

Encodes pi's discovery PREDICATE rather than a literal filename, so the
contract keeps holding across future renames, and adds the migration-006 test
matrix for retiring the stale gsd.cjs left in pre-fix installs.

Red until the fix lands.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): install pi's extension as gsd.js so pi actually discovers it

pi auto-discovers extensions/ entries via isExtensionFile(), which accepts
only .ts and .js and skips everything else silently. capabilities/pi declared
the dest as gsd.cjs, so the extension installed correctly and was then ignored
forever: no /gsd command, no error, no log line.

Install it as gsd.js. The in-repo source stays pi/gsd.cjs — tests require() it
directly and .cjs is unambiguous CommonJS; only the installed name has to
satisfy pi, and pi loads accepted files through jiti, which handles CJS and ESM
alike. (The reporter's premise that ~/.pi/agent/package.json declares
"type":"commonjs" does not hold — pi never writes that file.)

Renaming an installed artifact requires a migration record, so add 006 to
retire the stale gsd.cjs from pre-fix installs; without it the old path drops
out of the manifest and uninstall can never remove it. The migration plans
nothing for an unmanifested gsd.cjs: emitting remove-managed there would have
the executor downgrade it to preserve-user and mark it blocked, failing the
install for anyone who hand-placed their own file.

Also pins body-parser >=2.3.0 (GHSA-v422-hmwv-36x6). The advisory reaches the
production tree transitively via the Claude Agent SDK and fails the
npm-integrity gate, blocking any PR; pinned via the existing overrides idiom.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): address orthogonal review findings + register migration checksum

Code review:
- pi/gsd.cjs's install docstring still told readers to copy the file to
  extensions/gsd.cjs — the exact silently-broken state this PR fixes. Anyone
  following it recreated the bug.
- Two stale extensions/gsd.cjs comments in install-minimal-hooks.test.cjs.

Security review:
- _installNativePluginIfDeclared confined nativePlugin.dir but joined
  nativePlugin.file onto the validated directory unchecked, so a descriptor
  whose file carried .., an absolute path, or a NUL byte would have written
  outside configHome. Not reachable in a shipped build (descriptors are
  first-party and compiled into the capability registry), but file is exactly
  the field this PR changes. Confine the full dest path instead; for a
  well-formed descriptor this resolves identically to the previous
  mkdir(dir) + join(dir, file). Covered by four new write-confinement tests.

Also register migration 006 in the #670 EXPECTED_CHECKSUMS baseline — shipped
migration bodies are locked to a committed checksum and a new migration fails
CI until it is listed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): never dereference a symlinked managed path when snapshotting

fs.copyFileSync follows symlinks, so a managed path replaced by a link had the
REFERENT's bytes copied into the migration journal's rollback and backup trees
— a gsd.cjs symlinked at a private key would land that key's contents under
gsd-migration-journal/. Deletion was already safe (fs.rmSync unlinks the link,
never the target); the copy was not.

Nothing GSD installs is ever a symlink, so the faithful snapshot of a symlinked
managed path is the link itself. copyPreservingSymlink recreates it, which
keeps rollback fidelity (restore re-creates the same link) while never reading
the referent. Scoped the pre-delete to the symlink branch only, so the
regular-file path keeps copyFileSync's overwrite-in-place and a mid-restore
failure cannot destroy the destination. The restore-side existence check moves
to lstat, since existsSync follows a link whose target is gone and would
silently skip the restore.

This lives in the engine all six migrations share, so 000-005 are hardened too.

Also regenerates the pi golden-parity hash: correcting pi/gsd.cjs's own install
docstring changes the extension's content, which the golden suite caught.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): symlink-preserve the in-apply failure-recovery restore too

The previous commit routed three copy sites through copyPreservingSymlink but
missed a fourth: the catch block inside applyInstallerMigrationPlan, which
replays rollback snapshots taken earlier in the SAME apply attempt. Those
snapshots are symlinks precisely because of that commit, so the raw
copyFileSync there dereferenced them and wrote the referent's bytes to the LIVE
install path — worse than the journal-tree leak it was meant to fix, since it
is user-visible and at a predictable location.

Verified by experiment rather than assertion: with the pre-fix line restored,
the managed path comes back as a REGULAR FILE containing the referent's bytes;
with the fix it comes back as a symlink and the bytes appear nowhere.

The accompanying test injects the failure by letting the delete succeed and
then throwing once, modelling a later step failing after the delete. That
ordering is load-bearing — an earlier draft injected before the delete, which
leaves the live path in place, so the pre-fix copyFileSync hit a same-file
collision and threw instead of leaking. That draft passed against the bug it
was written to catch; this one fails against it.

Adds the missing rollback() coverage as well: a restored symlinked managed path
must come back as a link pointing at its original target.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2470): read the backup location from the journal, not the plan

The new backup-content assertion read backupRelPath off result.plan.actions,
where it is always null: the planner reserves the field and apply chooses the
concrete location, recording it in the journal. The assertion therefore failed
on "backup path must be recorded for the user" rather than on anything about
the behavior it was written to check.

Read it from the journal, which is the authoritative record. Verified by
executing all four new test bodies in-process against the built engine — the
backup file exists and holds the locally patched content.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2470): backfill changeset pr number to 2478

* chore(#2470): backfill changeset pr number to 2478

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 08:26:47 -04:00
Tom Boucher
953b8043ea fix(#2456): weight test chunks by measured cost and pack with LPT (#2463)
* fix(#2456): weight test chunks by measured cost and pack with LPT

scripts/run-tests.cjs guessed each test file's cost from its filename
(basename matching /^(?:install|codex-)/ scored 12, everything else 1).
Measured durations show that guess is wrong in both directions:
installer-migration-authoring.test.cjs scored 12 while running ~0.1s, and
the two most expensive files in the suite both scored 1 —
run-tests-harness.test.cjs never matched the prefix, and
release-tarball-smoke.install.test.cjs was missed because the regex is
anchored to the START of the basename.

Chunks were therefore balanced by file COUNT, not cost. On the real
shard 2/3 the two heaviest files packed into the SAME chunk, leaving the
slowest chunk 2.8x the lightest and sitting near the 600s per-chunk
timeout while other chunks idled.

Weight each file by its measured duration from a checked-in, regenerable
timings table and pack with LPT (heaviest first, into the lightest
chunk). On the same shard this drops the slowest chunk from 383s to 238s
and the imbalance from 2.79x to 1.00x, and separates the two heavy files.

Timings are advisory, never gated: an unknown file falls back to the
table's median weight, a missing or corrupt table falls back to uniform
weight, and a count-based floor guarantees the packer never produces
fewer chunks than plain count-based packing would.

Closes #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): harden chunk packing against degenerate knobs and table keys

Follow-up hardening found while reviewing the packer, fixed inline.

The chunk knobs are read from the environment with Number(), so a typo
(RUN_TESTS_MAX_FILES_PER_CHUNK=abc) yields NaN and an explicit 0 yields
0. Both flow into the new chunk-count arithmetic: NaN made Math.ceil
return NaN, Array.from({length: NaN}) produce zero bins, and packChunks'
retry loop spin forever — a hung CI job with no output. Zero made the
count Infinity and threw RangeError: Invalid array length. The previous
count-based packer degraded to a single chunk instead, so this was a
regression introduced by the LPT rewrite.

Normalize the knobs at the environment boundary (positiveNumberEnv:
anything not a positive finite number falls back to the default) and
guard packChunks itself, since it is exported and cannot assume its
caller normalized. Non-finite weights from an arbitrary weightOf are
clamped too. RUN_TESTS_CHUNK_TIMEOUT_MS gets the same treatment.

Also resolve timing-table lookups with Object.hasOwn: the table is
JSON-parsed, so a bare index would walk the prototype chain and return a
function for a file named constructor.test.cjs or toString.test.cjs.
The typeof guard already rejected that, but the lookup now resolves
correctly rather than relying on the downstream check.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): correct prototype-lookup rationale and guard generator keys

Two findings from independent security review, fixed inline.

The makeFileWeigher comment claimed a bare table lookup "would return a
FUNCTION for a file named constructor.test.cjs". That premise is false:
basename('constructor.test.cjs') is 'constructor.test.cjs', which is not
an Object.prototype key, and walkTestFiles only ever collects *.test.cjs.
The prototype chain was never reachable from a real selection, and the
existing typeof guard already rejected the function it would return, so
Object.hasOwn is defense-in-depth rather than a behavior change. The
comment now says that instead of asserting something untrue.

The accompanying test inherited the same false premise: it fed
constructor.test.cjs and asserted a median fallback that would have held
with or without the guard, so it passed for a reason unrelated to what
it claimed to prove. It now uses BARE keys (constructor, toString,
valueOf, hasOwnProperty, __proto__) — the only inputs that actually
resolve on Object.prototype — and asserts the real exported contract:
any key absent from the table weighs the median, never a function.

gen-test-timings.cjs built its output object by computed-key assignment
from basenames taken out of a reporter stream it does not control — the
js/prototype-polluting-assignment shape, and this repo has a CodeQL
barrier for exactly that pattern. It was not exploitable (the value is
always a rounded number, so the __proto__ setter is a silent no-op), but
it silently DROPPED such an entry rather than reporting it. Validate every
key against a test-basename pattern and fail loudly instead, and build
the table with a null prototype.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): replace tautological chunking tests and clamp chunk count

Six findings from independent correctness review, all reproduced and
fixed inline.

The two subprocess tests written to carry the #2088 guarantee forward
were tautological: every seeded file weighed exactly 1, so both passed
under the OLD prefix-heuristic packer and with the timings file deleted
entirely. Neither could fail for the reason it existed. Both are rebuilt
so the old algorithm produces a different packing and the assertion goes
red: the spread test now uses three expensive files named so the old
heuristic scored them 1 alongside three trivial `install-`-prefixed
files it scored 12 — inverted from real cost, giving {2,2,1,1} under the
old packer versus {2,2,2} under measured weights. The companion test
covers the other direction: four trivial `install-` files the old
heuristic split into four single-file chunks now stay in one.

packChunks clamped the chunk count from below but not above, so a
legitimate but tiny budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9, which
positiveNumberEnv accepts) asked for 637,000,000,000 bins and threw
RangeError. More chunks than files is never useful; the count now clamps
at one file per chunk.

The generator's basename-collision guard compared full dirnames, so two
OS lanes reporting the same file under different container roots
(/work/tests vs C:/work/tests) flagged every shared basename as a
collision — on the script's own documented multi-lane usage. Detection is
now scoped per stream, where the root is constant; a genuine same-lane
collision is still caught.

Also: the LPT tie-break compared raw paths, so a path separator (0x2F vs
0x5C) could order a subdir file differently per platform, contradicting
the documented byte-identical guarantee — it now normalizes separators.
loadTestTimings now honors schema_version instead of writing it and
never reading it, falling back to uniform weight on an unknown version.
A comment claiming an all-uniform suite "chunks exactly as it did
before" was false and contradicted by this PR's own test: the chunk
count is preserved, the composition is not. And the missing-table test
created a temp dir it never cleaned up, for a path that only needed to
not exist.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:51:35 -04:00
Tom Boucher
455ad49ae3 feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded

Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.

Red until the resolver, CLI flag, and manifest key land.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2296): config-gated provider escalation on quota-exceeded

The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.

- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
  capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
  Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
  the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
  validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
  passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
  every model tried once the ladder is spent.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2296): extract quota recovery to a reference fragment; regen goldens

The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.

That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.

Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2351): make the C1 orphan-reaping test load-independent

tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.

The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.

Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md

* chore(#2296): regenerate fixtures after rebase onto #2402

The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (b6e6a22fc) regenerated the same
artifacts. Conflict resolution picked a side to unblock the rebase; a true
regeneration on the combined tree then produced further drift, confirming the
resolved content was stale and would have dropped #2402's fixture changes.

Regenerated goldens, install-tree, size baseline, and INVENTORY-MANIFEST from
the merged tree. docs/INVENTORY.md keeps BOTH new reference rows.

execute-phase.md is 92782 LF bytes with both #2402's and this PR's extractions
applied — under the frozen ceiling (hard <93600, margin <=93400).

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 14:59:30 -04:00
Tom Boucher
b6e6a22fce fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer

Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage
(seven commits, never pushed) onto current origin/next as a single squashed commit.
The original work was substantial and correct; this commit preserves its full scope,
trimmed where rebase conflicts + workflow size budgets required it.

Three independent layers where response_language was being dropped are closed:

Layer 1 — orchestrator-facing directives across workflows. Adds the strong
"All user-facing output in this workflow MUST be presented in {response_language};
technical terms, code, paths, and subagent prompts stay in English" directive
to ~40 workflows that previously either lacked it entirely (verify-work,
new-project, new-milestone, quick, manager, and ~35 more) or carried only
the weak subagent-prompt-only form (plan-phase, execute-phase). The directive
covers narration between tool calls and banner output, not just the
AskUserQuestion prompts.

Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts
an optional responseLanguage parameter and renders the frame strings
("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.")
in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/
Chinese/Korean/Italian) with an alias table covering ~30 input variants
(en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads
config.response_language via loadConfig(cwd) and passes it through, so the
byte-for-byte block verify-work.md reprints verbatim is already localized
when written — preserving the anti-injection hygiene rule at verify-work.md
(the model is forbidden to translate after the fact). CJK display width is
computed by East Asian Width property ranges (W/F) so the right ║ border of
the banner stays aligned for full-width characters. English fallback is
byte-identical to the pre-fix behavior when response_language is unset or
unrecognized.

Layer 3 — literal English report templates in execute-phase. The top-of-
workflow directive covers all template sites (templates are a structural
source, not literal output). Inline render-language notes that previously
sat at each template site were removed during the squash because they
pushed execute-phase.md over its frozen pre-phase-6 byte ceiling
(93600 — ADR-857 Phase 6 capstone). The single top directive covers the
same surface with fewer bytes.

Also extends src/docs.cts and src/init.cts to propagate response_language
into the init JSON bundle of the additional workflows so the directive can
read it.

Tests added:
- tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls
  back to English default; recognized language swaps only the two frame
  strings while structural lines stay untouched; CJK display-width regression
  (independent recomputation of East Asian Width W/F ranges).
- tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language
  wiring through docs.cts/init.cts.

References: #2402; reporter's three-layer triage + Layer-4 follow-up; the
byte-for-byte anti-injection hygiene rule at verify-work.md (the reason
Layer 2 must be renderer-side, not model-translated).

This is a squash of the in-flight bot branch — seven commits representing
the original implementation plus its subsequent fix/CJK-padding/test/
changeset/regen cycles, none of which were ever pushed or PR'd. The squash
captures the final coherent state.

* chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md

* chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451)

Rebase conflicts were entirely in generated artifacts (golden-install-parity
fixtures + workflow-size-baseline.json). After taking theirs during rebase,
regenerated cleanly against the merged source tree.
2026-07-20 14:22:27 -04:00
Lanny Boarts
6c00cdad07 docs(#2343): list gsd-omp EoS integration (#2448)
* docs: list gsd-omp EoS integration

* chore: add gsd-omp registry changeset

* chore: update registry changeset PR reference

---------

Co-authored-by: AI Assistant <ai@example.com>
2026-07-20 07:59:18 -04:00
Tom Boucher
12e4d93b19 fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts (#2445)
* fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts

Bug: v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused
install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like
skills/, hooks/) was a pre-existing symlink, with no opt-out. Three
legitimate user-owned layouts were blocked:

  - (lars-hh) CLAUDE_CONFIG_DIR=~/.claude-personal with skills/hooks
              symlinked to a user-owned external dir
  - (Mamiki)  ~/.claude/skills is a Windows Junction to a shared skills dir
  - (Azd325)  ~/.claude itself is a symlink to a dotfiles repo (the early
              root-is-symlink return refused before the component loop ran)

Fix: add GSD_ALLOW_SYMLINKED_DEST env var (accepts '1' or 'true'). When
set, hasExistingSymlinkBetween follows symlinks instead of refusing them.
Cross-platform: fs.lstatSync().isSymbolicLink() returns true for both
POSIX symlinks and NTFS junctions (Node ≥ 16), so Mamiki's Junction case
is handled by the same code path.

Threat model preserved (these still refuse EVEN WITH opt-in):
  (a) path-traversal in the destSubpath string itself ('../../etc'-style)
      — ADR-1239 Phase B threat (a), untrusted destSubpath protection
  (b) a symlink whose resolved real path equals the install root itself
      — would let _removeGsdEntries wipe the root; #1704 threat (b)
  (c) broken symlinks (realpathSync throws) — fail-closed

What opt-in RELAXES specifically: the 'pre-existing symlink pointing
outside configHome' refusal — #1704 threat (c). The user has explicitly
asserted they own and trust the symlink target.

Error messages at all 4 call sites (installRuntimeArtifacts, _copyStaged,
migrateLegacyDevPreferencesToSkill, installOpencodeFamilySkills) updated
to (1) name the env var opt-in, (2) be accurate when the root itself is
a symlink (Azd325's complaint that the old message accused destDir of
'containing' a symlink when the root was the actual symlink).

Docs: docs/CONFIGURATION.md Environment Variables table updated.

Regression tests in tests/install-write-confinement.test.cjs cover:
  - child-symlink layout (lars-hh / Mamiki): default refuses, opt-in allows
  - root-is-symlink layout (Azd325): default refuses, opt-in follows
  - path-traversal '../../etc' refused EVEN WITH opt-in (threat a preserved)
  - resolved-target-equals-install-root refused EVEN WITH opt-in (threat b)
  - broken symlink refused EVEN WITH opt-in (fail-closed)

* test(#2393): import beforeEach/afterEach in install-write-confinement suite

The original file imported only { describe, test } from node:test. The new
#2393 opt-in describe block uses beforeEach/afterEach to manage the
GSD_ALLOW_SYMLINKED_DEST env var lifecycle — add them to the import.

* test(#2393): correct broken-symlink test — existsSync follows link → loop terminates early

Initial test expected broken symlinks to be refused even with opt-in. That
was wrong: fs.existsSync follows symlinks, so a broken symlink returns
false from existsSync and the component loop terminates before the symlink
check fires. Both default and opt-in paths share this behavior; the fix
preserves it.

Updates the test to pin the actual current behavior so a future refactor
(e.g. switching to lstatSync for existence) is a deliberate behavior change.

* fix(#2393): realpath the install root — guard against macOS /var ↔ /private/var

Code review (security subagent) flagged a HIGH-severity hole in the threat-(b)
preservation: realTarget (from fs.realpathSync) is fully symlink-resolved,
resolvedRoot (from path.resolve) is lexical-only. On macOS /var is a symlink
to /private/var, so resolvedRoot='/var/foo/.claude' but realConfigHome is
'/private/var/foo/.claude'. A symlink whose realtarget matches the install
root by real path would compare unequal to the lexical resolvedRoot —
defeating the wipe-protection guard exactly in the reporter's case (Azd325,
nix-darwin: ~/.claude is itself a symlink).

Fix: compute realRoot once via fs.realpathSync(resolvedRoot) at function entry
(with fail-closed fallback to lexical form on realpath failure — broken/missing
root, permission denied, exotic FS). Threat (a) path-traversal check above
still confines regardless. Compare against BOTH lexical and real forms in
both the root-symlink and component-symlink branches.

Also adds the reviewer's transitivity-trust clarification comment: once a
symlink is followed under opt-in, the walk continues from the resolved real
path WITHOUT re-checking further segments stay inside a confining boundary.
This is documented opt-in semantics — one opt-in trusts the whole reachable
tree — and the comment makes the design choice explicit so a future
maintainer doesn't add a 'follow one symlink only' expectation.

Regression test added for the macOS /var normalization case (spelled configHome
via os.tmpdir() lexically while pointing the test symlink through its realpath).
Test skips on non-darwin platforms and when os.tmpdir() has no symlink component.

* fix(#2393): root-symlink branch — do not apply threat-(b) check to root itself

Initial fix applied the wipe-threat-(b) check to the root-symlink branch
unconditionally. That was wrong: when root itself is a symlink (Azd325's
nix-darwin case), its realpath IS realRoot by construction — so the check
always fires, defeating the opt-in for exactly the case it was meant to
enable.

The wipe threat (b) does NOT apply to root being a symlink: destDir is a
CHILD of root, and resolving root just gives root's target. There is no
circular back-reference to root from a path that descends from a resolved
root. So the root-symlink branch should just follow the symlink under opt-in
and continue the walk, no threat-(b) check.

Threat (b) only fires in the COMPONENT loop, where a child symlink can
resolve back to the install root. That branch keeps the (b) check using BOTH
lexical and real forms of root (the macOS /var ↔ /private/var fix from the
prior commit).

* fix(#2393): apply opt-in at the 5 bin/install.js call sites + add env-var/transitive tests

Code review (correctness subagent) flagged a Critical coverage gap: the
initial fix updated only the 4 src/install-engine.cts call sites. Five
more call sites in bin/install.js still used the 2-arg signature, so the
opt-in env var was silently ignored on:

  - installCodexConfig (config.toml + agents/ dir + per-agent .toml paths)
    — Codex only
  - copyWithPathReplacement (the generic emit path: workflows, commands,
    staging) — ALL runtimes
  - resolveInstallRelativePath (path resolver used in various places)

Result: a user setting GSD_ALLOW_SYMLINKED_DEST=1 would see SOME refusals
disappear (engine path) and OTHERS remain (bin/install.js paths) — a
partially-applied install and a confusing UX, directly contradicting the
PR's headline claim.

Fix:
- Export isSymlinkedDestOptIn from src/install-engine.cts alongside
  hasExistingSymlinkBetween
- Import it in bin/install.js
- Update all 5 bin/install.js call sites to pass { allowOptInFollow }
- Update all 3 bin/install.js error messages to name the env var, matching
  the engine's phrasing

Also addresses reviewer's Medium test-adequacy findings:
- isSymlinkedDestOptIn env-var parsing now tested directly (accepts only
  documented '1' / 'true'; rejects 'TRUE', 'yes', 'on', '0', 'false',
  empty, unset)
- transitive symlink chain (configHome/outer → outside1 → outside2) test
  pins the documented 'transitive and unbounded' opt-in semantics so a
  future contributor can't accidentally narrow it

* chore(changeset): backfill pr:2445 in .changeset/eager-wasps-swim.md
2026-07-20 07:56:36 -04:00
Tom Boucher
517bae8d6d fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message

Bug: check.decision-coverage-plan's remediation message told the user to
cite decisions "(or body)" but extractPlanDesignatedSections only scanned
<objective>/<tasks>/<task>/<action>. A decision cited in <read_first>,
<behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to
the gate — false BLOCKING coverage gap, plus the message's own fix-hint
sent the user to "the body" where re-citing still failed.

Two-part fix (must change together — that drift was the bug):

1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also
   match <read_first>, <behavior>, <verify>, <acceptance_criteria>,
   <done>. These are all planner-canonical tags the planner is told to
   use (plan-phase.md:830-862, plan-phase.md:772). The body negative-
   lookahead mirrors the opening-tag set so each tag's body is captured
   independently.

2. Correct buildPlanMessage to name ONLY the surfaces the extractor
   actually scans (front-matter must_haves/truths/objective,
   designated markdown headings, and the nine planner-canonical tag
   bodies). The misleading "(or body)" clause is gone.

Also updates the planner's documented contract (agents/gsd-planner.md:69)
and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to
reflect the wider scan.

Regression tests in tests/decisions.test.cjs cover each newly-scanned
tag body, a control (no citation still uncovered), and a message/extractor
parity assertion that names every scanned surface — so the two cannot
drift apart again.

Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage
is a separate command (decision-coverage-verify) checking shipped
artifacts, not plan citations — untouched.

* chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures

gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision-
coverage contract (5 new scanned tag names + heading clarification).
Growth is justified: the contract surface is itself the fix — the prior
text under-described what the gate scans, which was the bug.

Updates:
- tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294)
- 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime)
- tests/fixtures/install-tree/*.json (regenerated by gen:golden)

* fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting

Code review (subagent) flagged a Medium edge-case regression from the
single-alternation regex: when a newly-scanned tag nests inside another
scanned tag, the alternation's negative lookahead halts the outer tag's
body at the inner tag — losing any D-NN citation in the outer tag's
prefix prose. Concretely:

  <action>per D-05 <verify>npm test</verify></action>

  → 3-tag alternation (old):  captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught
  → 9-tag alternation (bug):  captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST

Switches extractXmlTagBodies to per-tag matching: each tag gets its own
regex whose negative-lookahead tempers only against the SAME tag's
reopening. So <verify> inside <action> is absorbed into <action>'s body
(D-05 caught) AND <verify> is matched separately on its own pass.

Per-tag preserves both:
- the reporter's case (sibling tags inside <read_first>)
- nested-tag citations in outer-tag prefix prose
- ReDoS safety (each per-tag regex keeps the #2128 body tempering)

Also adds the reviewer's other requested edge-case tests:
- non-scanned tag (<name>) bearing D-NN must NOT count
- self-closing form <read_first /> safely ignored
- attribute form <verify type="...">D-NN</verify> (canonical planner shape)
- CRLF newlines in tag body do not break capture

* chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md
2026-07-19 23:07:13 -04:00
Tom Boucher
d16a66479a feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship

Adds a new  capability (#1950) that operationalizes GSD's
no-defer discipline as a tracked, enforced artifact:
accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths
across phases, and /gsd-ship blocks while any entry is open.

Implementation:
- src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR +
  I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed
  + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed
  error assertions. Windows-safe atomic rename with retry on transient
  EPERM/EBUSY/EACCES.
- gsd-tools.cjs: new  subcommand (status | append | waive | fixed),
  wired via routeWindows + HOST_COMMAND_ROUTERS.windows.
- capabilities/broken-windows/capability.json: one ship:pre gate with
  artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0.
  activationKey windows.enabled (default true) + sibling windows.enforce
  (default true, separate so tracking can precede enforcement).
- gsd-core/workflows/ship.md: capId==broken-windows branch in preflight,
  sibling to security — reads gsd_run windows status --raw, fails closed
  on open_count > 0 or unreadable ledger.
- agents/gsd-executor.md: extends the existing ## Known Stubs instruction
  to also append to WINDOWS.md via gsd_run windows append (best-effort,
  never blocks execution).
- agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify
  items in WINDOWS.md.
- gsd-core/workflows/progress.md: surfaces open + waived counts.
- docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md:
  document the gate, waiver mechanism, and new module.
- tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check
  roundtrip property; fail-closed on malformed ledger; security boundary on
  path traversal in --file.

Backward-compatible: a project with no .planning/WINDOWS.md reports
open_count: 0 and ships cleanly. Disable enforcement per-project with
gsd config-set windows.enforce false (tracking continues, gate stays open).

* chore(#1950): ratchet size baselines, defer verifier integration

- Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632
  (broken-windows preflight branch + open-windows surface).
- Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also
  appends to WINDOWS.md). gsd-verifier.md unchanged.
- LARGE_CAP (49152) preempted the planned verifier integration
  (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the
  documented 'real headroom'). Verifier integration deferred to a follow-up
  PR that extracts the VERIFICATION.md template (lines 739-859) to
  gsd-core/references/ — a pre-existing cap-tightness defect this PR
  exposed but does not expand scope to fix. Verifier integration is not in
  the issue's acceptance criteria (executor writes is; unmet-truths
  recording was an enhancement, not a gate).

* fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens

Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44
cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all
caps off → empty hooks'):

- capability manifest: rename windows.enabled+windows.enforce (default
  true) → single federated key workflow.windows_enforce (default FALSE,
  opt-in). Matches security's workflow.security_enforce convention and
  makes the adr857 all-caps-off test pass without modification (the test's
  buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps
  the gate out of the registry's default ship:pre resolution so existing
  loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId
  'security') stay valid; users opt in via
  gsd config-set workflow.windows_enforce true.
- drop activationKey (security doesn't have one either; workflow.* key
  doubles as the activation toggle).
- regenerate docs/reference/capability-matrix.md to include broken-windows
  (capability-matrix-sync test).
- regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) —
  installer now emits the new capability + lib file.
- update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md,
  agents/gsd-executor.md to use the new key name and /gsd:colon slash
  syntax (slash-command-namespace test).
- restore accidentally-regressed /gsd:capture in progress.md.

Tracking-only by default; enforcement is opt-in. Acceptance criterion
'/gsd-ship fails while any ledger entry is open' is met when
workflow.windows_enforce=true (test fixture enables it).

* test(#1950): update ship:pre structural invariants for 2-gate registry

- loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre
  (security + broken-windows), regardless of activation. Activation tests
  above still pin security-only or empty behavior via fixtures; these
  structural tests pin the REGISTRY shape, which has 2 gates as of #1950.
- workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce
  rename added 17 bytes).

* fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker

Adversarial isolated review (Step 6.3) found 2 HIGH findings that block
the PR and 3 mediums. All addressed:

H1 (HIGH): description containing the markdown 3-backtick fence would
terminate the ledger's JSON code block early inside JSON.stringify output
(JSON doesn't escape backticks), corrupting the file and bricking the
next parse. Fix: use a 4-backtick fence (json ... ) which
JSON.stringify cannot produce on its own, AND validate that no entry
text field contains a 4-backtick run (reject at append time with new
WINDOWS_INVALID_TEXT reason code). Locked by a regression test.

H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger',
silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate
would then pass on an unreadable ledger — the precise vector the
workflow doc claims is impossible. Fix: only ENOENT returns null;
every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the
gate blocks and the operator sees a real diagnostic. Locked by a
regression test that chmod 000s a ledger with open_count=1 and
asserts the result is never a false-green 0.

M1: writeLedgerAtomic left an orphaned .tmp file on rename failure.
Wrapped renameWithRetry in try/catch with best-effort unlink.

M2: validateLine silently coerced 'abc' → NaN → null, hiding type
drift. Removed the line === 0 special case (was undocumented) and
made the error message match the strict check. Now any non-positive-
integer line value throws, including strings.

M3: tests/broken-windows.test.cjs (with its fast-check property test)
was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate
src/broken-windows.cts but no test would catch the mutations,
producing false surviving-mutant scores. Added to the list.

L1 (dead throw e after error()), L7 (line boundary tests, H1/H2
regression tests, 4-backtick CLI test) also addressed.

* docs(#1950): inline concurrency + busy-wait notes (review L2+L3)

* fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test

gsd-test v4 caught two issues:
- goldens I regenerated earlier (commit 526682084) predated the L1
  routeWindows catch-block cleanup (commit dd844d565). Regenerated
  via 'npm run gen:golden' against current HEAD so the install
  parity hash for gsd-tools.cjs matches.
- 'append --line boundary' test expected --line 0 to succeed with
  null entry.line, but the M2 fix correctly rejects 0 (lines are
  1-indexed; 0 is not a valid source line). Updated the boundary
  test to assert --line 0 fails alongside -1 and 'abc'.

* chore(#1950): regen goldens after rebase onto next

* chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase)

* chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md

* fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization)

CodeQL flagged the markdown-table cell escaper:
  String(s ?? '').replace(/\|/g, '\\|')
— it escapes pipe but not backslash first. A description containing '\|'
would render as '\\|' which markdown parses as 'literal backslash' +
'cell separator', splitting the column.

Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|).
Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped
pipe), which markdown renders as a single '\|' inside the cell. The JSON
code block (the parse source-of-truth) was already correctly escaped via
JSON.stringify; only the display-only table was affected.

Locked by a regression test that:
1. Verifies the JSON block reparses with the description intact.
2. Walks the rendered table row counting unescaped pipes — must be
   exactly 11 (the row separators for 10 cells), proving no in-cell
   pipe added a split.
2026-07-19 20:24:21 -04:00
Behruz Nassre Esfahani
d04e287fa9 fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases

detectApiIntegration fired on any integration verb co-occurring anywhere on a
line with any API noun, treated / as a word boundary (so a first-party Next.js
src/app/api/... route path matched the noun "api"), and read any capitalized
word before API/SDK/REST/GraphQL as a service name behind a fixed stopword
denylist (so threat-model prose like "Resolver-only API" fired). Because the
verify:pre seal gate is BLOCKING, a phase touching no external API could not
reach UAT without fabricating a coverage matrix.

The compound rule now requires the verb and noun to share one clause (sentence
punctuation and table-cell walls end a clause) within a bounded word gap.
Non-prose spans are excluded before matching: fenced code (already), inline
code spans (new stripInlineCode in the markdown-sectionizer seam), and
path-shaped tokens. The <Service> API surface rule requires proper-noun
position — a clause-initial capitalized word is ordinary English and needs
dependency evidence (URL / package reference) on the same line — and rejects
compound modifiers ("Resolver-only", lowercase after the hyphen).

A phase that integrates no external API now has a first-class, reasoned way to
say so: a COVERAGE.md containing "No external API integration: <reason>"
satisfies the gate (declaration + rows is contradictory and blocks). The
true-positive path is pinned by regression tests: every default-vocabulary
positive still fires, including the widest word-gap pairing and the
surface-rule-only shape.

Fixes #2365

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2365): tighten api-coverage detector per Codex review (round 2)

Applies the Codex review findings on the initial #2365 fix:

- S-1: a COVERAGE.md "no external API integration" declaration is the human
  override for a fallible detector, so it must PASS even when detection still
  fires — but the contradiction is now SURFACED in the gate output (overridden
  signal count + terms) instead of passing silently.
- S-2: verb/noun pairing is now a term-group nearest-pair merge walk over
  precomputed word ordinals (computeWordStarts / minWordGap), not a match×match
  cross product — a hostile line repeating one pair thousands of times stays
  linear instead of going quadratic.
- FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`)
  are kept as noun/dependency evidence rather than being fully masked, so a
  genuine dependency reference inside code ticks still corroborates.
- C-1: the <Service> API surface rule now scans every candidate in every
  clause; a rejected first candidate no longer shadows a later genuine service.
- Cross-clause binding: a verb may bind a noun in the immediately following
  clause only when its own clause names a service object, within a tight gap —
  admits "Integrate Stripe, exposing its endpoints …" without re-admitting the
  unrelated-clauses false-positive class.
- Internal-descriptor negative evidence ("internal Payments API",
  "the internal endpoint") never pairs; URL/scheme matching generalized beyond
  http(s).

All 5 acceptance criteria still hold: the three reported false positives are
clean and "integrate the Stripe API" still fires. Built .cjs committed
alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) +
lint:regression-names clean; affected suites 256/256 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3)

Codex's second-round review found the round-2 tightening had over-corrected into
FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal
gate silently let through (the catastrophic class, since a missed API surface is
worse than a dismissable false positive). Retuned the detector to be explicitly
fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external
API integration" declaration dismiss the residual false positives.

Fail-open false negatives fixed (all now detect):
- F1 clause-initial `<Service> API` with a plain follower ("Stripe API for
  payment processing") — dropped the follower-allowlist / corroboration gate on
  clause-initial surfaces; a service that is not a stopword, descriptor, or
  compound modifier is a real name from any clause position.
- F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an
  alphabetic final label now contributes its API nouns; a first-party route
  path (no dotted host) still does not.
- F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the
  compound path no longer filters nouns on "internal"/"first-party" (Codex: the
  qualifier can describe the vendor's own API, not the consuming project's).
- F4 long single integration clause — removed the word-gap cap entirely: it
  could not separate a 21-word genuine clause from an 18-word internal one, so
  the clause boundary is now the whole relationship test.
- F5 lowercase cross-clause service — cross-clause binding no longer requires a
  capitalized "service object".

New false positives fixed (all now clean):
- F6 a URL token that swallowed a trailing clause comma, merging two clauses —
  trailing clause punctuation is kept literal so the split survives.
- F7 a capitalized internal component authorizing cross-clause binding — the new
  gate requires a dependent elaboration, not a new coordinate clause opened by a
  conjunction ("…, then document…").
- F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol
  and locality descriptors are rejected in the `<Service>` position.

- Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick
  runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms),
  semantics preserved (148 sectionizer tests unchanged).

Net simplification: the fail-closed model removed the round-2 minWordGap /
groupByTerm / follower / corroboration machinery (350 insertions vs 445
deletions across the touched files). Under fail-closed, three round-2 negative
tests now correctly detect (integration verb + "internal"-qualified noun, and
the distant-same-clause case); none were trek-e acceptance FPs.

Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown-
parsing) + lint:regression-names clean; all 8 review findings reproduced as
regression tests, both directions. Built .cjs committed alongside the .cts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4)

Codex's round-3 adversarial review found the fail-closed retune had introduced
new holes in both directions. Resolved:

Fail-open false negatives (now detect):
- External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an
  integration surface and contributes an endpoint noun even when the host names
  no vocabulary word. A bare domain link with no path ("https://example.com")
  stays a non-signal, so "Integrate … from example.com, document …" is still
  clean.
- Locality qualification ("internal", "private") no longer leaks across a
  sentence or clause boundary: only plain spaces may separate the descriptor
  from the service, so "The cache is private. Stripe API …" now detects.
- Cross-clause binding: the fragile head-word cap (which could not tell a
  genuine "Connect … to Stripe payments, exposing its endpoints" from an
  unrelated "Integrate … from URL, document …" — both 4 words after the verb)
  is replaced by a participial-continuation rule: a verb binds a noun in the
  next clause only when that clause begins with an "-ing" elaboration. This
  fixes the 4-word-head false negative AND the false positive below at once.

False positives (now clean):
- Cross-clause no longer binds a finite continuation regardless of separator:
  "Wire the settings form. Document endpoint props." / "…; document …" /
  "…, document …" are separate actions, not elaborations.

Perf (quadratic → linear):
- The trailing-punctuation peel is a backward char scan instead of an
  unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms).
- SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile
  "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms).

Consumer fail-open (blocking gate):
- readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is
  unreadable. On a read error the gate BLOCKS ("could not read the phase
  scope …") instead of silently certifying no-integration from partial scope —
  an unreadable plan could be the one describing the integration.

Documented fail-closed tradeoffs, now pinned with tests so they are not
"fixed" back into a fail-open: a clause-initial capitalized common word before
"API" ("Payment API", "Search API") reads as a service name; a long clause
pairs a verb with a distant noun; and a CommonMark inline code span that wraps
a newline is matched within-line only. Codex judged these acceptable because
the COVERAGE.md declaration is a cheap override.

One documented limitation remains out of scope: "Integrate Stripe, and
authenticate requests with its API" (a coordinate finite clause whose noun
refers back by pronoun) needs coreference resolution, beyond a lexical detector.

Verified: 379/379 affected + command-router tests pass (+14 new regression
tests covering every round-3 finding, both directions); tsc + eslint
(no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5)

Round-4 review confirmed the detector's two most complex features generate
findings in both directions no matter how they are tuned, because they need a
vendor dictionary + coreference the issue rules out in principle. Per the
operator's "ship the robust core" decision, both are removed and their gaps are
documented rather than chased further:

- Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused
  a fail-open on finite continuations ("Integrate Stripe; use its OAuth
  endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing
  endpoint terminology…" — wrongly fired). Detection is now same-clause only.
- URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint
  fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo
  link) and recreated routine UI-phase false positives. An external URL is
  evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1").

Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future
maintainer does not re-add the heuristics that caused the false positives above:
a service named only in a clause separate from its API noun, and a bare external
host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md
declaration and rare in real phase prose ("integrate the X API").

Also fixed from the round-4 review:
- Qualification now survives markdown emphasis ("The **internal** Payments API"
  stays clean) while still not crossing a sentence/clause boundary.
- readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the
  phase directory or reading the roadmap fallback — not only per-plan-file
  failures; a missing directory/section remains a legitimate no-op. The
  declaration-override path surfaces scope_read_error so an incomplete-scope
  override stays visible.
- SERVICE_SURFACE_API_RE length-bound comment no longer overclaims.

Net: the detector is same-clause verb+noun + `<Service> API` surface, with
path/code/inline masking and a fail-closed posture. All five acceptance criteria
hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) +
lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review)

The round-5 sanity review confirmed the detector simplification is sound (all
acceptance positives fire, all required negatives clean) and flagged one real
blocker plus a nit:

- Blocker: readPhaseScope's roadmap fallback could still silently pass an
  UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which
  returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent",
  no exception reached isRealReadFailure, and the blocking gate certified empty
  scope. Fixed at the source: read the roadmap directly and honor the function's
  OWN documented contract — null only on ENOENT (genuinely absent), otherwise
  throw. Both existing callers already wrap it in try/catch expecting that throw,
  and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by
  a new e2e test (unreadable roadmap fallback → block).

- Nit: the detectApiIntegration JSDoc still described the removed cross-clause
  participial binding and "every external hostname counts" — corrected to the
  actual same-clause-only behavior and the names-a-vocab-word URL rule.

Verified: full unit suite green; tsc + eslint + lint:regression-names clean.
Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): backfill changeset PR number (#2397)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): sync generated capability-registry + recapture install goldens

CI surfaced two generated-artifact staleness issues (all failing test shards +
lint-tests traced to these, not to a logic defect):

- gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the
  ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API
  integration" declaration section) but did not regenerate the registry, which
  embeds an inline copy of that fragment. Regenerated via
  `gen-capability-registry.cjs --write` — the diff is exactly the fragment text
  sync. Fixes `lint:generated-sync` and the "committed registry is in sync" +
  "registry integration" tests.

- The 18 golden-install-parity fixtures were stale by exactly one hash line each
  — `gsd-core/references/api-coverage.md`, which this PR edits and which is a
  hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that
  single hash per runtime and nothing else. Fixes the `golden parity — *` tests.

No source or behavior change — generated artifacts only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): flip representative-corpus manifest to assert the fixed behavior

The #2371 representative corpus (merged into next after this branch was cut) is a
known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test
fails loudly the moment #2365 is fixed, at which point — per its own contract in
representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the
assertion checks expectedDetected instead.

This is that moment. Removed currentBuggyOutput from the three detector fixtures
(nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now
asserts detected:false, which the fail-closed same-clause detector satisfies.
Notes updated to describe the fix rather than the bug. The #2366 matrix corpus
is left untouched — that tripwire belongs to its own PR (#2374).

Corpus test: 7/7 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): skip chmod-000 fail-closed e2e tests on Windows

The three fail-closed gate tests induce an unreadable plan / directory / roadmap
with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync
still succeeds, so the gate never reaches the read-error path and the assertion
fails on the windows-latest CI leg. The fail-closed LOGIC is platform-
independent (readError → block) and is fully exercised on the macOS/Linux legs;
only the method of inducing EACCES is POSIX-specific. Guard the three tests to
skip on win32 as well as root, mirroring golden-install-parity's win32 skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds

Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged;
this closes the test/process-hygiene findings.

Major:
- CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix
  relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate).
- Replaced the banned wall-clock assertion in the "hostile repeated-term line"
  test (Clock Seams rule — no elapsed-time asserts) with a deterministic
  signal-count assertion, which also directly verifies the term-dedup that keeps
  pairing linear (one signal for a 10k-pair line, not thousands).
- Rewrote the three fail-closed read-failure tests: instead of chmod 0o000
  (a no-op under root / on Windows, the pattern the repo's IO-failure convention
  avoids) they now exercise the newly-exported `readPhaseScope` in-process and
  inject the failure by monkeypatching fs.readFileSync/readdirSync to throw,
  restoring in finally. Deterministic and platform-independent (no skip needed),
  and they add the ENOENT-is-absence case that the chmod tests couldn't express.

Minor:
- Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40}
  service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN
  (200) on the declaration reason.
- Added a fast-check property that fuzzes the tokenizer / clause splitter /
  masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial
  tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector
  is total (never throws), shape-stable, holds detected <=> signals, and is
  deterministic.

readPhaseScope is exported for the in-process tests. Verified: 125 detector +
19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) +
lint-regression-test-names + lint-test-file-count clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 14:59:25 -04:00
Tom Boucher
d0bacc2517 fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout

Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10
hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates
exited 127 ("command not found") on such hosts, and the gates — which only
distinguish 0/124/other — misreported a passing build or test as a FAILURE.

Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in
gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU
`timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES,
128+signum on signal), inherits stdio so pipes/redirects work, and reaps the
whole process group so a watch-mode runner cannot outlive its budget. Runs
before gsd-tools' flag parsing so the wrapped argv stays opaque.

Hardened per adversarial review:
- On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant
  that traps SIGTERM was otherwise orphaned holding stdout, hanging captured
  gates (the exact watch-mode hang the feature prevents).
- Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it.
- Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the
  timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout).
- Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command
  position so prose "timeout 30 seconds" no longer false-positives.

Resolution lives once in the CLI; all 10 sites call the shared verb. A parity
guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if
a bare `timeout`/`gtimeout` execution reappears (the portable `command -v
timeout` probe form is intentionally allowed). Also fixes the identical bug in
the zh-CN checkpoints translation, updates the tests that asserted the old
strings, trims a redundant phrase in gsd-verifier.md to keep it under its size
hard cap, and refreshes the size baselines + golden install-parity fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2351): add changeset (#2426)

* chore: regenerate golden/size baseline after rebase onto next

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:37:15 -04:00
Tom Boucher
cd6665d73b fix(#2406): stop Codex config.toml from double-registering agent roles (#2432)
* fix(#2406): stop Codex config.toml from double-registering agent roles

generateCodexConfigBlock emitted an [agents.<name>] role table per agent
pointing config_file back at the standalone agents/<name>.toml Codex
already auto-discovers, so every install declared each role twice in
one config layer and Codex logged a duplicate-role warning per agent.
Remove the redundant role-table loop; the standalone per-agent TOML is
now the sole canonical registration source. The existing marker-truncate
and leaked-section stripping in mergeCodexConfig already clean up legacy
[agents.gsd-*] tables from prior installs, so updates converge to zero
duplicates without any new migration path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2406): regenerate fixtures + lint gate-prep

* fix(#2406): repair failing tests after gate verification

* docs(#2406): add changeset for Codex duplicate agent-role fix

Adds the missing .changeset/*.md fragment for the Codex config.toml
double-registration fix, closing the PR-gate finding from review.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2406): backfill changeset pr (#2432)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 12:53:22 -04:00
Tom Boucher
a7d83dc234 fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix

phase.add now returns a `warning` field when a description reads as
goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped,
instead of silently writing the whole paragraph verbatim as the
`### Phase N:` header. The CLI still creates the phase as-is (the
strict two-layer slash-vs-CLI interface is unchanged); the warning
just surfaces the gap.

Also clarifies six doc sites (command argument hints, workflow
detection steps, and how-to/reference docs) that described the
phase-number argument as "auto-detecting" the next unplanned phase --
that detection is an orchestrating-workflow/LLM step reading
ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase`
field), not a `gsd-tools.cjs` CLI feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2390): regenerate fixtures + lint gate-prep

* fix(#2390): repair failing tests after gate verification

* chore(#2390): add changeset (#2425)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 07:53:04 -04:00
Tom Boucher
1720aacf0c feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element

Red phase for issue #1949 (Design by Contract: <precondition> element
asserted before task execution). Tests assert:

- docs/reference/plan-md.md documents the new <precondition> element
- agents/gsd-planner.md @-references planner-preconditions.md and stays
  under the 49152-char cap (progressive-disclosure requirement)
- gsd-core/references/planner-preconditions.md exists and documents the
  three emission cases mandated by the issue (user_setup / prior-phase
  artifact / env-var) and the contract triad mapping
- agents/gsd-executor.md asserts <precondition> before task execution
  and routes unmet preconditions through existing checkpoint machinery
- cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both
  with and without <precondition> — the additive-validation guarantee
- Parity assertion: plan-md.md and planner-preconditions.md agree on the
  canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard)

Most prose-contract assertions are Red until the implementation lands.
The behavioral validator assertions pass immediately (regression guards
proving the validator already accepts unknown optional tags).

* feat(#1949): <precondition> task element — Design by Contract

Add an optional <precondition> element to <task> in PLAN.md (issue #1949,
The Pragmatic Programmer Topic 23). The front-of-task side of the plan
contract — preconditions (before) ↔ postconditions (<verify>/<done>/
<acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the
whole plan). Together with the tracer-bullet proposal (#1945), this closes
both ends of the 'outrunning your headlights' failure mode for an
autonomous AI executor.

Acceptance criteria met:
- <precondition> is an optional element on <task>; plans that omit it
  validate unchanged (cmdVerifyPlanStructure checks for presence of
  required tags, does not reject unknown optional tags).
- gsd-executor evaluates the precondition before any other task work.
  Unmet halts execution with a checkpoint:human-verify and no partial
  commit; met or absent produces no visible change to execution flow.
  Unmet is never auto-approved under AUTO_CFG=true — a missing
  prerequisite is a fact the executor cannot establish on its own.
- gsd-planner emits <precondition> in exactly the three cases the issue
  mandates: user_setup consumption, prior-phase artifact dependency, and
  env-var/runtime-config dependency.
- Tests cover met, unmet, and absent preconditions plus the additive-
  validator guarantee.

Files:
- gsd-core/references/planner-preconditions.md (NEW): full emission
  rules, the three cases with worked examples, format guidance,
  anti-patterns, the contract triad mapping, and the executor assertion
  contract. Progressive disclosure.
- agents/gsd-planner.md: slim <precondition> note in Task Anatomy with
  @-reference to the new file. To stay under the 49152-char agent-file
  cap (27-char headroom before this change), the inline
  <comment_text_discipline> and <region_scoped_negative_gate> summaries
  are compressed to one-line pointers — their full rules already live in
  planner-antipatterns.md, so no content is lost.
- agents/gsd-executor.md: new step 0 'Precondition check' in the
  execute_tasks loop, before the type dispatch, routing unmet through
  checkpoint_return_format.
- docs/reference/plan-md.md: new Preconditions section in the schema
  reference, with the canonical example and the three emission cases.
- CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet.
- docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new
  references/planner-preconditions.md (regen via gen-inventory-manifest).
- tests/precondition-element.test.cjs: failing-first tests covering
  schema docs, planner emission contract, executor assertion contract,
  reference-file presence + the three cases, behavioral additive-
  validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX-
  DIVERGENCE guard).
- .changeset/quick-hawks-bark.md: Added fragment.

Companion to #1945 (tracer bullets).

* chore(#1949): regen agent-size baseline + install-tree goldens

Documented baseline regenerations required by the feat(#1949) prose changes
(RULESET.AGENT_SIZE_BUDGET + golden-install-parity):

- npm run size:baseline — locks in the new gsd-executor.md size (+1050
  bytes: the precondition-check step 0 block). gsd-planner.md is net
  smaller (-142 bytes: compressed two inline summary blocks whose full
  rules already lived in planner-antipatterns.md to make room for the
  slim <precondition> pointer). No hard-cap breach.
- npm run gen:golden — pick up the new references/planner-preconditions.md
  + the two changed agent files across all 18 runtime install trees.

Both regens are CI-mandated after intentional agent/reference changes;
see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in
tests/golden-install-parity.test.cjs.

* fix(#1949): bound <precondition> checks to read-only (security review)

Apply the security-review finding (LOW, isolated /security-review subagent):
the executor's 'run the cheapest check' phrasing for a plan-author-controlled
prose line was broader than ideal — a hostile plan author could craft a
<precondition> whose 'cheapest check' is side-effecting (curl to an attacker
host under the guise of verification, rm -rf before checking, secret emission).

The risk is inherited from GSD's existing plan-trust model (<verify>, <action>,
<done> already direct the executor to run arbitrary shell), so <precondition>
does not materially expand it. But the new prose actively directs execution
('run the check') rather than passively consuming the element, so the bound
is worth making explicit.

Tightened across all four surfaces that describe the check shape:
- agents/gsd-executor.md step 0: 'Verify with read-only checks only — file
  existence, env var presence (no value output), idempotent GET /health-style
  pings. Do NOT run commands with side effects (writes, network POSTs, secret
  emission) as the check; if a side-effecting check seems required, halt and
  surface via checkpoint instead.'
- gsd-core/references/planner-preconditions.md Format section: same bound,
  plus the halt-and-surface escape hatch.
- docs/reference/plan-md.md Preconditions section: mirrored.
- CONTEXT.md Precondition glossary entry: mirrored.

Regenerated agent-size baseline (executor grew 46186 -> 46440; still under
the 49152 cap) and install-tree goldens.

* chore(#1949): backfill changeset pr number 2422

Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7:
backfill the placeholder pr:0 with the real PR number immediately after
gh pr create returns. Avoids the fail_invalid_fragment gate.

* fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456)

CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires
every // allow-test-rule: exemption on a NEW test file to carry an issue
reference (#NNN or URL). My earlier push omitted it.

Local 'npm run lint' (eslint) does NOT run this check — only 'npm run
lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs
lint:ci; a local pass is not the gate.' I should have run lint:ci before
pushing; correcting now.

Pattern matches the companion feature's test file:
tests/tracer-bullet.test.cjs:1  // allow-test-rule: source-text-is-the-product [#1945]
2026-07-19 07:52:36 -04:00
Tom Boucher
8d2f8bcb23 fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found

Adds requirements.ready-ids (execute-plan.md's update_requirements step)
so a requirement ID declared by multiple plans in a phase only marks
Complete once every declaring plan has produced a SUMMARY.md, and
requirements.revert-phase (execute-phase.md's gaps_found branch) so a
gaps_found verdict reverts the phase's own prematurely-Complete IDs
before the gap report renders. Single-plan IDs still mark immediately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2388): regenerate fixtures + lint gate-prep

* fix(#2388): repair failing tests after gate verification

* chore(#2388): add changeset (#2424)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 07:52:17 -04:00
Tom Boucher
b2f4aa9435 docs(#2420): clean stale get-shit-done/ path refs in translated docs (#2421)
After the package/repo rename in #604, the English docs were updated to
use gsd-core/... paths, but the four translated doc trees (ja-JP, zh-CN,
ko-KR, pt-BR) and .changeset/README.md were never updated and still
referenced the pre-rename get-shit-done/ runtime directory, which no
longer exists.

This commit brings the translations in line with the English docs:

  - docs/{ja-JP,zh-CN,ko-KR,pt-BR}/**/*.md (57 files):
      get-shit-done/ -> gsd-core/  (path references)
      #references-get-shit-donereferencesmd -> #references-gsd-corereferencesmd
                                            (anchor in INVENTORY -> ARCHITECTURE links)
  - .changeset/README.md:9 issue URL:
      open-gsd/get-shit-done-redux -> open-gsd/gsd-core

Legacy references intentionally preserved (historical record):
  - CHANGELOG.md, .changeset/archived/*, docs/RELEASE-NOTES-LEGACY.md
  - docs/cleanup-get-shit-done-cc.md, docs/adr/*, docs/research/*
  - docs/{ja-JP,ko-KR}/superpowers/plans/2026-03-18-* (developer's local paths)
  - docs/{INVENTORY,README,FEATURES,installer-migrations}.md (rename-history
    descriptions, some tagged <!-- gsd-allow-legacy-name -->)
  - Code/tests implementing or testing legacy-cleanup logic
    (bin/install.js, gsd-core/bin/lib/legacy-cleanup.cjs,
    scripts/lint-legacy-dir-name.cjs, migration sources/tests)

No source code changes — documentation only.

Fixes #2420
2026-07-18 23:11:15 -04:00
Tom Boucher
dd5a2211c9 enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests

Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests:
semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches
same-root-cause/different-wording cases), indexing resolved sessions at archive,
graceful degradation to keyword matching when MemPalace is absent,
knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 /
Matching Logic is semantic-first (the stale 'keyword overlap, not semantic
similarity' claim must go), and no new embedding/vector infra (reuse MemPalace).

Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation,
and the archive indexing step do not yet exist.

* feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback)

Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic
recall: at Phase 0 the debugger queries MemPalace with the current symptoms
and surfaces the top-k meaning-similar prior resolutions, catching the
same-root-cause/different-wording cases keyword overlap missed (the self-noted
'keyword overlap, not semantic similarity' limitation). Resolved sessions are
indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence
guard). knowledge-base.md remains the durable plain-text source of truth; when
MemPalace is absent the debugger falls back to keyword-overlap matching
(logged, never a silent skip). No new embedding/vector infrastructure —
MemPalace is reused.

Size-neutral agent edits: the Matching Logic section reframed (keyword-only ->
semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets
consolidated into one semantic-first bullet; one MemPalace-indexing step added
at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in
gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail)

- HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has
  no MCP tools. Added an Invocation section naming the Bash CLI
  (mempalace search --wing <wing>) + MCP-when-registered + wing resolution
  (config.mempalace.wing -> project_code -> project dir), matching every other
  MemPalace integration. Without this the feature silently degraded to keyword
  matching even when MemPalace was present.
- MEDIUM (security x2): index the agent-authored Resolution summary
  (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms —
  excludes attacker-controlled prose from the cross-session index AND reduces
  secret/PII leakage. Redact secret-shaped values before indexing. Stated the
  write order (KB append + commit MUST succeed before indexing).
- LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback;
  added a test asserting the fallback mechanics survived the Phase 0
  consolidation (Error patterns field, 2+ token overlap, identifiers,
  case-insensitive).

* chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197)

* chore(#1964): backfill changeset pr number (PR #2416)
2026-07-18 19:01:04 -04:00
Tom Boucher
c67f301867 feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests

Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless
5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent
error as 'why was that possible?'), the 'why wasn't this caught?' question,
the recurrence-guard taxonomy (regression test / assertion / lint rule / KB
pattern), the KB-entry why_not_caught + recurrence_guard fields with backward
compat, the session-manager prevention summary line, and the Zawinski
scope-boundary (a block, not a subsystem).

Failing-first: reference, archive_session edit, KB schema extension, and
session-manager summary do not yet exist.

* feat(#1963): emit blameless-postmortem Prevention block at resolution

Epic #1957 Phase 3B. At archive_session the debugger now produces a
Prevention block with three blame-free components: a branching 5-Whys causal
chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts
'why was that possible?', never blame), a 'why wasn't this caught?' answer
naming the missed gate (test/typecheck/lint/review/verify), and a concrete
recurrence guard (regression test / assertion / lint rule / KB pattern).

The knowledge-base entry gains two structured fields (why_not_caught +
recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not
just the prior fix. Additive: old entries without the fields still load. The
session-manager compact summary surfaces a one-line prevention summary.

Full rules extracted to gsd-core/references/debugger-prevention.md (slim
archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test)

- CRITICAL: the archive_session KB append template omitted Why not caught +
  Recurrence guard (only the Entry Format had them) — the feature's core
  deliverable silently did not happen. Added both fields to the append template
  the agent actually follows (nearest-instruction wins).
- HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were
  dead data. Extended the Phase 0 Evidence line to consume why_not_caught +
  recurrence_guard when present (absent on old entries — backward compat holds).
- MEDIUM: added a cross-section parity test (every Entry-Format field must also
  appear in the append template — the guard that would have caught the
  Critical) + a Phase-0-consumption assertion.
- MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses
  reasoning_checkpoint.candidate_causes across the four categories.
- MEDIUM: recurrence-guard taxonomy gains type refinement + config-default
  change; LOW: added 'build' gate to both surfaces for parity.
- NIT: compact-summary fallback shape ('no gate existed'); verify the guard
  artifact exists before recording it.

* test(#1963): anchor Phase-0 consumption test on the specific heading

The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file
(in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop.
Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars.

* chore(#1963): backfill changeset pr number (PR #2410)
2026-07-18 17:38:54 -04:00
Tom Boucher
36a311c5bb enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests

Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking
(fast-check/Hypothesis, minimized seed, manual-minimization degradation), the
four oracle types (specified/derived/metamorphic/implicit with implicit flagged
weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the
equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in
(minimized seed + real oracle => the mutation guardrail bites).

Failing-first: reference, agent cross-refs, and template field do not yet exist.

* feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries)

Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First
Debugging (oracle classification + boundary neighbors):
- Shrinking: wrap an input-space failing input in a property (fast-check JS/TS,
  Hypothesis Python) and store the MINIMIZED counterexample as the regression
  seed; degrade to manual minimization when no PBT framework is present.
- Oracle classification: state specified / derived (contract/model) /
  metamorphic / implicit (crash, weakest) before writing the assertion; record
  under Resolution.oracle_type; never default to implicit silently.
- Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed
  defect's equivalence class.

Together they turn the regression test into a root-cause check — what the Phase
1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/
debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline +
install-parity goldens + AGENTS.md + DEBUG template updated.

* fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple)

- HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade-
  to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) —
  the gauntlet violation the sibling references already honored.
- Medium: test-provenance caveat (the failing input often comes from the bug
  report — author the generator from a sanitized description, cross-ref
  debugger-fix-acceptance.md).
- Medium: oracle scope note — the 4 types cover deterministic bugs; non-
  deterministic failures re-route to stability-stress per bug-taxonomy.
- Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient;
  boundary neighbors close the adjacent-input escape; the sufficient triple is
  seed+oracle+neighbors.
- Low: preserve the original noisy repro as a secondary reference; operationalize
  'equivalence class' (the predicate the fix draws). Nit: degradation reworded.

* chore(#1962): backfill changeset pr number (PR #2409)

---------

Co-authored-by: sim <sim@local>
2026-07-18 15:46:42 -04:00
Tom Boucher
6baa2a8182 feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests

Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy
classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect,
Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/
deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append)
plus a routing-table specification object pinning the documented decisions
(SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam).

Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist.

* feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger

Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the
failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the
investigation technique via an explicit class->technique table (Kernighan: no
opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect;
Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical
sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai
ranking — the load-bearing 1B/2B seam); Concurrency -> the
atomicity/order/deadlock checklist first.

Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by
situation' table into a class-routed table; the 11 techniques remain as routed
targets. bug_class recorded in Current Focus (DEBUG template); common-bug-
patterns catalog cross-referenced to the taxonomy.

Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY
+ manifest + agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding)

- HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed
  'Phase 1.25' (matches the agent + SBFL reference).
- HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards,
  Differential, Comment-out, Follow-the-indirection) were orphaned by the
  situation-table reframe. Added a 'General (any class, situation-cued)'
  lane to BOTH the reference routing table and the agent's Technique
  Selection table that re-homes them — supersede-not-append now holds.
- MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before
  Phase 1.75 classification), so reframed the table column from 'Do NOT use'
  to 'Revoke if already run' + an explicit 'retroactive revocation, not
  proactive skip' note stating the ordering honestly.
- MEDIUM: contract tests are now row-scoped (parse the table by class, assert
  per-row) instead of presence-only; added a guard that the previously-
  orphaned techniques now have a General-lane route.
- LOW: pinned the canonical bug_class value form (lowercase-kebab:
  bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case).
- NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling
  timeouts) per the unbounded-subprocess gauntlet.

* chore(#1961): backfill changeset pr number (PR #2407)
2026-07-18 14:42:29 -04:00
Tom Boucher
f8b16d1874 enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests

Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone
>=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning
checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause
note, DEBUG template) plus behavioral schema-invariant checks on two fixtures:
two contributing causes (AND-gate yes) -> both recorded; single-cause
(AND-gate no) -> one root_cause, identical to today.

Failing-first: reference, agent edits, and template note do not yet exist.

* feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger

Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing
root_cause, the debugger enumerates candidate causes across >=2 Ishikawa
categories (code/config/environment/data) and explicitly answers an AND-gate
question. When the AND-gate fires, every contributing cause is recorded, so a
multi-cause fix no longer recurs via the unaddressed second cause.
Resolution.root_cause may hold one OR a small set (additive; single-cause
sessions are byte-identical to today). The Structured Reasoning Checkpoint gains
candidate_causes + and_gate fields; debugger-philosophy.md adds the
single-cause-bias trap.

Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim
Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated.

* fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples)

- Reference: the collapse rule now enforces AND-gate self-consistency —
  and_gate=yes with a single confirmed cause is flagged as incomplete
  (return to Phase 3); a race/timing note clarifies such bugs bridge
  categories; the 'byte-identical' backward-compat claim narrowed to
  'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every
  session'.
- DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface
  drift the reviewer flagged); new debug-session-management parity test pins
  the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe).
- Scalar-assuming consumers of set-valued root_cause updated: session-manager
  compact summaries (319/332), diagnose-only return (1062), archive entry
  (1216), ROOT CAUSE FOUND return (1322).
- Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the
  fixture describe block honestly as a schema-invariant specification.
- Phase 2 bullet phrasing clarified ('at hypothesis formation, before the
  Phase 4 commit').

* test(#1960): parity regex accepts word-form count ('seven-field' or '7-field')

* test(#1960): parity regex counts array-valued YAML keys (no inline value)

* chore(#1960): backfill changeset pr number (PR #2405)
2026-07-18 13:42:58 -04:00
Tom Boucher
13d181aedf fix(#2349): exclude status: superseded plans from phase completion counts (#2404)
Adds a status: superseded plan-frontmatter marker that scanPhasePlans excludes from both plan and summary counts, so a phase with a deliberately-unexecuted plan no longer reads incomplete forever (the plan-level analogue of #1514). Includes all-superseded completion handling and a bounded, symlink-safe frontmatter read. Fixes #2349.
2026-07-18 08:08:03 -04:00
Tom Boucher
56a5c6404c feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests

Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai
formula documented, Tarantula fallback, top-N seeding, no-coverage skip
logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai
formula-correctness section: bound [0,1], max-score invariant, a known-fault
fixture proving the fault ranks #1 (criterion 2), clean degradation on
zero failing tests, and two fast-check properties.

Failing-first: reference file and agent routing do not yet exist.

* feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger

Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists
(>=1 failing AND >=1 passing test), the debugger computes an Ochiai
suspiciousness ranking over the coverage spectrum and seeds the top-N
suspicious locations into Evidence as first-class hypothesis candidates,
narrowing the search space deterministically before LLM reasoning. Tarantula
documented as fallback. Degrades cleanly (logged, never silent) when there is
no test suite, no failing tests, or no per-test coverage, and is explicitly
not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy).

Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25
routing kept in the agent to respect the size cap). No new coverage framework
— reuses the project's existing test/coverage runner. INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* test(#1959): bound property generators to valid coverage counts

The [0,1] property generated failedExec independently of totalFailed, but
Ochiai's score is only bounded by 1 under the coverage invariant
failedExec <= totalFailed (a failing test that executed s is one of the
totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5)
make the formula correctly return >1. Bound failedExec by totalFailed via
fc.chain so the property tests the real domain. Also cleaned up the ranking
property (removed dead code).

* fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding)

- Replace vacuous ranking property (true-by-sort-construction) with a
  non-trivial monotonicity property: holding totalFailed + passedExec fixed,
  ochiai is non-decreasing in failedExec. An inverted formula would fail it.
- Add the missing 'no passing tests' degradation row (preconditions require
  >=1 passing test; Tarantula would divide by totalPassed=0).
- Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run,
  degrade-to-skip on timeout, never hang the debug session.
- Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not
  delete)' per Kernighan auditability.

* test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed)

* chore(#1959): backfill changeset pr number (PR #2403)
2026-07-18 07:55:28 -04:00
Tom Boucher
5e52350736 feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests

Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the
5-signal fix-acceptance guardrail contract (target test, mutation check,
no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful
degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file
recording, and subprocess bounding.

Failing-first: reference file and agent sections do not yet exist.

* feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger

Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test
(Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix
acceptance: target test, mutation check (Stryker), no-op/behavior-deleting
detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully
when Stryker or a test suite is absent (each skip logged, never a silent pass),
records per-signal results under Resolution.verification, and returns a
FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for
revise / accept-as-debt / abandon.

Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim
routing kept in the agent to respect the agent-size cap). Debug template +
INVENTORY + manifest + agent-size baseline + AGENTS.md updated.

* test(#1958): correct newline-tolerant assertion + regen install-parity goldens

The revert-and-reconfirm assertion collapsed whitespace before matching so
markdown line-wrapping does not break it. Regenerated the golden-install-parity
and install-tree fixtures (npm run gen:golden) to absorb the intentional
gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file
changes to the installed artifact tree.

* fix(#1958): tighten guardrail per orthogonal review

Addresses the isolated reviewer's findings:
- signal 5 now states its recorded-repro dependency and routes the no-repro
  case to the degradation row; revert mechanism specified (git stash / git
  revert -n); minimality flag tied to diff structure, not revert-ability.
- bounded-subprocesses section now bounds the git subprocess (5-30s) too,
  requires argv-array argument passing, and scopes Stryker to the driving
  regression test (a mutant killed only by a non-driving test is a finding).
- new test-provenance (security) clause: the driving test must be
  agent-authored; bug-report repro scripts are DATA, never executed verbatim.
- tightened 3 contract assertions to bind to specific clauses
  (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding).
- Goodhart framing softened to 'partially-independent'; DEBUG.md template
  verification field notes the nested map shape.

* chore(#1958): backfill changeset pr number (PR #2396)

* fix(#1958): add issue ref to allow-test-rule annotation (ADR-456)

CI lint-allow-test-rule-refs requires every allow-test-rule exemption to
carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation
lacked it; this adds (see #1958).
2026-07-18 00:39:28 -04:00
Tom Boucher
58028eaf56 fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point.

Closes #2341. Admin-merged (self-review bypass) with full green CI.
2026-07-17 15:20:19 -04:00
Tom Boucher
b302f53ee6 refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2)

Behavior-preserving relocation of the 706-line case 'capability': arm from
gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs
(sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are
rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is
now async (capability's install/upgrade/consent ops await the lifecycle); sync
routers (state/phase/…) pass through await unchanged. case 'capability':
removed; capability dispatches via HOST_COMMAND_ROUTERS.

Validated by the existing capability-lifecycle / -consent / -trust / -loader
test suites (no logic changed). Golden install-parity fixtures regenerated.

Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→
readHostVersion, capReadStrict dedup) deferred to a follow-up slice.

* fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row

The relocated capability arm references capabilityState (cmdCapabilityState,
resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both
module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep
scan missed. Added as sibling requires to capability-command-router.cjs. Also
adds the new cli module to docs/INVENTORY.md + regenerates the manifest.

* fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation

capHostVersion's VERSION/package.json paths were bin/-relative ('..' and
'..','..'); on relocation to bin/lib/ they resolved one level too deep,
so capHostVersion returned 0.0.0 and capability install failed the
engines.gsd gate (#1920). Added one more '..' to each (now resolves
gsd-core/VERSION and repo-root package.json correctly).

* test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd)

capability is async and does FS/config reads, so invoking it against the
unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the
invocation loop; capability is covered by the non-invoking registry-
ownership assertion + the dedicated capability-* test suites.

* chore: retrigger CI (no-changelog label now present)
2026-07-17 11:04:47 -04:00
Tom Boucher
67a9243cf1 chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants

The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.

Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:

- scripts/gen-adr-index.cjs generates the index between markers and validates
  the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
  successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.

Correct the lifecycle metadata the gate surfaced, without flipping any status:

- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
  Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
  the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
  capability system shipped and epic #857 is closed. Ratification is a
  maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
  the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: capture stderr via spawnSync; record ADR-0010 draft supersession

Two fixes surfaced by the first gsd-test run and by regenerating the index:

- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
  through the thrown error on non-zero exit. The `--write` path exits 0 while
  reporting outstanding violations on stderr, so the helper always saw ''.
  spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
  "earlier draft superseded by ADR-0011" while the file itself still said
  Proposed. Deriving the index from the files would have dropped that
  assertion and resurrected a superseded draft as a live decision, so it is
  recorded at its source, with the reciprocal Supersedes on ADR-0011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174

src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").

ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (11918dcc3), so the candidate can no longer resolve in any layout: a source
repo has no sdk/, and an install layout points it at ~/.claude/sdk/shared/, which
the installer never writes -- the original #3288 bug. It was dead weight implying
a package boundary this repo no longer has.

No test depends on it: the #3288 regression tests in tests/install.test.cjs write
their own synthetic old-path fixture and assert it throws.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: ratify nine shipped ADRs; record why ten others stay Proposed

The corpus carried 19 Proposed ADRs, most describing architecture that had
already shipped. A Proposed label on live architecture tells contributors and
agents the decision is an unbuilt idea -- the capability system and EoS were
both being misread that way.

Audited all 19 against the shipped tree and GitHub. Each candidate flip then had
to survive two independent reviewers instructed to refute it.

Ratified Proposed -> Accepted, each with a dated Ratification section carrying
the verified evidence (file:line, symbols, tests, issue state):

  857  capability system      894  declaration format   1244 capability ecosystem
  1577 injection boundary     1610 size-budget ratchet  1990 existing-code onboarding
  15   cross-AI convergence   22   plan-drift guard     0011 default reviewers

Held ten, each now carrying a "Why this is still Proposed" section naming the
blocker and its unblock condition, so the audit is not repeated:

  2264 its own headline acceptance criterion is unmet in the tree
  230  live branch protection contradicts the decided spec (1 approval, not 2)
  660  the namesake release/<version> re-cut is manual, not automated
  959  issue #2346 is approved and plans its graduation as its own ADR
  1213 the shipped writer's return shape differs from the decided interface
  443  the orchestrator override path has no live caller
  1143 / 1606 each states its own bar for acceptance; neither is met
  612 / 1671 legitimately open

Shipped code proved necessary but not sufficient: eight ADRs had every named
module, symbol, and test present with their epics closed, and still failed the
bar. That lesson is written into README.md's ratification procedure.

Also corrected ADR-857's "Supersedes (generalizes)" to "Subsumes": taken
literally it would have marked two live seams dead -- ADR-0011 (surface.cts:348)
and ADR-58 (runtime-artifact-install-plan.cts:82). Both keep Accepted status and
gain Subsumed-by pointers.

Index: Active 39->48, Proposed 19->10, Superseded/Legacy 7. 65 total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden gen-adr-index against hostile titles and non-ADR filenames (#2356)

Three findings from the pre-PR orthogonal security review, all confirmed:

- An ADR title containing the literal ADR-INDEX:END marker was emitted verbatim
  into its table cell, relocating the splice boundary so the NEXT --write
  spliced against the wrong marker and truncated README.md. Titles now render
  through cellText(), which escapes pipes and angle brackets -- making an HTML
  comment (and any other HTML) unformable from ADR-authored text.
- A docs/adr/*.md without a numeric prefix crashed on match(...)[1] of null.
  Such a file is also invisible to the index -- the very failure this gate
  exists to prevent -- so it is now reported as a naming-convention violation
  naming the file and the fix.
- Tests leaked their mkdtemp dirs. They now use helpers.createTempDir/cleanup
  via t.after(); helpers.cleanup carries the Windows-EBUSY retry budget that a
  raw fs.rmSync lacks (caught by local/no-raw-rmsync-in-tests).

Adds five regression tests: marker hijack, HTML injection, pipe cell-break,
non-conforming filename, and splice stability across repeated writes.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: close two gate false-passes; read ## Supersedes sections (#2356)

Second round of confirmed findings from the pre-PR orthogonal code review. Both
false-passes matter more than a false-fail: a gate that silently misses a
violation is worse than no gate, because it is trusted.

- A relation field mixing a link with a bare id silently dropped the bare claim:
  the check tested `rel.links.length` (does this field have ANY link?) instead
  of whether THAT id was linked. `Supersedes: [ADR-0001](...), ADR-0011` passed
  clean -- accepting exactly the ambiguous bare reference the rule forbids. Now
  each bare id is checked against the ids actually linked in the same field, so
  a repeat in trailing prose stays quiet while an unlinked claim is flagged.
- The ratification guard (`statusToken !== 'Accepted'`) skipped BOTH relation
  directions, which killed the IN check entirely: `supersedes.in` is only ever
  populated on an ADR whose status IS `Superseded`, so a dangling `Superseded by
  X` where X never claims it always passed. The guard now applies to OUT only --
  a prospective claim must not obligate its target, but an ADR's statement about
  ITSELF is always owed a reciprocal.
- Fixing that surfaced a parser gap: ADR-0174 declares its supersessions in a
  `## Supersedes` table SECTION, not a header field, and headerBlock() stops at
  the first `##`. The repo's best-documented supersession was invisible. Section
  form is now parsed for both relations.
- Replaced a vacuous test: the em-dash negation case passed whether or not
  NEGATED_RELATION_RE matched (a mutation to /$^/ survived). It now carries a
  link that would create a failing asymmetric relation if negation did not fire.

Also removes docs/adr/9401-test-target.md -- a synthetic fixture a reviewer
created in the worktree while reproducing a finding, swept in by `git add -A`.

Adds regression tests for each: mixed link+bare, linked-and-repeated-in-prose,
dangling superseded-by from a non-Accepted ADR, and the ADR-0174 section shape.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: escape backslashes before pipes in the ADR index cell renderer (#2356)

CodeQL js/incomplete-sanitization (high) on scripts/gen-adr-index.cjs: cellText()
escaped `|` -> `\|` without first escaping the backslash. Markdown's escape
character is the backslash, so the input `\|` became `\\|`, which renders as a
literal backslash followed by an UNESCAPED pipe -- re-opening the cell break the
pipe escape exists to prevent. Order is load-bearing: escape the escape
character first, then everything that emits one.

Same class as the index-marker hijack fixed earlier: ADR-authored text breaking
out of the cell it is rendered into.

Adds a regression test asserting a `\|`-bearing title leaves exactly the row's
own 5 unescaped delimiters and cannot forge a Status cell. Uses split(/\r?\n/)
per local/no-crlf-fragile-split -- a literal "\n" split is CRLF-fragile on the
Windows CI leg.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 10:51:58 -04:00
Tom Boucher
cf004df678 refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1)

Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host
dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in
runCommand's default case after capability/overlay dispatch, before the
unknown-command error. Migrates 'state' as the pilot: removes the hardcoded
case 'state': arm; state now dispatches default -> dispatchHostCommand ->
routeStateCommand, byte-identical to the old path (proven by the new
state-command-cutover equivalence test, 5-category template).

Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey)
— the capability registry stays reserved for toggleable feature bundles per
ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked;
the ADR is corrected here alongside the code that realizes it.

- gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand
  (prototype-pollution-safe); wired into default case; case 'state': removed;
  dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests.
- tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY
  equivalence (recording-mock + runGsdTools end-to-end + pollution guard).
- docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability-
  registry model (correction that did not land in the merged #2355).

Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/
verify using this proven template.

Closes #2360.

* test(#2360): regenerate golden fixtures + allowlist for state cutover

Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the
install-parity fixtures (gsd-tools.cjs content hash changed), and the new
tests/state-command-cutover.test.cjs is added to the lint-test-file-count
allowlist under the 'state' prefix.

* refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify)

Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS
(state landed in the pilot commit). init preserves its #1688 warnIfStaleBake
pre-hook; validate binds the output emitter. Cutover test extended to assert
all 6 are consumed + owned. Golden install-parity fixtures regenerated.
2026-07-17 09:19:28 -04:00
Tom Boucher
f1a91072b6 docs(#2357): fix Registry Discussions category name and state its format (#2361)
The submission process told an admin to create a `Registry` Discussions
category and never stated its format. Both were wrong in a way a correct
reading of the docs could not catch.

Name: the category is `EoS Registry`, and it carries threads for both
registries — `discussion` is a required field on Capability entries as well as
EoS entries. A reader following the old text would create a second, duplicate
category.

Format: `discussion` being required means the thread must exist before the
entry's PR, opened by the entry's author — an outside contributor holding
neither `maintain` nor `admin`. GitHub's Announcement format restricts starting
discussions to those two levels, so an admin could pick it, block every external
submission at the first step, and see nothing wrong. Records open-ended as the
required format, and why Announcement and Question/Answer do not work.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 08:20:02 -04:00
Tom Boucher
ed06b6a4b9 fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir

Red phase, empirically probed: global/local install lands in command/ (singular)
with 71 gsd-*.md files and no commands/; the manifest records 71 keys under
command/ and zero under commands/; all four declaring sites report 'command'.
Migration coverage is black-box (two sequential install runs against one
configDir) so it holds regardless of how the fix implements cleanup.

The Kilo guard passes today by design — a forward-looking no-collateral check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/

OpenCode discovers slash commands from commands/ (plural); the installer wrote
them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the
TUI. Five sites declared the directory and all had to agree:

- capabilities/opencode/capability.json: both artifactLayout destSubpath entries
  (global + local) and hostBehaviors.flatCommandDir
- bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal,
  so the manifest would have diverged from the descriptor even after a rename. It
  now derives from _hostBehaviors(runtime).flatCommandDir.
- src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target,
  which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was
  a fifth site the issue did not list — without it the descriptor change alone
  would not have moved a single file.

Migration: an upgrade over a pre-fix install removes only manifest-proven
GSD-managed files from the legacy command/ dir and rmdirs it once empty.
Unmanifested user files are preserved, never deleted.

Kilo shares the opencode family install path and is explicitly unaffected —
pinned by a no-collateral test.

Note on the tests: the migration cases originally built their legacy fixture by
running the installer and relying on it to produce command/ — i.e. they depended
on the bug to set up the fixture, and became unsatisfiable the moment it was
fixed (block 1 requires command/ to be absent after a fresh install). They now
fabricate the legacy layout explicitly, including rewriting the manifest keys to
the command/ prefix — which is load-bearing, since the migration only removes
manifest-proven files and an unrewritten fixture would silently no-op and pass
even against a broken migration.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): regenerate opencode install golden after rebase onto next

The golden conflicted on rebase because #2322 also regenerated it. Resolved by
regenerating from the merged source rather than hand-merging a generated file;
the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): update stale tests that pinned opencode's singular command/ dir

Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the
descriptor's flatCommandDir, the install-integration contract, and the
resolveRuntimeArtifactLayout golden). They passed in the red phase precisely
because they pinned the buggy singular dir; the fix intentionally changes that
contract, so these are stale-test corrections, not regressions.

Kilo shares the opencode family install path and is deliberately NOT changing —
it stays on command/ (singular). The shared opencode/kilo test is now split via
an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own
layout test is untouched. tests/opencode-command-dir-plural.test.cjs
independently pins Kilo unchanged end-to-end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): changeset for opencode commands/ dir fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): backfill PR number 2354 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): correct the changeset — do not assert opencode ignores command/

The changeset repeated the issue's stated mechanism ("OpenCode discovers them
from commands/ ... a clean install produced no usable commands in the TUI at
all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts
globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc
still calls .opencode/command/ typical. Shipping that claim as a release note
would document a mechanism that does not exist.

The change is still right, for the stronger reason: OpenCode's config docs list
plural as the convention and singular as backwards compatibility, so GSD was
shipping on the alias the vendor may withdraw. Reworded to describe it as the
alignment it is, decided on OpenCode's source and docs rather than on bug reports
in either repo.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened

Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the
install destination to a surface the first-time baseline scan does not cover:
000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command',
'skills','agents'] — no 'commands'.

installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its
destination before writing the fresh set (install-engine.cts:870-873), with zero
manifest or migration involvement. The only thing that protects a pre-existing
file is assertInstallerMigrationsUnblocked, which runs before materialization and
halts when the baseline scan flags an unknown file at a KNOWN surface.

Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits
0). The identical file under the legacy, already-baselined command/ surface
correctly halts the install with "installer migration blocked pending user
choice". So this PR would have traded a protected surface for an unprotected one.

Fixed with a NEW fix-forward migration rather than editing 000, per
docs/installer-migrations.md:131-134 — an applied migration never re-runs, so
editing 000 would only protect fresh installs and leave every existing machine
exposed. A new id runs for both populations and drifts no shipped checksum;
adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions.
All five pre-existing shipped checksums verified byte-identical.

Kilo is excluded by the migration's runtimes filter and keeps command/.

This was previously deferred as a PR-body note claiming "low impact — nothing
else acts on baseline-scan misses". That claim was never probed and was wrong.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): drop the parenthetical product description from the changeset

The product-name purity guard (#1777) rejects "Kilo (which still uses
command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product
name must not carry a parenthetical. Reworded to a plain sentence; the meaning
is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 08:14:59 -04:00
Tom Boucher
15b3cc8690 docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a
two-layer dispatch (registry families + leaf-verb table filling the prepared
_dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions
ADR-959 leaves open: full dissolution, family/leaf classification rule, shared
parseFamilyArgs, and the capability-arm extraction shape. Phased under epic
#2345 (P1-P4). Behavior-preserving; each cutover proven by the
audit-command-cutover equivalence template.

- docs/adr/2346-command-dispatch-completion.md (new)
- docs/adr/959-*.md: Status Proposed -> Accepted + amendment section
- docs/adr/README.md: index rows for 959 + 2346
- docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub
- CONTEXT.md: seed glossary entry

Closes #2346 (docs-only; no production code).
2026-07-17 07:19:17 -04:00
Tom Boucher
ada79bee97 fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md

Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md
unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and
per-workstream milestone state already lives in the workstream's own STATE.md /
ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design —
whichever workstream ran new-milestone last silently won the shared heading.
Step 4 is now skipped when a workstream is active; step 6 no longer stages
PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing
when a staged path is unchanged).

Also fixes a second defect found while diagnosing this, same root cause (the
workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and
the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at
the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x`
suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped,
violating the routing-propagation contract. Step 1 now parses --ws using the
established idiom from verify-work.md.

Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not
export the latter and it is only priority 2 of 5 in resolution, so it would miss
the --ws flag case that is the actual repro.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2308): regenerate install goldens for the new-milestone workflow change

gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content
hash is pinned in all 18 runtime golden fixtures. Only that hash changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests

Independent review found the first pass was partly cosmetic:

1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in
   step 1's shell and each step's bash block runs in its own shell — this file
   already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file
   for exactly that reason (#2288). The guard read an unset variable, always took
   the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS
   in step 6, the branch is removed entirely: step 4 Part A's guard is what
   protects the shared heading, so post-guard the only content PROJECT.md can
   carry is Part B's idempotent Evolution backfill — which must be staged, not
   stranded. A regression test now asserts no cross-step GSD_WS branch returns.

2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a
   shared, idempotent backfill that is not workstream state. A pre-Evolution
   project running only `--ws` would never get the section that transition and
   complete-milestone expect. Step 4 is now split: Part A (milestone-state write)
   is workstream-guarded; Part B (Evolution) always runs.

3. The tests were tautological prose-pinning — including one asserting a comment
   mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it
   passed on the inert guard it existed to catch. Replaced with executable tests
   that extract the step-1 and step-6 fences and run them under bash with stubbed
   gsd_run, asserting real parse and --files behavior.

4. --ws is now stripped from the milestone name (step 1 previously left
   "--ws search" in the remaining text), and documented in argument-hint,
   help/modes/full.md, and docs/COMMANDS.md.

5. Changeset no longer overstates: --ws reaches the prose guard and routing hints
   only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone
   still take no ${GSD_WS} — out of scope here).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change

skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md,
so documenting --ws in the argument-hint made it stale (caught by lint:ci's
gen-plugin-skills --check). Regenerated it plus the install goldens and workflow
size baseline, since commands/, skills/, and gsd-core/workflows/ all ship.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2308): backfill PR number 2338 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:13 -04:00
Tom Boucher
ff9cb6069f fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active'
but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller
outside their own CLI router, and execute-phase.md declared an
execute:wave:pre hook point that the workflow body never rendered — so
claude_orchestration.enabled:true had zero effect on real runs.

Approach B (maintainer-chosen):
- execute-phase.md now renders the execute:wave:pre hook
  (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75,
  immediately before each wave's Agent() dispatch — fixing the latent
  dead-hook gap for any pre-wave capability.
- Move the claude-orchestration contribution execute:wave:post ->
  execute:wave:pre (a pre-wave backend selector belongs before dispatch,
  not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md
  with prose instructing the orchestrator to call resolve-wave-dispatch
  before step 3. Unrelated wave:post contributions (ui.safety-gate, drift,
  external-job, mempalace) untouched.
- New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend
  + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result;
  exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is
  a real non-CLI-router, non-test caller of both functions.

Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool
absent, SDK below floor, execution_backend:inline, malformed input) or an
emit failure resolves to inline with a byte-identical result shape — no
regression to the default-off execute-phase path.

Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor
BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a
fast-check composition property, capability.json contribution assertions,
and a source-contract guard that execute:wave:pre is now actually rendered.
Dependent registry-shape assertions updated in-scope.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 19:18:28 -04:00
Tom Boucher
4a9833d3e3 fix(#2278): use Edit() not Write() for Claude allow-permissions + migrate legacy (#2302)
GSD_CLAUDE_ALLOW_PERMISSIONS pre-populated Claude Code settings.json
with Write(.planning/*) and Write(STATE.md). Claude Code has no
standalone Write permission gate — file-editing tools are gated
collectively via Edit(pattern) — so those rules never matched, fresh
installs still hit first-run approval prompts for .planning/* and
STATE.md, and Claude Code emitted a session-start warning about the
unmatched rules.

Swap the two entries to Edit(.planning/*) / Edit(STATE.md). Add a
GSD_CLAUDE_LEGACY_ALLOW_PERMISSIONS list of the retired Write(...) forms,
consulted by mergeClaudePermissions (actively remove stale entries when
adding current ones, idempotent, user entries preserved) and by the
uninstall cleanup filter (still removes the legacy form). Sample
settings.json in docs/USER-GUIDE.md corrected to match.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 13:46:26 -04:00
Tom Boucher
315d94f6d4 feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate

Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode.

- gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top.
- gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer.
- --no-tracer flag wired through plan-phase workflow/command/help/skill.
- CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled.
- tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1945): backfill changeset PR number to 2294

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:41:36 -04:00
Tom Boucher
c4237df8e6 docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a
conceptual Embeddable Orchestration System (EoS) explanation
(docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md
with a v1.7.0 feature section, and wire both new docs into the docs index
(docs/README.md) and the root README.

Covers the release's marquee changes: the ADR-1239 Host-Integration Interface /
EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS
discoverability registries, the gsd-mcp-server companion, model-catalog advances
(GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format,
plus a themed summary of the 100 fixes and 4 security hardenings.

Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay
now documents the #2009 fail-open behavior for a load-failed gate-declaring
capability (previously described as fail-closed).

American house style; no parity-gated reference docs hand-edited.

Refs #2276, #1678

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:49:44 -04:00
Cody Anderson
20ff405cb3 feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline

New statusline.state_format config, enum full|compact (default full —
existing rendering untouched). "compact" renders the state segment as
"<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 ·
executing" — dropping the milestone name and progress bar (the two
biggest width costs) and collapsing narrative statuses to a single
keyword. Per the #2162 approval conditions, the keyword set is the
canonical vocabulary from normalizeStateStatus() in state-document.cjs
(discussing/planning/executing/verifying/completed/paused) — no
parallel hand-rolled list, so the vocabularies can't drift — and the
canonical stuck state "paused" renders uppercase as PAUSED (no new
"blocked" lifecycle state). Statuses the normalizer passes through
unrecognized fall back to their first word capped at 16 chars.
Lifecycle scenes preserved: active_phase wins over the body phase
number, milestone completion renders "complete", idle-with-next-action
renders "next <action> <phases>".

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2162): changeset fragment for PR #2175

* fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format

- register statusline.state_format in the fix-1628 coercion-bypass matrix
- 15/16/17-char boundary tests for the shortGsdStatus fallback cap
- changeset body ends with the (#2162) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage

- compact renderer gates the milestone-complete scene behind the absence of
  an in-flight phase id, mirroring formatGsdState's if/else precedence
  (Scene 1 beats Scene 3); regression test covers the non-atomic
  active_phase + percent=100 STATE.md shape
- direct config-set accept/reject test for statusline.state_format plain
  strings (ENUM_KEYS matrix covers only the JSON coercion shapes)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests)

Re-review Major: gating done on !phaseId held completion back for the
legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100
regardless of phaseNum, so compact must too. Gate is now !s.activePhase.
The phaseNum-only test now expects 'complete' and cross-checks the full
renderer; a parity test feeds identical inputs to both renderers.
Re-review Minor: shortGsdStatus gets fast-check property coverage
(totality, canonical fixed points, separator safety, fallback shape).
Golden fixtures regenerated for the hook byte change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-14 21:09:24 -04:00
Cody Anderson
eec9efc351 feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge (#2173)
* feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge

Claude Code appends " (1M context)" to the model display name in
long-context sessions, eating 12 characters of statusline width. Collapse
it to " (1M)" — the signal stays, the width doesn't. Any other display
name passes through unchanged.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2160): changeset fragment for PR #2173

* fix(#2160): review fixes — ctx variant, boundary test, changeset format

- broaden the suffix match with a context|ctx alternation (approval-condition
  variant the regex missed)
- pin non-context parentheticals ((beta), (deprecated)) as untouched
- changeset body ends with the (#2160) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-14 20:37:43 -04:00
Tom Boucher
fc913b37a5 refactor(#2268): gen:golden one-command fixture regenerator (#2275)
Phase 3 (convenience form) of golden-parity redesign (epic #2264). Adds npm run gen:golden (regenerates both fixture sets) and points the golden-parity/tree failure messages at it. Full CI-auto-comment deferred (documented in ADR-2264). Closes #2268.
2026-07-14 20:18:34 -04:00
Tom Boucher
89b1bef881 refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267.
2026-07-14 19:22:13 -04:00
Tom Boucher
ef5a5bc15d docs(#2265): ADR-2264 golden-install-parity redesign (#2270)
Phase 0 of golden-install-parity redesign epic #2264. Adds docs/adr/2264-golden-parity-redesign.md + index entry. Closes #2265.
2026-07-14 15:32:19 -04:00
Tom Boucher
8b70db343b fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207

completePhaseCore was writing the overloaded bare 'Milestone complete' on the
last phase — the same string space the milestone-close verb owns for terminal
state. Per ADR-2207, phase-completion now writes the existing intermediate
value 'All phases complete' (already used in gsd2-import.cts). Milestone
termination ('<version> milestone complete' / 'Awaiting next milestone')
remains solely with milestoneCompleteCore.

Status lifecycle: Ready to plan → All phases complete → <version> milestone
complete → Awaiting next milestone.

Changes:
- src/state-transition.cts: completePhaseCore status value
- src/phase.cts: #2028 guard comment
- tests/state-transition.test.cjs: assertion + test name
- tests/phase.test.cjs: 8 assertion updates (positive + negative)
- tests/state.test.cjs: normalizeStateStatus test case + reset regex
- tests/workstream.test.cjs: fixture status to terminal value
- gsd-core/workflows/progress.md: Route D label
- gsd-core/workflows/transition.md: Route B label
- CONTEXT.md: Status lifecycle glossary entry (ADR-2207)
- .changeset/brave-geese-jump.md

* test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline

Workflow file edits (progress.md, transition.md) changed install payload
hashes and pushed past the committed workflow-size baseline. Regenerated
all 17 golden-install-parity fixtures + claude-local via the standalone gen
script (which now also covers the local-scope claude layout). Updated
workflow-size-baseline.json and agent-size-baseline.json via size:baseline.

* fix(#2204): correct claude-local golden hashes + document gen-script limitation

The gen-script's claude-local generation produces macOS-specific hashes
incompatible with Linux CI (local-scope install embeds platform-varying
node-runner paths). Reverted to manual update using Linux FAILURES.md
+actual hashes for the 2 changed workflow files. Added explanatory
comment in the gen script.

* test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary

Addresses orthogonal code-review findings (Medium #1 + #2):
- Add isCompletedInventory test cases for ADR-2207 status lifecycle
  (terminal 'milestone complete' → true; intermediate 'All phases
  complete' → false; archived → true; active statuses → false)
- Clarify CONTEXT.md glossary: note that isCompletedInventory
  intentionally excludes the intermediate value

* docs: backfill changeset PR number (#2259)

* docs(#2204): add Status lifecycle table to state-md reference (ADR-2207)
2026-07-14 14:47:03 -04:00
Tom Boucher
2cbf186420 chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3

Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112,
#2118) were already fixed tactically on next; this introduces the reusable
structural contracts and rewires the primary #2140 site onto them.

- src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the
  per-surface write-set (`WriteOutcome {surface, applied, requirement?}`,
  `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports
  `Result` from here (single source; distinct from command-routing-hub's Result).
- requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT,
  per-surface write-set; `write_set_complete` is true only if every surface of
  every requirement applied — structurally forbidding the #2140 OR-into-one-flag
  masking, including across a multi-ID batch (adversarial-review regression).
  Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed).
- deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial
  null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow;
  RoadmapProgress return contract unchanged.
- commit --files (#2112) and milestone complete --dry-run (#2118) left as-is
  (single-surface commit / pre-mutation preview — not genuine multi-surface writes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2244): backfill changeset PR number (#2251)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:39:58 -04:00
Tom Boucher
efd04716da chore(#2143): withSection bounded-mutation seam + phase.cts migration — Phase 2 (#2250)
* chore(#2143): bounded-mutation seam (withSection/withPhaseSection) + phase.cts migration — Phase 2

Phase 2 of epic #2143 (ADR-2143 §4): add a bounded-mutation primitive so a
per-phase ROADMAP edit is structurally confined to that phase's own section,
and migrate the phase-scoped mutation sites in `phase.cts` onto it.

- `src/markdown-sectionizer.cts`: `withSection(content, target, edit, opts?)` —
  resolves a section via `collectSection` and applies `edit` to ONLY that
  section's body, re-serialising via `replaceSection`. The edit callback sees
  only the section body, so any regex it runs is physically confined.
- `src/roadmap-parser.cts`: `withPhaseSection(content, phaseId, edit)` —
  resolves a phase's `### Phase N` detail-section heading via the #2121
  phase-id source and delegates to `withSection`. Heading match is anchored to
  the heading start (a sibling phase whose title mentions the number is not
  hijacked) and bounds at the next ATX heading of any level (`levelBounded:false`).
- `src/phase.cts`: `mutateMilestonePhase`'s plan-count and per-plan-checkbox
  writes now route through `withPhaseSection` — structurally retiring the
  #2130 / #2067 / #2080 boundary-crossing class for these sites. The phase-LIST
  checkbox is intentionally left milestone-slice-scoped (it lives outside any
  `### Phase N` detail section). Cross-phase renumbering is untouched.
- Property test (fast-check): editing phase k leaves every sibling section
  byte-identical; regression tests for title-collision + mixed heading depth.

Behaviour-preserving (verified by old-vs-new differential runs on real
fixtures). Extend-never-mutate (ADR-2143 §2). Registration: CONTEXT.md +
docs/INVENTORY.md export lists.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2243): backfill changeset PR number (#2250)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:39:22 -04:00
Tom Boucher
d49ac81306 chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1

Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a
canonical seam and migrate the pilot reader.

- Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed
  {columns, rows} addressed by column NAME; ragged rows are typed parse
  errors, not silent), a single-source TABLE_SCHEMAS registry
  (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with
  variants under one id), matchTableSchema, and findTableBySchema. Result<T>
  is scoped to this seam (distinct from the dispatch Result).
- Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the
  position-anchored regex to name-based resolution via the seam — fixes #2137
  (the 5-column milestone-grouped Progress table previously returned all-null).
- Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route
  fast.md's log_to_state through it, retiring the inline `awk NF-2` column
  arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped
  (| and newlines) and the STATE.md read-modify-write is atomic under
  readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230).
- Writer/reader/template parity test guards TABLE_SCHEMAS against drift
  (ADR-2143 §3 Generative-Fix-Divergence).

Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md +
INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md.

Behaviour-preserving for the canonical 4-column Progress table; the named
bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2242): backfill changeset PR number (#2248)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): escape backslash before pipe in markdown-table cell escaping

CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not
the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow
unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl.
literal backslashes) round-trip exactly. Added backslash round-trip tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan

Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself
"pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via
a new seam helper findTableWithColumns (first table whose header is a superset of
Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads
cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact
TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while
staying seam-based and preserving its `## Progress` scoping (#2012/#1445).
Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale
state.test.cjs assertion that predated the Phase-1 migration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:36:14 -04:00
Tom Boucher
db725e49e0 Merge pull request #2183 from arakasi1/feat/2163-statusline-git-segment
feat(#2163): opt-in git branch/status segment in the statusline
2026-07-13 16:38:56 -04:00
Tom Boucher
1047d27320 Merge branch 'next' into docs/612-bracket-phase-id-convention 2026-07-13 16:29:26 -04:00
Cody Anderson
d6672ff926 feat(#2163): opt-in git branch/status segment in the statusline
New statusline.show_git config (default false). When enabled, a git
segment renders after the directory: current branch plus compact
work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓
when clean and in sync), e.g. " │ main+2~1?3".

One git status --porcelain=v2 --branch spawn per render via execFileSync
with a fixed argument array (no shell), a 1.5s timeout, and the
workspace dir passed with -C. Fails silently — segment absent outside a
repo, without git, or on timeout. Default output is unchanged when the
flag is absent.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 11:33:32 -06:00
Cody Anderson
ad7111e50b feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter

New statusline.show_context_tokens config (default false). When enabled,
the context meter shows the absolute token total after the percentage,
e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read,
and output tokens from context_window.current_usage (matching /context).

Default output is byte-for-byte unchanged when the flag is absent or
false. The .planning config is now read once per render and shared with
the last-command/position block instead of being re-read.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2161): changeset fragment for PR #2174

* fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format

- formatTokens promotes to the M branch when k-rounding reaches 1000
  (999,500-999,999 rendered "1000k" instead of "1.0M")
- boundary tests at 999499/999500/999999/1000000/1000001
- Number() guards on the four usage fields (silent string-concat gap)
- changeset body ends with the (#2161) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style

- config-set accept/reject tests for statusline.show_context_tokens
  (mirrors the post-planning-gaps precedent the issue scope names)
- changeset + docs no longer claim parity with /context: the suffix sums
  four fields while the meter %% derives from used_percentage (three), so
  the figures can diverge slightly
- module.exports one entry per line

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:25:47 -04:00
Tom Boucher
ffd353f080 docs(#2240): clarify PLAN.md <execution_context> is install-relative, record #2238 wontfix (#2241)
The <execution_context> reference presented @~/.claude/gsd-core/... as canonical
and described the paths as files "the executor reads before starting". Both are
misleading: the prefix is install- and runtime-relative (Claude global vs Cursor
.cursor/gsd-core vs an absolute --local path), so a committed plan is not
clone-portable, and /gsd-execute-phase loads the workflow from its own installed
copy rather than gating on the committed block.

- docs/reference/plan-md.md: describe the install-relative, non-clone-portable
  nature of the block and contrast it with repository-relative <context>.
- .out-of-scope/plan-md-execution-context-portability.md: record the #2238
  wontfix decision (PLAN.md is a machine artifact; #2158 precedent) with a
  revisit-if condition.

Refs #2238. Closes #2240.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 11:58:02 -04:00