Commit Graph

7 Commits

Author SHA1 Message Date
Tom Boucher
03b7125293 enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks

Binds the four fabrication sites found by executing the surfaces (ADR-3889
failure class (c)), each with a positive control so an over-firing fix goes red:

- the blocking api-coverage.verify-pre gate certifying "no external-API
  integration" from a zero-byte phase scope
- the assumption-delta query route scanning an unresolvable phase section as
  the empty string and reporting it as an examined negative
- both capability fragments' probe fallbacks, which append a fabricated
  verdict rather than replacing, and fire on the legitimate exit-1 negative

Verification runs on the remote runner.

Refs #3909

* enhance(#3909): a probe that could not run no longer asserts a verdict

ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a
confident negative; each now reports what it could not establish.

- check api-coverage.verify-pre: a phase with no plan body and no roadmap
  section ran detection over zero bytes and PASSED the blocking seal gate,
  certifying "no external-API integration" from input it never read. It now
  holds with scope_unavailable. The discriminator is bytes examined, never
  signals found, so a phase whose plans are real and simply carry no API
  vocabulary passes exactly as before.
- query assumption-delta scan: an unresolvable phase section was scanned as
  the empty string and reported as an examined negative. It now returns
  {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded
  result in the payload, leaving the gsd-tools exit projection to P8.
- both capability fragments: `|| echo '{"detected":false}'` appended rather
  than replaced, and fired on the legitimate exit-1 negative, so a correct
  answer and an honest skip both arrived as two concatenated objects. They now
  keep the probe's own payload and manufacture only an explicit
  probe_unavailable skip when the probe produced nothing at all.

Every registered outcome is more restrictive on a blocking gate, so this can
turn a false green red and never a red green.

Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md
seal-time outcome table, and a new how-to for the reason-code vocabulary.

Verification runs on the remote runner.

Closes #3909

* test(#3909): correct the stale unknown-phase assertion

`unknown phase → detected:false, no throw (graceful)` scanned phase 999
against a two-phase roadmap and asserted `detected === false`. That pinned
the fabrication as intended behavior: the phase does not exist, so the
detector was handed the empty string and its "no core assumption changed"
answer described nothing that was ever read.

It now asserts the skipped-with-reason shape. The graceful-degradation
contract the test was actually protecting — the query succeeds and does not
throw on an unknown phase — is unchanged.

Found by code review, not by the author.

Refs #3909

* docs(#3909): author the FEATURES entry in its generator source

`docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the
per-feature fragments in `docs/features/`. The API-coverage entry was edited
in the generated file, so the next regeneration silently dropped it.

The text now lives in `docs/features/api-coverage-gate.md` and
`docs/FEATURES.md` is regenerated from it, leaving the shipped file
byte-identical and its content actually derivable.

Caught by `lint:generated-sync`.

Refs #3909

* test(#3909): bind the skip to "not found", and pin the discriminator

The first verification run went red on one case, and the case was wrong
rather than the code.

`getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a
missing ROADMAP.md, but for a section whose body is whitespace-only it returns
the heading line alone — which is not empty. So a body-less section WAS found,
and reporting `detected:false` over its heading is a real negative, not a
fabrication. The test had assumed the resolver yielded `''` there.

Correcting the test rather than the resolver keeps `skipped` bound to the
distinction the issue asks for — found versus not found — and avoids diverging
`assumption-delta scan` from `roadmap.get-phase`, which the fragment documents
as sharing one resolver.

Also adds the seeded property the test matrix had promised: for any plan body,
the scope read back is whitespace-only exactly when the body was. That pins the
gate's discriminator to bytes examined, so it cannot quietly become "no signals
found", across unicode whitespace and CRLF.

`docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table —
surfaced by the co-change gate, not by a lint failure.

Refs #3909

* chore(#3909): backfill the changeset PR number

Refs #3909

---------

Co-authored-by: sim <sim@local>
2026-08-27 15:50:12 -04:00
Tom Boucher
9410f7e6e6 enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4

ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the
other three RED before any fix.

Rung 2 — the install marker has four readers and resolveRuntime is not one.

  resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no
  marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled
  readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with
  test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in
  hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule.

  Fixtures and seam names mined from PR #3382 rather than re-derived; it
  implemented this rung and was closed "not on the merits".

Rung 3 — the sandbox map, and the fallback that was the real defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit:

    - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly,
      zero disagreements — the map carries nothing the contract does not
    - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit

  So the map is redundant and the silent fallback is the defect. The maintainer
  chose to derive but hold those 16 at read-only pending the question of whether
  Codex enforces sandbox_mode or merely advises; HALT.md records it.

  T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not
  in aggregate — an aggregate passes while one role silently widens, which is
  the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a
  stale hold, so the hold list cannot rot into the subset map being deleted.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId`
  returns only documentation commits. That was the wrong instrument. Direct
  inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and
  the tests match that code rather than a guess at its semantics — including
  first-write-wins on a duplicate short form.

  T43 asserts at the consumer's output: the emitted `waves` map from the real
  CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is
  dropped. A unit assertion on resolveDependencyId would have passed throughout
  this defect's life.

Observed RED, this tree:
  rung 2   11/11 fail — no marker rung, no seams
  rung 3   T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML
           whose sandbox_mode disagrees — it checks presence only)
  rung 4   T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1

Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's
unresolvable-token warning and wave-verdict suppression, and #3785's
display-mapping passthrough. If the third tier over-reaches, those go red — that
is their job.

Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed)
cannot be isolated behaviorally, because planMap always masks it. It is a
non-crash boundary pin, weaker than the other rows, and is recorded as such
rather than presented as equivalent.

Design:      .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md
Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md
Decision:    .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier

ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the
other three.

Rung 2 — the install marker had four readers, and resolveRuntime was not one.

  resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no
  marker, while bin/install.js writes one (#2297) and four hand-rolled
  readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with
  seams), hooks/gsd-agent-isolation-guard.js, and twice in
  hooks/gsd-cursor-subagent-start.js.

  model-resolver's was already the house idiom, so it was promoted rather than
  replaced: src/runtime-slash.cts now owns it, and model-resolver plus both
  hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam
  lint-hooks-runtime-build-seam enforces. No import cycle existed - checked
  both directions before moving anything.

  The marker is the THIRD rung: env > project config > marker > 'claude'.

  N1 was checked rather than assumed, and my first reading of it was wrong. A
  marker holding an unknown name comes back essentially verbatim, which looked
  like a validation gap. Measured against the env rung with the same inputs -
  including "../../etc/passwd" and "claude;rm -rf /" - the two are identical,
  because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that,
  and it is met. The residual (the shared normalizer normalizes shape, it does
  not validate against the known-runtime set) is pre-existing on the env rung
  and plausibly deliberate, since a new runtime should not need a code change.
  The marker also does not widen the trust boundary in any real sense: it lives
  inside the install tree beside the code, so anyone who can write it can write
  runtime-slash.cjs itself.

Rung 3 — the map was redundant; the silent fallback was the defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their
  mapped value exactly, zero disagreements. The map carried nothing the contract
  did not already have, so it is DELETED rather than reconciled. What was
  actually broken is `|| 'read-only'`, which silently under-granted 24 of 35
  roles.

  16 of those 24 declare Write or Edit and would widen under derivation. Per the
  maintainer's decision (HALT.md), they are held at read-only pending the
  question of whether Codex enforces sandbox_mode or merely advises. Emitted
  TOML is therefore byte-identical for all 35 roles - asserted per role, not in
  aggregate, because an aggregate passes while one role silently widens.

  The hold list self-invalidates. A hold whose role no longer derives broader
  fails, and so does a hold naming a role with no agents/<name>.md. Without
  that it would rot into exactly the hand-maintained subset map being deleted,
  and this commit's own ledger claim would become false over time. Both cases
  were proved by injecting them and watching them throw.

  Two committed tests asserted the deleted map's existence and contents. They
  were pinning the thing being removed, so the tests moved rather than the
  production code: the 11 role-value pairs survive as a test-local
  PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real
  derivation against real agents/*.md. The coverage is preserved; only its
  source moved out of production code.

  validate agents gains checkCodexSandboxPosture, mirroring the existing
  checkCodexModelPosture: each installed TOML's sandbox_mode must equal the
  role's expected value, failing with role, expected and found. It previously
  checked file presence and manifest completeness only, so a TOML whose
  sandbox_mode disagreed passed.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: git log -S returns only
  documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^
  carries five occurrences, and the implementation here matches that code rather
  than a guess at its semantics - including first-write-wins on a duplicate
  short form, deterministic from the sorted plan order.

  It resolves the bare plan number: depends_on: ["01"] now reaches
  26-01-auth-hardening. That is a control-flow change, not a diagnostic one -
  plans that silently collapsed into a single wave 1 now execute in their
  declared waves, and execute-phase.md consumes those wave values.

  In-phase only, by construction: the map is built from this phase's rawPlans,
  so a same-named short form in another phase does not resolve.

  #3785's display-mapping passthrough and #3885's unresolvable-token warning and
  wave-verdict suppression are untouched and stay green. If the third tier had
  over-reached, those are what would have caught it.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open I introduced, and wire the posture check to its command

Two blockers from review. Both are mine, and one is a security regression my own
change created.

1. A held role could escape its hold by editing its own frontmatter.

  The Codex install loop set the sandbox identity from the agent's frontmatter
  `name:` field rather than from its filename, so the hold lookup keyed off a
  value the file itself declares:

    deriveCodexSandboxMode('gsd-doc-writer',   <real file>)          -> read-only
    deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write
    deriveCodexSandboxMode('GSD-Doc-Writer',   <same file, name: recased>) -> workspace-write

  What makes this a blocker rather than a nit is the DIRECTION. The deleted
  CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an
  allowlist: an unmatched key fell back to read-only, which is safe. The new
  scheme derives workspace-write from the tool contract and uses the hold as a
  subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk
  into a fail-open one and did not notice; the isolated reviewer proved it by
  execution.

  Neither safety net caught it. validateCodexSandboxHolds only checks that
  <key>.md exists, never that a file's derived identity matches its key.
  checkCodexSandboxPosture looks the canonical source up by the installed TOML's
  filename, finds nothing for a renamed agent, and treats it as a custom
  non-roster agent — silently no violation.

  The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds
  already validates and what an attacker editing frontmatter cannot change
  without renaming the file — at which point the existing validator catches it.
  The lookup is case-insensitive so a recase does not slip past either. The
  frontmatter name still drives the TOML body and filename, unchanged; only the
  sandbox identity moved.

  All 35 roster files were checked: name matches filename stem everywhere, so a
  stricter "they must agree or throw" invariant would have been safe against real
  content. It is deliberately NOT added — it would abort an install on a tampered
  file where emitting a correctly-derived read-only TOML is the safer outcome.
  Recorded as a fork rather than decided silently.

2. checkCodexSandboxPosture was exported and never called.

  cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and
  checkCodexModelPosture only; grep for the sandbox check in that file returned
  nothing. So criterion 3 — "validate agents fails on semantic drift, not only on
  missing files" — was unmet, and `validate agents` behaved exactly as before.
  That is ADR-3473 Decision 2's named shape: a declared policy with no executor.

  It also meant the T28 test asserted at the helper's return value while the
  COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to
  close, committed by me while enforcing it elsewhere in the same epic.

  Now wired as an additive `sandbox_posture` field beside `codex_posture`,
  following the sibling precedent exactly. Drift is report-only, not a non-zero
  exit, because that is what checkCodexModelPosture does — two sibling posture
  checks disagreeing about whether a violation is fatal would be its own defect.
  The choice is recorded in a comment rather than left implicit. A consumer-output
  test now drives the real CLI and asserts on the emitted JSON, and was shown
  failing before the wiring and passing after.

Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3
was not in this deliverable, written while it was halted and false once the
maintainer unblocked it.

Verified after both fixes: the three bypass probes all return read-only, the
per-role table is 35/35 byte-identical, and both hold self-invalidation cases
still throw.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on

Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install
marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture
field; docs/reference/plan-md.md documents that depends_on accepts the bare plan
number.

Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a
concurrent PR hand-allocating a section number, regenerated into FEATURES.md.

ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style,
recording what was measured and built against the section's 2026-08-26 correction
- including the qualification that checkAgentsInstalled itself still checks
presence only, and the semantic assertion lives in a sibling wired into validate
agents rather than folded into it.

No how-to. Both user-visible changes are zero-step: a non-Claude install resolving
its own runtime, and plans executing in their declared waves, both happen without
the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a
DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does
not touch, and was deliberately left alone rather than edited by association.

No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it
documents Claude-side tools frontmatter, never Codex sandbox_mode, and the
emitted tools contract did not change.

The prompt layer documents depends_on only by example, not by schema, so nothing
there needed editing - and few-shot-examples/plan-checker.md already showed
depends_on: ['01'], which now actually resolves.

Translated copies of plan-md.md are untouched; the project treats translations as
community-maintained.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser

The full suite came back with 26 failures across four files. Three distinct
causes, mapped individually rather than assuming the first explained the rest.

A. Requiring bin/install.js printed the GSD banner to stdout and corrupted
   `validate agents` JSON.

     Unexpected token '', "[36m   ██"... is not valid JSON

   checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring
   bin/install.js, whose module load prints the ASCII banner. So the command
   emitted banner bytes before its JSON and every JSON consumer broke, including
   ten tests that predate this branch. src/ reaching into bin/ was backwards
   layering that happened to also be loud.

   The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML
   domain module, no new module and no six-gate ripple - and both bin/install.js
   and src/agent-install-check.cts import it. One owner, which is §8.3's rule
   applied to the fix for §8.3.

B. The stale-hold throw fired on a legitimate partial source dir, and masked a
   security assertion.

   validateCodexSandboxHolds treated "this hold's .md is absent from the install
   SOURCE dir" as a stale hold and threw. A test fixture, or any partial install
   source, legitimately contains a couple of agents. Worse, it threw BEFORE the
   path-escape check, so a test asserting that a `../../evil` frontmatter name is
   rejected got my unrelated error instead of the traversal rejection it was
   written for. A fail-closed check of mine was hiding a real security check.

   The "no stale holds, shrink-only" invariant is a property of the repo's
   canonical agents/ roster, not of whatever directory an install happens to read.
   It is off the runtime path and enforced where it belongs, in the tests that
   already existed for it. A partial source dir now installs cleanly, and the
   evil-name case throws with its own escapes-configHome message again.

C. T8 depended on ambient process.env state.

   The marker/env parity assertion round-tripped through live process.env. It now
   compares against resolveExplicitRuntime's already-exported dependency-injection
   parameter - deterministic and hermetic, same claim. Proven still falsifiable
   rather than assumed: with the marker rung's normalization temporarily bypassed
   the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd")
   and the assertion fails, then passes again once reverted.

One correction folded in along the way. The first version of the move added
private _extractFrontmatterAndBody/_extractFrontmatterField helpers to
codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph
already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893).
Adding a third inside the epic whose thesis is one implementation per rule is not
defensible. deriveCodexSandboxMode no longer parses anything: it takes
(identity, toolsValue) and each caller supplies the tools value using the
extractor it already has. Both helpers are deleted. The identity argument is
still the filename stem, so the fail-open fix is untouched.

Verified after all three: `validate agents --raw` emits parseable JSON with no
banner and both posture fields; the four hold-bypass probes still return
read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9
workspace-write; the hold list is still 16.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test

Suite down to 7 failures from 26. Three more causes, mapped individually.

A. My extractor import dragged in a script that does not exist in an installed
   tree.

     Cannot find module '../../../scripts/fix-slash-commands.cjs'

   Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs,
   which requires command-roster.cjs, whose line 36 requires
   ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not
   in an install, so every test exercising a synthetic install dir died at module
   load. I picked that extractor for convenience without checking what it pulls
   in - the same mistake that produced the banner bug, one layer further out.

   agent-install-check now uses a single-purpose extractToolsLine on
   codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser:
   we deleted those helpers a commit ago for good reason, and this reads one
   line. Verified from outside the repo root that requiring either module prints
   nothing and does not throw.

B. A test pinned the deleted name-based fallback.

   'defaults unknown agents to read-only' called generateCodexAgentToml with a
   fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent
   with a writing contract correctly derives workspace-write - design row S6, a
   new writing role gets the contract, not the pin. The behavior it asserted was
   the silent fallback this rung deleted; identity no longer decides the sandbox.

   Replaced with two rows rather than a flipped string: no tools declared ->
   read-only (absence is not a grant), and Write/Edit declared -> workspace-write.
   Strictly more coverage than the row it replaces.

C. The stale-hold check still threw per derivation call.

   Last commit took the roster-existence check off the install path, but
   deriveCodexSandboxMode itself still threw when a hold's role did not derive
   broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture
   for a held role.

   The throw is gone, and it cost nothing: if a held role's content does not
   derive broader, the hold pins read-only and derivation returns read-only
   anyway, so the hold is a no-op and there is nothing to fail about. The
   staleness invariant is a property of the real agents/ roster, and
   validateCodexSandboxHolds still enforces it there - confirmed against the real
   roster after the change, not assumed.

   deriveCodexSandboxMode is now total: every (identity, toolsValue) including
   undefined and null returns read-only or workspace-write, never throws.

Verified: validate agents emits parseable JSON; the four hold-bypass probes
return read-only; the per-role table is 35/35 at 26 read-only / 9
workspace-write.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path

The ADR entry and the feature fragment both ended their rung-3 explanation with
"see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That
directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at
read-only was reachable only from the machine that produced it. A reader of the
ADR got a pointer to nothing.

Both now carry the reasoning inline: the criterion asks both that the sandbox
derive from the declared tool contract and that no role gain a broader sandbox,
and those cannot both hold, because a faithful derivation widens 16 roles the
deleted map never listed and that fell through its silent read-only default. The
resolution is derive-and-hold - the derivation owns the rule now, each hold is
released as its enforcement question is answered, and a hold is reversible where
a widened sandbox that turns out to be enforced is not.

Checked before assuming this was a defect class: CONTEXT.md cites
.gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight
module entries, and four other shipped docs do the same. Citing a phase artifact
is an established convention here, so those are left alone. What was wrong was
specific to these two: they put load-bearing rationale behind the pointer instead
of provenance.

docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs
rather than hand-edited.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration

Two orthogonal reviews on the shipped sha. Three of the findings are the same
failure class this epic exists to close, committed inside it.

1. BLOCKER - the sandbox was decided for one identity and applied to another.

   bin/install.js derived sandbox_mode for the filename stem and then wrote the
   result to `${name}.toml`, where name comes from the file's own frontmatter.
   Make the two disagree and a HELD role's artifact goes wide:

     rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer
       -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml
     add any gsd-*.md whose frontmatter name: is a held role
       -> clobbers that role's toml with workspace-write

   Both emit read-only on origin/next, because the deleted map was an allowlist
   and a miss fell back safe. This is a regression my change introduced. The
   previous review round moved the HOLD KEY off frontmatter to the filename stem
   and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985
   calls that value attacker-editable, four lines above the line that uses it as
   the filename.

   The decision is now made over BOTH candidate identities, most-restrictive
   wins: if either the stem or the emitted name is held, the mode is read-only.

2. MAJOR - hold matching was toLowerCase() only, so confusables escaped.

   Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline,
   ./ and ../agents/ all slipped the hold and emitted workspace-write.
   Identities are now basenamed, trimmed of NBSP/zero-width/control characters,
   NFKC-normalized and lowercased - and anything still carrying a character
   outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not
   enumerate confusables; every shipped roster file is ASCII, so refusing to
   widen on an identity we cannot recognize is fail-closed with no false
   positives on real content.

3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY.

   shortFormToId keyed on the last dash-segment of any canonical id with no
   constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md
   made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the
   worst shape in the epic: the unresolvable-token warning fires on a DROPPED
   token, so a MIS-RESOLVED one is invisible and the tool reports a confident
   wave assignment built from a wrong edge. A wrong edge is worse than a missing
   one.

   The segment must now match /^\d+$/, which is exactly the contract
   docs/reference/plan-md.md already documents. This tier was recovered verbatim
   from the retired SDK lineage, which carried the same defect; we are
   deliberately NOT preserving it bug-for-bug, and the comment says so, so the
   next reader does not "restore" it.

4. MAJOR - the derivation was reading a declaration as an absence.

   extractToolsLine read one line, so a YAML list-form tools: block returned only
   its first item. Two roster files use list form, and gsd-nyquist-auditor
   declares Write and Edit there - parsed as "- Read", found no write tool, and
   emitted read-only. Rung 3's headline claim is that sandbox_mode derives from
   the declared tool contract; that claim was false for 2 of 35 roles and
   materially wrong for 1. Reading a declaration as an absence is the silent-drop
   class this epic exists to close.

   Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now
   derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per
   the standing derive-and-hold decision - so emitted TOML stays byte-identical
   at 26 read-only / 9 workspace-write while the hold list finally records every
   role that would widen. A previous pass declined this fix because it moved the
   count; that inverts the priority. Byte-identity is preserved THROUGH the hold,
   not by leaving a parser broken.

   Divergence check, because this is where that bug hides: both paths feeding
   sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now
   route through the one extractor. The tools readers in
   runtime-artifact-conversion and install.js's other frontmatter call sites
   serve Claude-side emission and do not feed sandbox derivation.

Also fixed, each real: the posture check's `found` used a naive whole-file regex
where its own sibling uses the block-aware scanner, so prose inside
developer_instructions produced a false violation; `found` skipped
truncatePostureValue and leaked a 300-char value into validate agents output;
deriveCodexSandboxMode's absolute never-throws claim was false for an object with
a throwing toString; T49 could not falsify cross-phase leakage (its target phase
had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded
table and pinned the FIXTURE size, so a 36th agent would be silently unchecked;
three tests reimplemented the code they were testing instead of importing it; and
T2-T4 deleted GSD_RUNTIME without restoring it.

Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and
emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer
resolves while ["01"] still does, both identity-bypass cases and every confusable
vector emit read-only.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the hold list is 17, and the reason the 17th was missing

The count read 16 because the derivation could not read the declaration it
claimed to derive from: the tools reader was single-line, so a YAML list-form
tools: block returned only its first item and gsd-nyquist-auditor's declared
Write and Edit were read as an absence.

Both the ADR entry and the feature fragment now carry the corrected count and the
reason for it, rather than a silently updated number. Deriving from a declaration
you cannot parse is not deriving, and a flattering count is worse than a wrong
one because it looks settled.

docs/FEATURES.md regenerated from the fragment.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3897): backfill changeset pr number

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 15:19:01 -04:00
Tom Boucher
1e67ec9737 enhance(#3908): the scanners distinguish an empty diff from one they could not compute (#3937)
* feat(#3908): the scanners distinguish an empty diff from one they could not compute

collect_files ended 2>/dev/null || true, which destroyed the evidence three ways: the redirect discarded git's diagnostic, the pipe replaced git's status with grep's, and || true forced success regardless. Four distinct conditions - an established-empty diff, a bad ref, no repository, and a repository with no commits - all reported clean, and a secret scanner reporting clean because git failed is indistinguishable from an all-clear to any gate consuming it.

git now runs separately from the filter so its status and diagnostic both survive. An established-empty diff exits NO_INPUT; a scope that could not be established exits UNAVAILABLE; the usage sites move off 2 to USAGE. || true is retained on the filter alone, where it is correct: a diff of only images is empty, not failed.

Codes are sourced from a generated shell fragment rather than written into three scripts, so a re-allocation cannot desync them, and a missing fragment fails loudly instead of falling back to literals. The security workflow is updated in the same change: without it, a docs-only PR would newly fail the job.

* fix(#3908): keep scanner stderr out of the file list, and drop try/finally from test bodies

Capturing git and find output with 2>&1 was right for the failure path but wrong for the success path: a warning emitted alongside a successful diff flowed into the file list and was treated as a filename. stderr is now captured separately, forwarded as a warning on success and as the diagnostic on failure, and never folded into the list.

Also converts the control tests' try/finally blocks to t.after(), which CONTRIBUTING bans inside a test body because it masks failures.

* chore(#3908): backfill changeset pr number

* docs(#3908): record the scanners' four-outcome exit contract

SECURITY.md is root-level, so the docs gate correctly held: a Changed fragment owes a file under docs/. The contract also belongs where the feature is described, as REQ-SCAN-INJ-05.

docs/FEATURES.md is GENERATED from per-feature fragments (#3840) - the first edit went into the generated file and gen-features --check caught it, which is the same edit-the-output drift this epic exists to close. The fragment is the source; FEATURES.md is regenerated.

---------

Co-authored-by: sim <sim@local>
2026-08-27 13:11:13 -04:00
Tom Boucher
929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00
Tom Boucher
e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00
Tom Boucher
fb2d122d7f feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb

only this package publishes. The path-based branches — a project-local install,
a runtime config directory — had no such guarantee; they trusted their
configured location. This closes them.

Mechanism: once resolution finishes, and before any verb runs, the preamble
probes the tool it picked with `runtime-identity --raw` and matches the answer
with a shell `case` pattern ANCHORED to the start of the compact payload
(`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts
the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which
any colliding package could publish. The outcome is exported as the two-valued
`GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE
rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling:
`unverified` prints one line naming BOTH causes and continues, because
`no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core`
older than the verb, and at rollout the old-version case is the common one.

The blocker was byte budget, not design. The preamble is inlined into 112
shipped files and several sat within single-digit bytes of frozen ceilings
(`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a
first attempt broke five of them. What made room was collapsing the resolver's
twenty near-identical `elif [ -f … ]` arms into one candidate-list helper
(`_gsd_at`), which buys far more than the assertion costs. The preamble is now
2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every
capped file moved away from its ceiling rather than toward it. No cap raised, no
size-budget exception added, no override token emitted.

Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source
fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all
preserved byte-for-byte in substring terms; the snippet still begins with
`_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor
on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal.

Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both
described an `[ -x ]` guard as the load-bearing re-source defense. That guard
was tried and REMOVED in #3831 — it rejected the bare function name, fell
through every branch, and hit `exit 1`, which kills a sourced caller's shell.
`unset -f gsd_run` is the actual mechanism.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3841): pair the anchor's brace by requiring a closed identity payload

The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md
has unbalanced braces: net depth 2" — plus a knock-on report from its parent
`bug #1516` describe, which is the same failure counted once at the child and
once at the block.

Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and
increments on `{`, decrements on `}`, with no awareness of shell quoting. It
scans `new-project.md` PLUS every `new-project/steps/*.md`, and both
`new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy
— hence net 2 from a snippet that was off by exactly one. The unpaired brace was
the `{` inside the single-quoted `case` pattern of the identity anchor, which is
correct shell and invisible to a text scanner.

Fix in the snippet, not the guard. The pattern now anchors at BOTH ends:
`'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that
does real work rather than a cosmetic pair — a truncated payload whose prefix
matches now fails too, where before it verified. Safe for any future additive
field: a JSON object's own closing brace is always the last character, whatever
type the last value has, which is pinned by two negative-space tests (a nested
object and an array-valued last key must both still verify). Cost: +3 bytes,
against the 1,873 the resolver fold already gave back.

The alternative considered and rejected was dropping the literal `{` for a `?`
glob. It balances too, but weakens the anchor from "must be an opening brace" to
"must be any one character", and the anchor is the entire point.

Two guards added so this cannot recur silently:
- runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next
  edit to that pattern fails on the file it broke instead of surfacing three
  files downstream in a test whose name mentions neither the launcher nor this
  issue. It also asserts depth never goes negative, since a `}` preceding its
  `{` nets to zero while being unbalanced at every prefix.
- runtime-identity gains behavioral truncated-payload and trailing-garbage
  fixtures, so the added `}` is proven load-bearing rather than merely present.

Verified: snippet 51/51 braces; new-project combined net depth 0; the seven
other preamble-bearing files with nonzero depth are unchanged from merged next
(their own prose, not the preamble, and not in any guard's scan set); all 112
inlined copies and the resolver reference re-synced byte-equal; sync:launcher
idempotent on the second run.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3841): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 01:05:53 -04:00
Tom Boucher
36375513b9 feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments

docs/FEATURES.md was hand-maintained, and every feature PR wrote into two
shared mutable cells: the '### N.' heading whose integer was hand-allocated at
authoring time, and the hand-maintained table of contents. Concurrent PRs all
picked the same next integer, and two PRs adding differently numbered features
still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across
successive rebases, each collision also costing a full matrix verification run
because the sha-keyed pass marker dies with the rebase.

Mechanism: one fragment per feature at docs/features/<slug>.md carrying
id/title/group (and an optional order) in frontmatter, consolidated by
scripts/gen-features.cjs --write|--check into a marker-delimited region of
docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings
and their order are derived too - a group sorts by its lowest-ordered member -
so there is no shared registry to edit either; optional per-group prose lives in
docs/features/_groups/<slug>.md. A contributor adds exactly one new file.
Wired into regen:derived and lint:generated-sync alongside the eight existing
generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting.

Migration froze all 168 existing numbers verbatim: identical section set,
identical order, identical bodies. Two defects found in the tree are fixed
inline rather than carried forward - the '## Related' block had been spliced
into the middle of the document, orphaning §142's Reference line, and four
inbound anchors were already broken on next (FEATURES.md#runtime-identity in
two files, and #143-spec-phase-edge-completeness-probe off by one). Since the
repo has no link checker, --check now validates every inbound
FEATURES.md#anchor by resolved target, so that class cannot ship silently
again; locale FEATURES.md files resolve elsewhere and stay out of scope.

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3840): carry upstream §69 delta into its fragment and harden the generator

Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus
origin/next. Root cause was a stale base, not extraction loss: those lines
landed in 394bf384b (#3844) AFTER this branch forked at 63abcface, and
'git diff 63abcface origin/next -- docs/FEATURES.md' is exactly that hunk.
Merging origin/next auto-applied the hunk into the GENERATED region, which
--check immediately reported as stale; the delta is now carried in
docs/features/statemd-consistency-gates.md and regenerated from there.

--write is now fail-closed. It previously rendered the region even with
violations outstanding, warning only on stderr and exiting 0, so a
'--write && git commit' chain could commit a FEATURES.md carrying two
colliding sections. It now refuses and exits 1; --force is the explicit
override and says so in the report. The test that pinned the old behavior now
pins the refusal, plus the --force override and its scoping.

Marker forgery is rejected at two layers. A fragment body containing
'<!-- FEATURES:START' or '<!-- FEATURES:END' is a typed
body_forges_region_marker violation (fragments and group notes alike), and
spliceIntoFeatures anchors the end boundary with lastIndexOf instead of
indexOf, so a marker that reaches the document by any other route can only
make the generated region grow, never shrink. Matching is on marker PREFIXES,
so a decorated variant comment cannot slip past.

Symlinked corpus entries are refused with a typed dirent_not_regular_file
rather than read. A fork PR could otherwise commit docs/features/evil.md as a
symlink to any readable path and have the generator inline those bytes into
the committed docs/FEATURES.md on the next regen.

Equivalence re-verified with a method that cannot cancel out. The first
check extracted both operands with the same body-normalising helper, so
anything that helper dropped was dropped on both sides. The replacement runs
two independent passes: a global content-line multiset diff with no
per-section logic at all (0 gained, 19 lost, all 19 the stale hand-written
mini-TOC links this change deliberately deletes), and a per-section
byte-exact body diff carrying a coverage assertion that fails loudly per file
when the extractor accounts for fewer lines than the file contains. That
assertion caught two blind spots in the checker itself. 168/168 sections
present, order identical, one intended body difference (§142 regains the
Reference line orphaned by the misplaced '## Related' block).

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3840): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 22:49:00 -04:00