8b41d855e002ee7deffd90f162858b9f435f9318
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
03b7125293 |
enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks Binds the four fabrication sites found by executing the surfaces (ADR-3889 failure class (c)), each with a positive control so an over-firing fix goes red: - the blocking api-coverage.verify-pre gate certifying "no external-API integration" from a zero-byte phase scope - the assumption-delta query route scanning an unresolvable phase section as the empty string and reporting it as an examined negative - both capability fragments' probe fallbacks, which append a fabricated verdict rather than replacing, and fire on the legitimate exit-1 negative Verification runs on the remote runner. Refs #3909 * enhance(#3909): a probe that could not run no longer asserts a verdict ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a confident negative; each now reports what it could not establish. - check api-coverage.verify-pre: a phase with no plan body and no roadmap section ran detection over zero bytes and PASSED the blocking seal gate, certifying "no external-API integration" from input it never read. It now holds with scope_unavailable. The discriminator is bytes examined, never signals found, so a phase whose plans are real and simply carry no API vocabulary passes exactly as before. - query assumption-delta scan: an unresolvable phase section was scanned as the empty string and reported as an examined negative. It now returns {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded result in the payload, leaving the gsd-tools exit projection to P8. - both capability fragments: `|| echo '{"detected":false}'` appended rather than replaced, and fired on the legitimate exit-1 negative, so a correct answer and an honest skip both arrived as two concatenated objects. They now keep the probe's own payload and manufacture only an explicit probe_unavailable skip when the probe produced nothing at all. Every registered outcome is more restrictive on a blocking gate, so this can turn a false green red and never a red green. Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md seal-time outcome table, and a new how-to for the reason-code vocabulary. Verification runs on the remote runner. Closes #3909 * test(#3909): correct the stale unknown-phase assertion `unknown phase → detected:false, no throw (graceful)` scanned phase 999 against a two-phase roadmap and asserted `detected === false`. That pinned the fabrication as intended behavior: the phase does not exist, so the detector was handed the empty string and its "no core assumption changed" answer described nothing that was ever read. It now asserts the skipped-with-reason shape. The graceful-degradation contract the test was actually protecting — the query succeeds and does not throw on an unknown phase — is unchanged. Found by code review, not by the author. Refs #3909 * docs(#3909): author the FEATURES entry in its generator source `docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the per-feature fragments in `docs/features/`. The API-coverage entry was edited in the generated file, so the next regeneration silently dropped it. The text now lives in `docs/features/api-coverage-gate.md` and `docs/FEATURES.md` is regenerated from it, leaving the shipped file byte-identical and its content actually derivable. Caught by `lint:generated-sync`. Refs #3909 * test(#3909): bind the skip to "not found", and pin the discriminator The first verification run went red on one case, and the case was wrong rather than the code. `getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a missing ROADMAP.md, but for a section whose body is whitespace-only it returns the heading line alone — which is not empty. So a body-less section WAS found, and reporting `detected:false` over its heading is a real negative, not a fabrication. The test had assumed the resolver yielded `''` there. Correcting the test rather than the resolver keeps `skipped` bound to the distinction the issue asks for — found versus not found — and avoids diverging `assumption-delta scan` from `roadmap.get-phase`, which the fragment documents as sharing one resolver. Also adds the seeded property the test matrix had promised: for any plan body, the scope read back is whitespace-only exactly when the body was. That pins the gate's discriminator to bytes examined, so it cannot quietly become "no signals found", across unicode whitespace and CRLF. `docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table — surfaced by the co-change gate, not by a lint failure. Refs #3909 * chore(#3909): backfill the changeset PR number Refs #3909 --------- Co-authored-by: sim <sim@local> |
||
|
|
9410f7e6e6 |
enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4 ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the other three RED before any fix. Rung 2 — the install marker has four readers and resolveRuntime is not one. resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule. Fixtures and seam names mined from PR #3382 rather than re-derived; it implemented this rung and was closed "not on the merits". Rung 3 — the sandbox map, and the fallback that was the real defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements — the map carries nothing the contract does not - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit So the map is redundant and the silent fallback is the defect. The maintainer chose to derive but hold those 16 at read-only pending the question of whether Codex enforces sandbox_mode or merely advises; HALT.md records it. T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not in aggregate — an aggregate passes while one role silently widens, which is the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a stale hold, so the hold list cannot rot into the subset map being deleted. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId` returns only documentation commits. That was the wrong instrument. Direct inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and the tests match that code rather than a guess at its semantics — including first-write-wins on a duplicate short form. T43 asserts at the consumer's output: the emitted `waves` map from the real CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is dropped. A unit assertion on resolveDependencyId would have passed throughout this defect's life. Observed RED, this tree: rung 2 11/11 fail — no marker rung, no seams rung 3 T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML whose sandbox_mode disagrees — it checks presence only) rung 4 T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1 Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's unresolvable-token warning and wave-verdict suppression, and #3785's display-mapping passthrough. If the third tier over-reaches, those go red — that is their job. Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed) cannot be isolated behaviorally, because planMap always masks it. It is a non-crash boundary pin, weaker than the other rows, and is recorded as such rather than presented as equivalent. Design: .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md Decision: .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the other three. Rung 2 — the install marker had four readers, and resolveRuntime was not one. resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no marker, while bin/install.js writes one (#2297) and four hand-rolled readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with seams), hooks/gsd-agent-isolation-guard.js, and twice in hooks/gsd-cursor-subagent-start.js. model-resolver's was already the house idiom, so it was promoted rather than replaced: src/runtime-slash.cts now owns it, and model-resolver plus both hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam lint-hooks-runtime-build-seam enforces. No import cycle existed - checked both directions before moving anything. The marker is the THIRD rung: env > project config > marker > 'claude'. N1 was checked rather than assumed, and my first reading of it was wrong. A marker holding an unknown name comes back essentially verbatim, which looked like a validation gap. Measured against the env rung with the same inputs - including "../../etc/passwd" and "claude;rm -rf /" - the two are identical, because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that, and it is met. The residual (the shared normalizer normalizes shape, it does not validate against the known-runtime set) is pre-existing on the env rung and plausibly deliberate, since a new runtime should not need a code change. The marker also does not widen the trust boundary in any real sense: it lives inside the install tree beside the code, so anyone who can write it can write runtime-slash.cjs itself. Rung 3 — the map was redundant; the silent fallback was the defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements. The map carried nothing the contract did not already have, so it is DELETED rather than reconciled. What was actually broken is `|| 'read-only'`, which silently under-granted 24 of 35 roles. 16 of those 24 declare Write or Edit and would widen under derivation. Per the maintainer's decision (HALT.md), they are held at read-only pending the question of whether Codex enforces sandbox_mode or merely advises. Emitted TOML is therefore byte-identical for all 35 roles - asserted per role, not in aggregate, because an aggregate passes while one role silently widens. The hold list self-invalidates. A hold whose role no longer derives broader fails, and so does a hold naming a role with no agents/<name>.md. Without that it would rot into exactly the hand-maintained subset map being deleted, and this commit's own ledger claim would become false over time. Both cases were proved by injecting them and watching them throw. Two committed tests asserted the deleted map's existence and contents. They were pinning the thing being removed, so the tests moved rather than the production code: the 11 role-value pairs survive as a test-local PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real derivation against real agents/*.md. The coverage is preserved; only its source moved out of production code. validate agents gains checkCodexSandboxPosture, mirroring the existing checkCodexModelPosture: each installed TOML's sandbox_mode must equal the role's expected value, failing with role, expected and found. It previously checked file presence and manifest completeness only, so a TOML whose sandbox_mode disagreed passed. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: git log -S returns only documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^ carries five occurrences, and the implementation here matches that code rather than a guess at its semantics - including first-write-wins on a duplicate short form, deterministic from the sorted plan order. It resolves the bare plan number: depends_on: ["01"] now reaches 26-01-auth-hardening. That is a control-flow change, not a diagnostic one - plans that silently collapsed into a single wave 1 now execute in their declared waves, and execute-phase.md consumes those wave values. In-phase only, by construction: the map is built from this phase's rawPlans, so a same-named short form in another phase does not resolve. #3785's display-mapping passthrough and #3885's unresolvable-token warning and wave-verdict suppression are untouched and stay green. If the third tier had over-reached, those are what would have caught it. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open I introduced, and wire the posture check to its command Two blockers from review. Both are mine, and one is a security regression my own change created. 1. A held role could escape its hold by editing its own frontmatter. The Codex install loop set the sandbox identity from the agent's frontmatter `name:` field rather than from its filename, so the hold lookup keyed off a value the file itself declares: deriveCodexSandboxMode('gsd-doc-writer', <real file>) -> read-only deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write deriveCodexSandboxMode('GSD-Doc-Writer', <same file, name: recased>) -> workspace-write What makes this a blocker rather than a nit is the DIRECTION. The deleted CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an allowlist: an unmatched key fell back to read-only, which is safe. The new scheme derives workspace-write from the tool contract and uses the hold as a subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk into a fail-open one and did not notice; the isolated reviewer proved it by execution. Neither safety net caught it. validateCodexSandboxHolds only checks that <key>.md exists, never that a file's derived identity matches its key. checkCodexSandboxPosture looks the canonical source up by the installed TOML's filename, finds nothing for a renamed agent, and treats it as a custom non-roster agent — silently no violation. The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds already validates and what an attacker editing frontmatter cannot change without renaming the file — at which point the existing validator catches it. The lookup is case-insensitive so a recase does not slip past either. The frontmatter name still drives the TOML body and filename, unchanged; only the sandbox identity moved. All 35 roster files were checked: name matches filename stem everywhere, so a stricter "they must agree or throw" invariant would have been safe against real content. It is deliberately NOT added — it would abort an install on a tampered file where emitting a correctly-derived read-only TOML is the safer outcome. Recorded as a fork rather than decided silently. 2. checkCodexSandboxPosture was exported and never called. cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and checkCodexModelPosture only; grep for the sandbox check in that file returned nothing. So criterion 3 — "validate agents fails on semantic drift, not only on missing files" — was unmet, and `validate agents` behaved exactly as before. That is ADR-3473 Decision 2's named shape: a declared policy with no executor. It also meant the T28 test asserted at the helper's return value while the COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to close, committed by me while enforcing it elsewhere in the same epic. Now wired as an additive `sandbox_posture` field beside `codex_posture`, following the sibling precedent exactly. Drift is report-only, not a non-zero exit, because that is what checkCodexModelPosture does — two sibling posture checks disagreeing about whether a violation is fatal would be its own defect. The choice is recorded in a comment rather than left implicit. A consumer-output test now drives the real CLI and asserts on the emitted JSON, and was shown failing before the wiring and passing after. Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3 was not in this deliverable, written while it was halted and false once the maintainer unblocked it. Verified after both fixes: the three bypass probes all return read-only, the per-role table is 35/35 byte-identical, and both hold self-invalidation cases still throw. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture field; docs/reference/plan-md.md documents that depends_on accepts the bare plan number. Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a concurrent PR hand-allocating a section number, regenerated into FEATURES.md. ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style, recording what was measured and built against the section's 2026-08-26 correction - including the qualification that checkAgentsInstalled itself still checks presence only, and the semantic assertion lives in a sibling wired into validate agents rather than folded into it. No how-to. Both user-visible changes are zero-step: a non-Claude install resolving its own runtime, and plans executing in their declared waves, both happen without the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does not touch, and was deliberately left alone rather than edited by association. No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it documents Claude-side tools frontmatter, never Codex sandbox_mode, and the emitted tools contract did not change. The prompt layer documents depends_on only by example, not by schema, so nothing there needed editing - and few-shot-examples/plan-checker.md already showed depends_on: ['01'], which now actually resolves. Translated copies of plan-md.md are untouched; the project treats translations as community-maintained. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser The full suite came back with 26 failures across four files. Three distinct causes, mapped individually rather than assuming the first explained the rest. A. Requiring bin/install.js printed the GSD banner to stdout and corrupted `validate agents` JSON. Unexpected token '', "[36m ██"... is not valid JSON checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring bin/install.js, whose module load prints the ASCII banner. So the command emitted banner bytes before its JSON and every JSON consumer broke, including ten tests that predate this branch. src/ reaching into bin/ was backwards layering that happened to also be loud. The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML domain module, no new module and no six-gate ripple - and both bin/install.js and src/agent-install-check.cts import it. One owner, which is §8.3's rule applied to the fix for §8.3. B. The stale-hold throw fired on a legitimate partial source dir, and masked a security assertion. validateCodexSandboxHolds treated "this hold's .md is absent from the install SOURCE dir" as a stale hold and threw. A test fixture, or any partial install source, legitimately contains a couple of agents. Worse, it threw BEFORE the path-escape check, so a test asserting that a `../../evil` frontmatter name is rejected got my unrelated error instead of the traversal rejection it was written for. A fail-closed check of mine was hiding a real security check. The "no stale holds, shrink-only" invariant is a property of the repo's canonical agents/ roster, not of whatever directory an install happens to read. It is off the runtime path and enforced where it belongs, in the tests that already existed for it. A partial source dir now installs cleanly, and the evil-name case throws with its own escapes-configHome message again. C. T8 depended on ambient process.env state. The marker/env parity assertion round-tripped through live process.env. It now compares against resolveExplicitRuntime's already-exported dependency-injection parameter - deterministic and hermetic, same claim. Proven still falsifiable rather than assumed: with the marker rung's normalization temporarily bypassed the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd") and the assertion fails, then passes again once reverted. One correction folded in along the way. The first version of the move added private _extractFrontmatterAndBody/_extractFrontmatterField helpers to codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893). Adding a third inside the epic whose thesis is one implementation per rule is not defensible. deriveCodexSandboxMode no longer parses anything: it takes (identity, toolsValue) and each caller supplies the tools value using the extractor it already has. Both helpers are deleted. The identity argument is still the filename stem, so the fail-open fix is untouched. Verified after all three: `validate agents --raw` emits parseable JSON with no banner and both posture fields; the four hold-bypass probes still return read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9 workspace-write; the hold list is still 16. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test Suite down to 7 failures from 26. Three more causes, mapped individually. A. My extractor import dragged in a script that does not exist in an installed tree. Cannot find module '../../../scripts/fix-slash-commands.cjs' Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs, which requires command-roster.cjs, whose line 36 requires ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not in an install, so every test exercising a synthetic install dir died at module load. I picked that extractor for convenience without checking what it pulls in - the same mistake that produced the banner bug, one layer further out. agent-install-check now uses a single-purpose extractToolsLine on codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser: we deleted those helpers a commit ago for good reason, and this reads one line. Verified from outside the repo root that requiring either module prints nothing and does not throw. B. A test pinned the deleted name-based fallback. 'defaults unknown agents to read-only' called generateCodexAgentToml with a fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent with a writing contract correctly derives workspace-write - design row S6, a new writing role gets the contract, not the pin. The behavior it asserted was the silent fallback this rung deleted; identity no longer decides the sandbox. Replaced with two rows rather than a flipped string: no tools declared -> read-only (absence is not a grant), and Write/Edit declared -> workspace-write. Strictly more coverage than the row it replaces. C. The stale-hold check still threw per derivation call. Last commit took the roster-existence check off the install path, but deriveCodexSandboxMode itself still threw when a hold's role did not derive broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture for a held role. The throw is gone, and it cost nothing: if a held role's content does not derive broader, the hold pins read-only and derivation returns read-only anyway, so the hold is a no-op and there is nothing to fail about. The staleness invariant is a property of the real agents/ roster, and validateCodexSandboxHolds still enforces it there - confirmed against the real roster after the change, not assumed. deriveCodexSandboxMode is now total: every (identity, toolsValue) including undefined and null returns read-only or workspace-write, never throws. Verified: validate agents emits parseable JSON; the four hold-bypass probes return read-only; the per-role table is 35/35 at 26 read-only / 9 workspace-write. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path The ADR entry and the feature fragment both ended their rung-3 explanation with "see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at read-only was reachable only from the machine that produced it. A reader of the ADR got a pointer to nothing. Both now carry the reasoning inline: the criterion asks both that the sandbox derive from the declared tool contract and that no role gain a broader sandbox, and those cannot both hold, because a faithful derivation widens 16 roles the deleted map never listed and that fell through its silent read-only default. The resolution is derive-and-hold - the derivation owns the rule now, each hold is released as its enforcement question is answered, and a hold is reversible where a widened sandbox that turns out to be enforced is not. Checked before assuming this was a defect class: CONTEXT.md cites .gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight module entries, and four other shipped docs do the same. Citing a phase artifact is an established convention here, so those are left alone. What was wrong was specific to these two: they put load-bearing rationale behind the pointer instead of provenance. docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs rather than hand-edited. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration Two orthogonal reviews on the shipped sha. Three of the findings are the same failure class this epic exists to close, committed inside it. 1. BLOCKER - the sandbox was decided for one identity and applied to another. bin/install.js derived sandbox_mode for the filename stem and then wrote the result to `${name}.toml`, where name comes from the file's own frontmatter. Make the two disagree and a HELD role's artifact goes wide: rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml add any gsd-*.md whose frontmatter name: is a held role -> clobbers that role's toml with workspace-write Both emit read-only on origin/next, because the deleted map was an allowlist and a miss fell back safe. This is a regression my change introduced. The previous review round moved the HOLD KEY off frontmatter to the filename stem and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985 calls that value attacker-editable, four lines above the line that uses it as the filename. The decision is now made over BOTH candidate identities, most-restrictive wins: if either the stem or the emitted name is held, the mode is read-only. 2. MAJOR - hold matching was toLowerCase() only, so confusables escaped. Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline, ./ and ../agents/ all slipped the hold and emitted workspace-write. Identities are now basenamed, trimmed of NBSP/zero-width/control characters, NFKC-normalized and lowercased - and anything still carrying a character outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not enumerate confusables; every shipped roster file is ASCII, so refusing to widen on an identity we cannot recognize is fail-closed with no false positives on real content. 3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY. shortFormToId keyed on the last dash-segment of any canonical id with no constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the worst shape in the epic: the unresolvable-token warning fires on a DROPPED token, so a MIS-RESOLVED one is invisible and the tool reports a confident wave assignment built from a wrong edge. A wrong edge is worse than a missing one. The segment must now match /^\d+$/, which is exactly the contract docs/reference/plan-md.md already documents. This tier was recovered verbatim from the retired SDK lineage, which carried the same defect; we are deliberately NOT preserving it bug-for-bug, and the comment says so, so the next reader does not "restore" it. 4. MAJOR - the derivation was reading a declaration as an absence. extractToolsLine read one line, so a YAML list-form tools: block returned only its first item. Two roster files use list form, and gsd-nyquist-auditor declares Write and Edit there - parsed as "- Read", found no write tool, and emitted read-only. Rung 3's headline claim is that sandbox_mode derives from the declared tool contract; that claim was false for 2 of 35 roles and materially wrong for 1. Reading a declaration as an absence is the silent-drop class this epic exists to close. Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per the standing derive-and-hold decision - so emitted TOML stays byte-identical at 26 read-only / 9 workspace-write while the hold list finally records every role that would widen. A previous pass declined this fix because it moved the count; that inverts the priority. Byte-identity is preserved THROUGH the hold, not by leaving a parser broken. Divergence check, because this is where that bug hides: both paths feeding sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now route through the one extractor. The tools readers in runtime-artifact-conversion and install.js's other frontmatter call sites serve Claude-side emission and do not feed sandbox derivation. Also fixed, each real: the posture check's `found` used a naive whole-file regex where its own sibling uses the block-aware scanner, so prose inside developer_instructions produced a false violation; `found` skipped truncatePostureValue and leaked a 300-char value into validate agents output; deriveCodexSandboxMode's absolute never-throws claim was false for an object with a throwing toString; T49 could not falsify cross-phase leakage (its target phase had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded table and pinned the FIXTURE size, so a 36th agent would be silently unchecked; three tests reimplemented the code they were testing instead of importing it; and T2-T4 deleted GSD_RUNTIME without restoring it. Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer resolves while ["01"] still does, both identity-bypass cases and every confusable vector emit read-only. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the hold list is 17, and the reason the 17th was missing The count read 16 because the derivation could not read the declaration it claimed to derive from: the tools reader was single-line, so a YAML list-form tools: block returned only its first item and gsd-nyquist-auditor's declared Write and Edit were read as an absence. Both the ADR entry and the feature fragment now carry the corrected count and the reason for it, rather than a silently updated number. Deriving from a declaration you cannot parse is not deriving, and a flattering count is worse than a wrong one because it looks settled. docs/FEATURES.md regenerated from the fragment. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3897): backfill changeset pr number Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1e67ec9737 |
enhance(#3908): the scanners distinguish an empty diff from one they could not compute (#3937)
* feat(#3908): the scanners distinguish an empty diff from one they could not compute collect_files ended 2>/dev/null || true, which destroyed the evidence three ways: the redirect discarded git's diagnostic, the pipe replaced git's status with grep's, and || true forced success regardless. Four distinct conditions - an established-empty diff, a bad ref, no repository, and a repository with no commits - all reported clean, and a secret scanner reporting clean because git failed is indistinguishable from an all-clear to any gate consuming it. git now runs separately from the filter so its status and diagnostic both survive. An established-empty diff exits NO_INPUT; a scope that could not be established exits UNAVAILABLE; the usage sites move off 2 to USAGE. || true is retained on the filter alone, where it is correct: a diff of only images is empty, not failed. Codes are sourced from a generated shell fragment rather than written into three scripts, so a re-allocation cannot desync them, and a missing fragment fails loudly instead of falling back to literals. The security workflow is updated in the same change: without it, a docs-only PR would newly fail the job. * fix(#3908): keep scanner stderr out of the file list, and drop try/finally from test bodies Capturing git and find output with 2>&1 was right for the failure path but wrong for the success path: a warning emitted alongside a successful diff flowed into the file list and was treated as a filename. stderr is now captured separately, forwarded as a warning on success and as the diagnostic on failure, and never folded into the list. Also converts the control tests' try/finally blocks to t.after(), which CONTRIBUTING bans inside a test body because it masks failures. * chore(#3908): backfill changeset pr number * docs(#3908): record the scanners' four-outcome exit contract SECURITY.md is root-level, so the docs gate correctly held: a Changed fragment owes a file under docs/. The contract also belongs where the feature is described, as REQ-SCAN-INJ-05. docs/FEATURES.md is GENERATED from per-feature fragments (#3840) - the first edit went into the generated file and gen-features --check caught it, which is the same edit-the-output drift this epic exists to close. The fragment is the source; FEATURES.md is regenerated. --------- Co-authored-by: sim <sim@local> |
||
|
|
929e02cb2c |
enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking answer. Three families do exactly that today; this commit pins each one RED. Measured on this tree, 2026-08-27: intel query, .planning/intel/file-roles.json nested 12000 deep -> exit 1, "Error: Maximum call stack size exceeded" searchJsonEntries / matchesInValue carry no depth parameter at all. The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it. same fixture nested 48 and 49 deep -> both return total=1 at exit 0, truncated=undefined Nothing distinguishes "searched to the bottom" from "stopped looking". query phase-plan-index, a plan whose depends_on names an unresolvable token -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] The token is never mentioned. computeDependencyLevels drops the edge with `if (!resolvedDep) continue;`, every plan becomes a root, and the tool then reports the author's correct wave: as the thing that is wrong. countPhasePlansAndSummaries with fs.readdirSync throwing EACCES -> hasContext:false, indistinguishable from a phase that simply has no CONTEXT.md. context_read_error is undefined. The shapes these tests assert against, chosen here so the implementation has a target rather than inventing one later: `truncated: boolean` on the intel query result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and `context_read_error: string | null` per analyzed phase. Deliberately green, and they must stay that way — each stops the fix from over-firing: depth 48 is found and NOT flagged truncated (the ceiling is inclusive) a shallow miss reports no truncation (noise control, N1) 10,000 siblings at depth 2 are unaffected (the bound is DEPTH, N2) a genuine wave: mismatch on a fully-resolved DAG still warns (N3) a genuinely missing directory is absent, not an error the emitted depends_on display mapping still passes an unresolved token through verbatim — already pinned by the existing #3785 test, so no duplicate was added T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the real CLI and reads the emitted JSON, because a unit assertion on computeDependencyLevels would have passed throughout #3427's life. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3885): no silent swallow, and no verdict manufactured from dropped data Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed into an output that reads as authoritative. The recursion bound, restored but NOT verbatim (src/intel.cts) MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and matchesInValue, which carried no depth parameter at all. The bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it — §8.3's "a consolidation may not delete an invariant along with the surface that held it", demonstrated. Measured before: a .planning/intel file nested 12000 deep exits 1 with "Error: Maximum call stack size exceeded". Reachable from a project document. The original returned a bare `false` at the ceiling. Restoring that verbatim would trade a crash for a silent "no match" when the truth is "I stopped looking" — the same class this epic exists to close, and ADR-3473 Decision 4 forbids it. So the bound carries a truncation signal: nesting 47 -> found, truncated false nesting 48 -> found, truncated false (the ceiling is inclusive) nesting 49 -> not found, truncated TRUE nesting 12000 -> exit 0, truncated TRUE, no RangeError A shallow document that simply has no match reports truncated FALSE — the flag means "I stopped early", never "I found nothing", or it would be noise. The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected. The dropped edge is named, and stops being blamed on the author (src/phase.cts) computeDependencyLevels dropped every unresolvable depends_on token with a bare `continue`. Each drop makes a plan a root, so the whole phase collapses to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as the thing that was wrong. Before: warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] After: warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not resolve to any plan in this phase — edge dropped, wave placement for this plan may be unreliable"] The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is deliberately not built here. The emitted depends_on display mapping still passes an unresolved token through verbatim (#3785). No artifact from failed inputs (gsd-core/workflows/review.md, #3352) A failed lane leaves no result file, so "every lane failed" is exactly "the aggregate JSONL has zero lines" — the gate condition already existed as a byproduct. REVIEWS.md is no longer written in that case, and the commit step is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT counted as a failure. Per-lane output and non-empty .err are preserved to .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that the lanes failed at all; the commit step names one file, never a glob, so the diagnostics are not swept in. Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2) Four callers collapsed an EACCES on a phase directory into [] and reported hasContext:false — byte-identical to a phase that simply has no CONTEXT.md. Each now names the directory it could not read. A genuinely missing directory stays absent rather than becoming an error, which is what keeps the fix from over-firing. Fatal errno folded into a retry set: audited, no defect found Reported as a verified negative rather than padded with a change. withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776; atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they return or rethrow the final error rather than swallowing it. estimate-cli's sole caller surfaces that rethrow as write_error in its JSON output. Manufacturing a diff to make the checkbox look worked-on is the Goodhart outcome Decision 6 exists to prevent. Disclosed: R46 (the commit step names one file, never a glob) is a real regression guard but is NOT independently failing-first — the commit fence is byte-identical pre- and post-fix, so it only fails pre-fix through its shared extraction dependency. Recorded rather than claimed as fail-first. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence Two review findings, both real, both in my own change. An isolated adversarial review found the evidence-preservation block never checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in a SEPARATE fenced block. A disk-full or unwritable phase directory therefore still destroyed the only copy of the failed lanes' output — reintroducing the exact #3352 data loss this item exists to stop, inside the fix for it. Preservation and cleanup are now one block, because each fenced block is a separate execution and a shell variable cannot carry between them. mkdir -p and each cp are exit-checked; cleanup runs only when preservation succeeded, and a failure warns naming the intact run directory. "Nothing to preserve" is not a failure and still cleans up. Driven three ways: success removes run_dir, failure leaves it intact with the warning, nothing-to-preserve removes it. The failure is induced by a file-vs-directory conflict rather than chmod 0o000, which root bypasses. The new unresolved-depends_on warning embedded a user-authored token verbatim: warnings: ["Plan 03-02: depends_on token \"evil Plan 03-01: FORGED WARNING\" does not resolve ..."] The JSON wire form is safe, and the security reviewer judged it non-exploitable for that reason. It is escaped anyway through formatDiagnosticToken — the helper #3884 added one phase earlier for exactly this class. warnings[] is an array a consumer naturally prints line by line, and not reusing the sibling fix is the generative-fix-divergence shape this epic exists to close. The same treatment is applied to context_read_error / phase_dir_read_error, which embed a phase directory path a repository can choose, and to the fs error message, which echoes the raw path itself. Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per negative space N2, disclosed rather than left to be discovered. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot" Blocker from the round-2 isolated review, and it is my own inconsistency: this phase applied "unreadable is not absent" to phase directories and left it broken in the file it was already editing. chmod 000 .planning/intel/file-roles.json gsd-tools intel query <term> -> {"matches":[],"total":0,"truncated":false} exit 0 safeReadJson swallowed every read failure and returned null, so an EACCES was byte-indistinguishable from an absent file AND from a genuine no-match. Now it separates three states: ENOENT stays silently absent, because not every project has every intel file and intelQuery loops over all of them expecting misses; EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel file previously read as "no matches" too — same defect, same fix. Threading that outcome through the other three callers found something worse than the reported case. intelDiff returned no_baseline:true for a corrupt or unreadable snapshot — not a silent failure but an actively FALSE verdict, telling the caller they never took a snapshot when they did. That is §8.5's headline case, so it is fixed and tested rather than noted. intelStatus and intelApiSurface collapsed the same way; intelApiSurface additionally printed a "not yet populated" banner that was simply untrue. Every row is failing-first, including the absent-file ones — the field is new, so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin that the fix does not OVER-fire on the ordinary absent case, which is what would turn this into noise on every project lacking an intel file. IO failure is injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as a test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): build the pathological intel fixture as text, not by stringifying a nested object The remote runner came back red on Linux with two failures, both T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on macOS. The product was never at fault. writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned. Linux's container stack is smaller than macOS's, which is the whole of the platform difference. Measured, with the same document built as JSON TEXT so nothing in the building process recurses: depth=100 rc=0 truncated=true depth=5000 rc=0 truncated=true depth=12000 rc=0 truncated=true depth=60000 rc=0 truncated=true V8 parses this shape iteratively; only stringify recurses. The bound works at every depth tried. The fixture is now built by string concatenation. That is also the more faithful input — a real deeply nested JSON document on disk is exactly what the bound guards, where a stringified object was only ever a way to produce one. The depth stays 12000. Lowering it would have made the test pass by weakening it to accommodate a fixture bug, and 12000 is a legitimate pathological input the product handles. T4 remains a genuine fail-first: rebuilt against the parent of the commit that added the bound, the string-built depth-12000 fixture still drives the CLI to rc=1 with "Error: Maximum call stack size exceeded". A comment records why the fixture is text, so it is not "simplified" back into a macOS-green / Linux-red test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3885): backfill the changeset PR number Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): normalize path separators before splicing into the workflow's bash CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the remote runner were all green. AssertionError: commit must name the single REVIEWS.md file; got: --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The harness spliced an OS-native temp path into the extracted bash, and bash consumes \U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory survived — which is the other two assertions. This is a fixture defect, not a product one, and that was checked rather than assumed. In production the phase directory is toPosixPath-normalized at every call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and the run directory is created by "mktemp -d" running inside the bash block itself (gsd-core/workflows/review.md:163), which emits POSIX-style output even under Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it. The file's pre-existing #3034 harness splices raw native paths too, but only ever inside double-quoted assignments, so it never tripped this — my new harness followed that convention faithfully into the one place where it does not hold. Both now splice through toPosixPath from shell-command-projection, the established seam, which is a no-op on POSIX and mirrors what production does. No assertion was weakened. "commit must name the single REVIEWS.md file" and "the run dir must still be destroyed" still assert exactly that; only how the fixture supplies its path changed. Nothing is skipped on Windows — a t.skip() here would have hidden the question of whether the exposure was real, which is the question that mattered. Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input produces a byte-identical shape, proving the normalization is idempotent. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): stop the harness making the deleted run dir its own cwd Windows shard 3/3 stayed red after the separator fix, on two assertions the separator fix never touched: AssertionError: the run dir must still be destroyed AssertionError: nothing to preserve is not a failure — run dir must still be removed The separators were a real bug and fixing them fixed the --files assertion. They were not this bug, and two CI cycles went into the wrong axis before I stopped converting path forms and looked at what the harness actually does. runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's working directory WAS the directory the block under test then removes with rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live process's working directory cannot be removed. So on Windows the directory survived and both assertions failed, on macOS and Linux it vanished and they passed. Nothing to do with slashes. Harness-only. Production never cd's into the run directory; every reference is by absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself (gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged. Fix: the child now runs with its cwd in an unrelated temp directory that the block under test never deletes. Neither assertion was weakened, and nothing is skipped on Windows — the tests in this file carry no platform guard and run there unconditionally, which is how this surfaced at all. Honest limit: the Windows failure mode cannot be reproduced on macOS, because POSIX permits the very thing Windows refuses. The diagnosis is grounded in that documented divergence and in the fact that only the Windows lane failed, but the green outcome on windows-latest is unverified until CI runs it. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e20744eacb |
enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb2d122d7f |
feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb only this package publishes. The path-based branches — a project-local install, a runtime config directory — had no such guarantee; they trusted their configured location. This closes them. Mechanism: once resolution finishes, and before any verb runs, the preamble probes the tool it picked with `runtime-identity --raw` and matches the answer with a shell `case` pattern ANCHORED to the start of the compact payload (`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which any colliding package could publish. The outcome is exported as the two-valued `GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling: `unverified` prints one line naming BOTH causes and continues, because `no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core` older than the verb, and at rollout the old-version case is the common one. The blocker was byte budget, not design. The preamble is inlined into 112 shipped files and several sat within single-digit bytes of frozen ceilings (`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a first attempt broke five of them. What made room was collapsing the resolver's twenty near-identical `elif [ -f … ]` arms into one candidate-list helper (`_gsd_at`), which buys far more than the assertion costs. The preamble is now 2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every capped file moved away from its ceiling rather than toward it. No cap raised, no size-budget exception added, no override token emitted. Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all preserved byte-for-byte in substring terms; the snippet still begins with `_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal. Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both described an `[ -x ]` guard as the load-bearing re-source defense. That guard was tried and REMOVED in #3831 — it rejected the bare function name, fell through every branch, and hit `exit 1`, which kills a sourced caller's shell. `unset -f gsd_run` is the actual mechanism. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3841): pair the anchor's brace by requiring a closed identity payload The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md has unbalanced braces: net depth 2" — plus a knock-on report from its parent `bug #1516` describe, which is the same failure counted once at the child and once at the block. Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and increments on `{`, decrements on `}`, with no awareness of shell quoting. It scans `new-project.md` PLUS every `new-project/steps/*.md`, and both `new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy — hence net 2 from a snippet that was off by exactly one. The unpaired brace was the `{` inside the single-quoted `case` pattern of the identity anchor, which is correct shell and invisible to a text scanner. Fix in the snippet, not the guard. The pattern now anchors at BOTH ends: `'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that does real work rather than a cosmetic pair — a truncated payload whose prefix matches now fails too, where before it verified. Safe for any future additive field: a JSON object's own closing brace is always the last character, whatever type the last value has, which is pinned by two negative-space tests (a nested object and an array-valued last key must both still verify). Cost: +3 bytes, against the 1,873 the resolver fold already gave back. The alternative considered and rejected was dropping the literal `{` for a `?` glob. It balances too, but weakens the anchor from "must be an opening brace" to "must be any one character", and the anchor is the entire point. Two guards added so this cannot recur silently: - runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next edit to that pattern fails on the file it broke instead of surfacing three files downstream in a test whose name mentions neither the launcher nor this issue. It also asserts depth never goes negative, since a `}` preceding its `{` nets to zero while being unbalanced at every prefix. - runtime-identity gains behavioral truncated-payload and trailing-garbage fixtures, so the added `}` is proven load-bearing rather than merely present. Verified: snippet 51/51 braces; new-project combined net depth 0; the seven other preamble-bearing files with nonzero depth are unchanged from merged next (their own prose, not the preamble, and not in any guard's scan set); all 112 inlined copies and the resolver reference re-synced byte-equal; sync:launcher idempotent on the second run. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3841): backfill changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
36375513b9 |
feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments docs/FEATURES.md was hand-maintained, and every feature PR wrote into two shared mutable cells: the '### N.' heading whose integer was hand-allocated at authoring time, and the hand-maintained table of contents. Concurrent PRs all picked the same next integer, and two PRs adding differently numbered features still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across successive rebases, each collision also costing a full matrix verification run because the sha-keyed pass marker dies with the rebase. Mechanism: one fragment per feature at docs/features/<slug>.md carrying id/title/group (and an optional order) in frontmatter, consolidated by scripts/gen-features.cjs --write|--check into a marker-delimited region of docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings and their order are derived too - a group sorts by its lowest-ordered member - so there is no shared registry to edit either; optional per-group prose lives in docs/features/_groups/<slug>.md. A contributor adds exactly one new file. Wired into regen:derived and lint:generated-sync alongside the eight existing generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting. Migration froze all 168 existing numbers verbatim: identical section set, identical order, identical bodies. Two defects found in the tree are fixed inline rather than carried forward - the '## Related' block had been spliced into the middle of the document, orphaning §142's Reference line, and four inbound anchors were already broken on next (FEATURES.md#runtime-identity in two files, and #143-spec-phase-edge-completeness-probe off by one). Since the repo has no link checker, --check now validates every inbound FEATURES.md#anchor by resolved target, so that class cannot ship silently again; locale FEATURES.md files resolve elsewhere and stay out of scope. Refs #3840 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3840): carry upstream §69 delta into its fragment and harden the generator Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus origin/next. Root cause was a stale base, not extraction loss: those lines landed in |