46e84d5e39a9d4ebe6e8dbbc4cb01ebbf6fb695b
4807 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
46e84d5e39 |
chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat (#2823)
* chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat Phase 2 of epic #2782 under ADR-2782. Delivers D1, D2, D3, D7, D8 and the four Phase-1 vocabulary amendments (A1-A4). - VALID_ROLES gains "reviewer"; the reviewer body is admissible on role:runtime (a host that is also a reviewer keeps one manifest) and on the new role:reviewer (a lane that is not an install target). A reviewer body on role:feature is an error: declaring one is an assertion of lane-ness. - validateReviewerBody + validateLaneProbe + validateLaneInvoke: nine closed enums, a transport discriminator selecting mutually-exclusive invoke sub-shapes, bounded probes (D7), and outputArg required-iff outputChannel is file-arg and forbidden otherwise. - Absent-safe (D4.1): only `undefined` is absent. null/{}/[]/false/0 are malformed assertions and error. 39 of 39 shipped capabilities depend on this. - collectReviewerWarnings: an unknown field inside the body warns, never errors, so a forward-built manifest degrades visibly instead of failing the build. - D8 uniqueness (slug / flags / reviewsSection) lives in validateCrossCapability, so it is enforced at build time over first-party AND at load time over the merged first-party union overlay set, with first-party-wins falling out of the loader's existing ordering rather than a new provenance check. - Config harvest widened past the role==="feature" branch in both the generator and the ownership loop. The often-cited cause of the stranded reviewer config keys -- the runtime body forbidding feature-only fields -- is not the mechanism: `config` is not in FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME. The cause is two harvest sites that never read it. Verified inert: no shipped capability declares config on a non-feature role, and the generated registry is unchanged. Three ADR corrections are folded in (Phase 1 set the precedent of amending in-phase): the misattributed config-stranding cause, D3's inverted profile-membership claim, and the specified capability folder names for lm_studio / llama_cpp, which would have failed the id kebab-case invariant. Closes #2795 * chore(#2795): collapse nine enum checks into one validateEnumField helper Standards-axis review findings, both applied: - Duplicated Code: the enum-membership + enumerate-the-members error shape repeated near-verbatim at nine call sites. Routing them through one helper makes "the error names its valid members" structural rather than a convention repeated nine times, where it would drift. That property is load-bearing until Phase 6 ships the prose reference, because these errors are currently the only documentation of the vocabulary. - Speculative Generality: the isReservedName() pre-check on every enum field was inert. A VALID_* set never contains __proto__/constructor/prototype, so membership alone already rejects them, and "must be one of: ..." is more actionable than "is a reserved name". The literal guards remain where they do real work -- the key-derived write sites in the registry generator and the claim() accumulator. The reserved-name test now asserts all three reserved names are rejected via enum membership, rather than one name via a branch that no longer exists. * fix(#2795): align lane slug grammar with Phase 1 and wire the load-time diagnostic channel Spec-axis review findings, both applied. (1) The slug grammar had diverged from Phase 1's core descriptor. Phase 1 exports LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/ (leading digit permitted); the manifest validator required a leading LETTER. A slug the core descriptor accepts -- a model-named lane such as 4o-mini -- would have been rejected by the manifest validator, which is exactly the translation layer ADR-2782 exists to delete. It was inert only because all eleven shipped slugs begin with a letter, so nothing else would have caught it until a third party shipped such a lane. The grammar cannot be reduced to one definition: Phase 1's module compiles to gitignored build output, and capability-validator.cjs is a committed plain .cjs that must load on a fresh worktree before build:lib has ever run. That makes this the repo's DEFECT.GENERATIVE-FIX class, so the duplication now carries a parity assertion -- laneSlugGrammarMatchesPhase1Descriptor -- which compares both the source grammar and the accept/reject verdict for a shared input set, and fails if the two ever drift again. (2) collectReviewerWarnings had exactly one caller: the build-time generator, which only ever sees first-party in-repo manifests. The real third-party overlay loader never called it and ValidatorModule did not declare it, so ADR-2782 D4.3 -- an unknown field inside a reviewer body is ignored WITH A WARNING -- surfaced nowhere at runtime, which is precisely the case D4.3 exists for. loadRegistry now collects those diagnostics on the accept path, behind a typeof-guard (an older built validator without the function still loads) and a try/catch (ADR-1244 D2's never-crash contract outranks a diagnostic). They land in a NEW OverlayMeta.diagnostics field rather than OverlayMeta.warnings, because warnings records capabilities that were SKIPPED and a consumer treating every entry as inactive would mislabel a working lane. Covered end-to-end by overlayLaneWithUnknownFieldIsAcceptedAndDiagnosed, which drives a real global-scope overlay through loadRegistry and asserts the lane is accepted, produces no skip warning, and yields a diagnostic naming the field. * fix(#2795): make the reviewer validators honour their documented totality contract Isolated adversarial review finding (MAJOR), reproduced by execution. validateReviewerBody documents itself as "TOTAL: returns an array of error strings for ANY input and never throws", and the overlay loader contracts every validator to RETURN errors -- #1461 OVL-1 records a validator that THREW and would have crashed every consumer of loadRegistry. The contract was false at ten sites: JSON.stringify throws on a BigInt and on a circular structure, and every enum/scalar rejection path interpolated the rejected value into its own rejection message. Reading the value could throw too, before any message was built, via a throwing getter or a Proxy get/ownKeys trap. Not reachable through a capability.json today -- every ingestion path is a plain JSON.parse of file text, which cannot express any of those shapes. Fixed anyway: the contract is stated on an EXPORTED function, and a caller must not have to re-derive today's reachability analysis to know whether it holds. Two layers, because serialization safety alone is insufficient: - describeValue() renders any value without throwing, so messages stay useful (a BigInt now reads "got: 10n" rather than degrading to a generic fallback). - A structural try/catch around validateReviewerBody and collectReviewerWarnings makes the guarantee absolute rather than argued, covering read-time throws that fire before any message exists. The same review found the property test guarding this contract was FALSE CONFIDENCE, which is the more important half. fc.anything() at default constraints emits no BigInt, no circular reference, no getter and no Proxy -- 20,000 sampled draws produced zero of each -- so the test was named for a contract its generator could not reach. Even withBigInt is insufficient under whole-value fuzzing, because the defect needs an exotic value in a specifically NAMED field and random key names never land on one. The property is now field-targeted across all twelve reviewer fields, and a companion test enumerates the shapes fast-check cannot generate at all (BigInt, circular, throwing getter, symbol, function, null-prototype) across scalar positions, array-element positions, and read-time traps. Verified red-before-green: with the fix reverted both property tests fail; with it restored all 119 pass. * chore(#2795): backfill changeset pr number to 2823 |
||
|
|
cfdfdf0b4c |
fix(#2695): deliver the complete four-file codex hook set for every profile (#2822)
* test(#2695): failing-first regression for codex hook worker/registry omission * fix(#2695): deliver the complete four-file codex hook set for every profile * test(#2695): regenerate codex install-tree + acknowledge emitted hook ripple * test(#2695): pin core config.toml/hooks.json wiring + fix agent-extension guard (review) * docs(changeset): backfill #2695 PR number to 2822 * fix(#2695): SessionStart wiring assertion is extension-portable (Windows .cmd shim) |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9624167eec |
fix(#2810): accept the documented effortSurface axis on EoS registry entries (#2813)
* fix(#2810): accept the documented effortSurface axis on EoS registry entries The EoS registry schema required an exact eight-key `interactions.axes` object, while `docs/registries/README.md` and `CONTEXT.md` both documented nine keys including `effortSurface`. An entry that faithfully mirrored its upstream descriptor's `effortSurface` key was rejected outright. `effortSurface` reached the runtime-descriptor vocabulary through ADR-1239 amendment #2481 (`HOST_INTEGRATION_AXES`), but the registry's hand-maintained copy of that vocabulary never picked it up. The runtime-descriptor surface is guarded by tests/host-integration-validator-parity.test.cjs; the registry copy had no equivalent guard, which is what let the two drift. Accept `effortSurface` as an OPTIONAL ninth axis validated against the canonical ['argv','none'] rather than a required one: registry entries mirror their upstream registry/eos-entry.json byte-for-byte, so requiring it would retroactively invalidate every entry published before the amendment. Adds tests/registry-axes-parity.test.cjs, which asserts that every key shared between the registry vocabulary and HOST_INTEGRATION_AXES has an identical enum array, plus limit-1/limit/limit+1 boundary coverage on the axes key set. Closes #2810 * test(#2810): fail when a canonical axis is added but never mirrored The enum-equality assertion compares only keys the registry and HOST_INTEGRATION_AXES already share, so it is blind to the exact drift that produced #2810: a new canonical axis appears and the registry copy is never told. Verified by simulation — mutating an enum is caught, adding a new canonical key is not. Assert instead that every HOST_INTEGRATION_AXES key is either modeled by the registry or named in an explicit NOT_MODELLED allowlist (subagentToolkit and isolation, both dispatch sub-fields the registry collapses into its free-form dispatch summary). Adding a canonical axis now fails until someone decides which bucket it belongs in. The allowlist is itself guarded against going stale. Refs #2810 * fix(#2810): harden the axis value lookup with the CodeQL barrier pattern Both orthogonal reviews flagged the same line: `AXES[key] !== undefined` is not an own-property test, and the bracket reads are shaped like a prototype-pollution sink even though the unknown-key gate above provably makes them unreachable. Switch the presence test to `Object.hasOwn` and add the repo's inline literal guards (`capability-state.cts:146-155`, "Prototype-pollution guard (inline literal, CodeQL barrier)"), which CodeQL can follow where it cannot follow the `.includes()` filter that actually does the work. Behavior is unchanged — re-verified all five axes key-count shapes plus a genuine own `__proto__` property built through JSON.parse (the shape a third-party registry PR would submit): it is rejected as an unknown key and Object.prototype is untouched. Refs #2810 * chore(#2810): backfill changeset PR number |
||
|
|
46bae2f9ff |
fix(#2624): write the .gsd-source marker before staging reads it (#2811)
* test(#2624): failing-first regression for stale .gsd-source marker read-before-write
Adds failing-first regression proving a Claude-global upgrade currently lets
staging read a stale prior-version .gsd-source marker before install() rewrites
it. Pre-seeds a stale marker pointing at a still-existing fake source, spies on
findInstallSourceRoot to capture which source staging resolves, and asserts
staging never resolves the stale path. Also covers fresh-install and ghost-marker
negative space.
* fix(#2624): write the .gsd-source marker before staging reads it
A Claude-global upgrade silently installed skill content from the PREVIOUS
version. findInstallSourceRoot prefers <configDir>/.gsd-source, and install()
used to rewrite that marker AFTER staging had already read it — so on an
upgrade the marker still pointed at the prior version's source (an npx cache
dir that still exists on disk) and every converted skill was generated from
the OLD commands/gsd, with generateManifest then recording the stale content's
hash as correct.
Extract the marker write into _writeGsdSourceMarker(runtime, targetDir, src,
isGlobal) and call it BEFORE the staging pass (before the _isSkillsRuntime
branch), preserving the original sourceMarkerFile && isGlobal guard, the
half-published-package existsSync guard, and the non-fatal write-failure warn.
This closes the read-before-write hole for every findInstallSourceRoot
consumer (skills, commands, /gsd-surface, capability-state) in one move.
Long-standing (marker write added in
|
||
|
|
7f13ee5373 |
enhance(#2793): ADR-2782 — reviewer lane becomes a declared capability surface (#2809)
* docs(#2793): add ADR-2782 — reviewer lane capability surface Design lock for epic #2782. Declares a reviewer lane as capability data rather than a core patch across three unrelated surfaces. Key decisions: - D2 transport discriminator (spawn | openai-http) — a survey of all twelve lanes found three that are HTTP endpoints with no binary, which invalidated the single-invoke-shape draft. - D4 the reviewer body is optional and absent-safe at every layer. - D5 a fourth executable-surface disclosure class covering the lane binary or host AND its egress payload classes. - D6 handler is a closed first-party enum, upholding ADR-1016; the consequence — third-party lanes are data-only — is stated plainly. - D7 probe kinds wider than existence, and every probe bounded. Amends ADR-857, ADR-894, ADR-1016, ADR-1244. Also records the D7/D8-extended-by-ADR-1244 marker on ADR-857 that ADR-1244 D8 promised but never added. Closes #2793 * docs(#2793): address orthogonal review findings on ADR-2782 Two blockers from the isolated adversarial pass: - D5 disclosed the spawn binary but not its args, reopening the #1459 bug class already fixed for MCP servers (binary python3 + args -c <program>). args are now disclosed and signature-bound. - hostConfigKey resolves from .planning/config.json, which is mutable after consent with no integrity check, so a lane consented against localhost could be silently redirected to a remote host by an ordinary PR. The resolved host is now consent-bound and re-verified on the invocation path; a mismatch blocks the lane. Majors and spec gaps: - D4 gains an explicit-selection carve-out. Absent-safe governs discovery, never a lane the user named; the current selector records that as info, which Phase 1 now corrects. - D4 gains a warning delivery channel. - D6 enumerates the handler closed-enum members; a closed enum whose membership is left to the implementing phase is not closed. - D6 records aider and plandex as concrete lanes the vocabulary cannot express, rather than claiming sufficiency it did not verify. - D2 gains evidenceClass, requiresBinaries, promptBudgetKey for per-lane divergence that was only prose, and motivates the one-member outputChannel enum. - reviewer.requires renamed requiresBinaries — it collided with the envelope requires (capability deps) at a different nesting depth. - Antigravity two-level timeout: Context cited it then dropped it; now explicitly delegated to the handler. - D9 gains a per-key ownership table, including three keys that stay central because they are policy across lanes, not lane properties. - Phase table maps every decision D1-D9 to a delivering phase; D6 handler modules and D5 invocation-time re-verification were previously unclaimed. - American English per house style. |
||
|
|
84bfef0818 |
docs(#2533): refresh gsd-cursor EoS metadata (#2792)
Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> |
||
|
|
1e3c995e6f |
fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it Every input to `diffEmitted` is base-relative -- `baseline` vs `current`, `changedPaths` from `git diff base...HEAD` -- except the ack set, which was read absolutely, from the working tree only. A differential machine consulting a non-differential input. So `staleAcks` asks exactly one question, "did a delta consume you?", and that cannot distinguish an ack that never explained anything (an authoring mistake) from one whose ripple is now absorbed into the base (the ack's SUCCESS condition). After merge an ack is in the second state but reports as the first. The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request events only, so a push to `next` falls through to origin/next -- the very commit under test. Both sides build identical content, no deltas remain, and every live ack is reported stale. PR #2768 acked a deliberate 40866 -> 42020 byte growth, was green on its own lane, and reddened `next` the moment it merged. It also reds every PR branching off the poisoned base, and since publish-emitted-baseline is gated on the test job, it blocked baseline publication too. Give the ack the base side it was missing. `diffEmitted` now takes `baseAck` -- the same document at the base ref, via `readAckFileAtRef`. An entry already present there is SPENT: it may no longer consume a delta and is never reported stale, only surfaced as `spentAcks` for tidying. An entry new or reworded in this diff stays live, and if nothing consumes it that genuinely fails, with blame on the author who just wrote it. This closes a hazard the IMPLEMENTATION named but could not prevent -- a leftover ack silently pre-clearing the next ripple on its path. (ADR-2719 §3 asserted only that TOUCHING the file is the alarm; its residual-risk list never covered pre-clearing, and §3 now carries an amendment.) Verified against the two-PR laundering sequence -- land an innocuous ack, then change the artifact -- which passed silently before and now fails on both the hash pass and the size ratchet. Three things the design has to get right, each of which was wrong first: - A read failure on the base document THROWS; only absence-at-the-ref returns null. Returning null on error LOOKS armed (every entry stays live) but a live entry's defining power is that it CONSUMES a delta, so null is armed on the staleness axis and DISARMED on consumption -- silently the whole pre-#2789 gate. `git show` cannot tell absence from fault, so absence is established with `ls-tree`. - Re-arming a spent ack costs actual PROSE. Internal whitespace and the zero-width family collapse, and `runtime` is not compared: a doubled space, an invisible character, or a decorative field would otherwise re-arm an ack whose justification still describes the previous ripple, showing a reviewer nothing. - `baseAck` is REQUIRED once an ack declares entries -- omission is an error, not a silent "inherit nothing" -- so a dropped argument fails loudly instead of quietly restoring this bug with the suite green. Because a corrupt document ON THE BASE is expensive (the loud base-side failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs blocks one from landing. It is standalone rather than importing parseAck -- scripts/ ships in the npm package and tests/ does not -- so a parity test runs both surfaces over one corpus and fails on divergence; it caught one immediately, a `null` document, now classed as policy rather than schema. Deadlock is separately foreclosed: a tree carrying no ack never reads the base, so the PR that DELETES a corrupt file still lands. `readAckFileAtRef` takes an injected git runner so all four branches are tested deterministically; it never executes in the remote runner, where the real-tree test skips for want of a base ref. It also refuses an option-shaped ref, since execFileSync's array form stops shell metacharacters but not git's own option parsing. Rejected: skipping the differential when base == HEAD. It treats the symptom, costs real coverage on the push-to-next lane, and does nothing about the downstream PRs the same flaw was reddening. Deletes the now-spent tests/emitted-drift-ack.json, and updates the CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a LIVE entry is, and a spent one is inert. Closes #2789 * chore(#2789): backfill changeset PR number |
||
|
|
d626dbc6e3 |
fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans (#2802)
* test(#1883): failing-first regression for findContextMdIn / listMilestoneArchiveDirs swallowing EACCES
Adds failing-first regression tests proving an unreadable dir is currently
swallowed as empty/null instead of surfacing the permission error. Covers
EACCES, EIO, the unchanged ENOENT empty path, the array fast-path, and both
CONTEXT.md forms. listMilestoneArchiveDirs is exercised in-process via a new
_listMilestoneArchiveDirs test seam (the validate command runs in a subprocess,
so an fs monkeypatch in the test process cannot reach it).
* fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans
findContextMdIn (src/planning-workspace.cts) and listMilestoneArchiveDirs
(src/verify.cts) catch-alled every readdirSync error into the empty marker
(null / []), conflating a genuine ENOENT ('nothing there') with an EACCES/EIO
failure ('can't read this'). An unreadable phase dir was silently reported as
'no CONTEXT.md' (discuss/plan gates wrongly skipped context) and an unreadable
milestones/ dir as 'no archives' (active-milestone resolution / archived-phase
filtering misbehaved).
Narrow each catch to ENOENT only — keep the long-standing null/[] contract for
genuine absence (Hyrum: empty path unchanged) and re-throw every other error so
it propagates to the caller's existing try/catch. All six findContextMdIn
callers either pass a pre-read string[] (no readdir) or sit inside a try block
that already handles readdir failures; the two listMilestoneArchiveDirs callers
live in the validate command path where errors reach the command error handler.
Exposes a _listMilestoneArchiveDirs test seam so the permission-error path can
be unit-tested in-process (the validate command runs in a subprocess, so an fs
monkeypatch in the test process cannot reach the private helper).
* fix(tests): delete stale emitted-drift ack for gsd-phase-researcher.md
Pre-existing base-branch defect, not part of #1883: commit
|
||
|
|
5296ff152d |
docs(#2533): list gsd-cursor EoS integration (#2581)
* docs: add gsd-cursor to the EoS registry Appends the gsd-cursor EoS entry to docs/registries/eos.json and regenerates docs/registries/eos-registry.md. Discussion: open-gsd/gsd-core#2578. * docs(#2533): regenerate eos-registry.md (id-sorted) + add changeset Addresses review on #2581: the generator sorts entries by id, so gsd-cursor (< gsd-omp) must appear first in both the table and the detail sections. Also adds the required .changeset Added fragment (mirrors precedent #2448). * docs(#2533): regenerate eos-registry.md via gen-registry.cjs (markdown escaping) The hand-written markdown left ( ) and _ unescaped; scripts/gen-registry.cjs escapes them (\(...\), model\_profile\_overrides), so gen-registry.cjs --check (part of lint:ci / the lint-tests job) failed. This is the generator's exact output; node scripts/gen-registry.cjs --check now passes. * docs(#2533): regenerate eos-registry.md for corrected compat floor Regenerated via scripts/gen-registry.cjs after the enginesGsd fix. node scripts/gen-registry.cjs --check passes (exit 0). * docs(#2533): correct gsd-cursor enginesGsd floor to >=1.8.0 (real release) The >=1.39.0 floor referenced a legacy-lineage version number that no current @opengsd/gsd-core release has ever reached (latest is 1.8.0). Per trek-e's review, corrected to >=1.8.0 — the current release, which ships the model_profile_overrides.<runtime> config key gsd-cursor uses. * docs(#2533): regenerate eos-registry.md (enginesGsd >=1.0.0, install v1.0.1) * docs(#2533): mirror gsd-cursor enginesGsd >=1.0.0 + repoint install to v1.0.1 Per review: enginesGsd must mirror what the package declares. Corrected engines.gsd upstream in gsd-cursor to >=1.0.0 and cut v1.0.1; this mirrors that range and pins install at the new tag. |
||
|
|
1b41083220 |
enhance(#2151): probe interactive-control for loading and error states (#2575)
* feat(#2151): probe interactive-control for loading + error states The ui-consideration-probe taxonomy mapped interactive-control to only one consideration (long-text), so a control-only UI surface (e.g. a theme toggle) lifted no loading or error consideration — the verifier never asked what a control shows while its action is in flight or when it fails, and a spec omitting those states could PASS. Add 'interactive-control' to the loading and error entries' elements in UI_TAXONOMY so control-only surfaces are probed for in-flight and failure states. empty is deliberately excluded (a control is not data-bearing; empty would be Goodhart noise). No new categories, no cue-map change, no probe-core change — a widening within the closed shape-rooted 8 (ADR-550), an independently-versionable predicate- generator adapter (ADR-857). Regenerated the reference-doc coverage table and the 19 golden-install- parity fixtures (reference-doc hash). Regression test added first (RED: interactive-control yielded only long-text; GREEN: now error+loading+long-text). Closes #2151 * feat(#2151): add changeset for interactive-control loading/error coverage --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a8b40fa53f |
fix(#2547): fail closed on crashing and path-shadowing Kimi payloads (#2595)
* fix(#2547): fail closed on a malformed Kimi edit list in normalizeKimiPayload
`normalizeKimiPayload` rebuilt old_string/new_string with
`String(e.old ?? '')`. `??` guards the value, not the dereference, so a
nullish entry in a Kimi `edit` list threw a TypeError at the top of the
handler, before any tool dispatch. Each guard's outer
`catch { process.exit(0) }` swallowed that crash and emitted the same exit
code as "nothing to report" — turning a should-BLOCK call into a silent
allow.
Two hard blocks were bypassable:
* gsd-worktree-path-guard's cross-git-root write block (#260) — a
StrReplaceFile write whose path resolves to a different git root is
correctly blocked with a well-formed edit list, and silently allowed
with `edit: [null]`.
* gsd-workflow-guard's force-add block on agent-* branches — a Shell
payload carrying a spurious `edit: [null]` field walks past it. The
Bash path never reads `edit`; the field only has to be present to
trigger the crash.
Fixed with `e?.old` / `e?.new`, landed identically in all five copies so
tests/kimi-guard-normalization-parity.test.cjs's byte-identity assertion
still holds.
The crash boundary is nullish specifically, not "non-object": `('x').old`
and `(7).old` are legal reads yielding undefined, so string/number entries
never threw. Both are kept as controls proving the fix did not change
their behaviour.
Regression coverage is folded into the owning suites per CONTRIBUTING.md
(no new bug-* files). Negative-controlled: the nullish cases exit 0
against pre-fix guards and exit 2 after, with positive controls (the
equivalent well-formed payload blocks) and negative controls (in-worktree
writes and benign commands still pass) alongside.
Refs #2547
* test(#2547): exercise the production Kimi payload shape in read-guard tests
The `#2304: Kimi tool vocabulary engages the read guard` cases send
payloads with no `session_id`, and runHook injects none. A live Kimi turn
always carries one — kimi-cli's hooks/events.py `_base()` sets it
unconditionally, and soul/kimisoul.py calls `set_session_id()` at the top
of every turn before tool dispatch, so the ContextVar's `default=""` never
reaches a tool call.
gsd-read-guard treats any non-empty `data.session_id` as "Claude Code
already enforces read-before-edit, skip" (#2520). So the advisory those
tests assert fires only for a shape production never sends: the tests were
green, and the guard was dormant on Kimi. A sibling #2520 case in the same
file asserts the skip when `session_id` IS present — both passed, and the
production shape hits the skip.
Two changes, test-validity only:
* Retitle the #2304 block to say what it proves — the tool VOCABULARY is
normalized through to the Write/Edit branch — with a comment warning
not to read it as production evidence.
* Add a #2547 block asserting behaviour against the production shape
(session_id populated), including a case that pins the delta directly:
the same payload fires without session_id and is silent with it.
The #2547 block characterizes a known gap; it does not endorse it.
Redesigning how the guard discriminates runtimes is explicitly out of
scope for #2547. If a later change makes the advisory fire on Kimi these
tests are supposed to fail — update them then rather than dropping the
coverage.
Refs #2547
* docs(#2547): scope the Kimi guard-engagement claim to what Kimi enforces
#2518 engaged the guards' Kimi matchers and the release notes describe the
result as "All seven guard hooks now engage on Kimi", singling out the
prompt-injection read scanner as "the security-relevant guard" taken "from
silently dormant to engaged". That is not achievable for the scanner at
the emit layer.
gsd-read-injection-scanner.js is a PostToolUse hook, and kimi-cli's
dispatch never inspects PostToolUse hook results: src/kimi_cli/soul/
toolset.py awaits PreToolUse and honours `result.action == "block"`, but
fires PostToolUse via asyncio.create_task() and returns the ToolResult
without awaiting it — the done_callback only retrieves the task's own
exception. So no output shape the scanner emits can block or flag a Kimi
tool call, and `security.injection_blocking` cannot take effect there.
Reshaping the scanner's output would not change this; the enforcement gap
is in kimi-cli's PostToolUse handling, which is out of scope here.
This corrects the claim rather than the code — there is no gsd-core emit
fix that would make it true:
* .changeset/2304-kimi-guard-tool-name.md — the fragment is unreleased,
so it would otherwise ship this as a CHANGELOG security claim.
Headline narrowed to "normalize Kimi's payload shape" and a scope
paragraph added naming what actually blocks on Kimi (the two
PreToolUse blocks) versus what cannot.
* docs/migration/kimi-to-kimi-code.md — the scanner was listed under
"Every GSD `PreToolUse` guard"; it is PostToolUse. Corrected, and the
"What about the dormant guards?" section now splits enforceable from
not-enforceable instead of saying Phase 0 "fixed all seven".
* hooks/gsd-read-injection-scanner.js — the same scope note in the
file's own Kimi rationale comment, where the next contributor to touch
the normalization will actually read it. Comment only; the shared
normalization block is untouched and byte-identity still holds.
Refs #2547
* chore(#2547): regenerate golden install-parity fixtures for the guard fix
The golden install-parity fixtures record a content hash per installed
file, so changing the five guard hooks changes their hashes across every
runtime's fixture. Regenerated with the full sweep (build, gen:golden,
size:baseline) rather than a single generator — running gen:golden alone
leaves tests/workflow-size-baseline.json stale and loses CI jobs to a
regeneration that looked complete.
The size baselines came out unchanged (no workflow or agent bodies
touched) and the hash delta is confined to exactly the five guards:
gsd-prompt-guard, gsd-read-guard, gsd-read-injection-scanner,
gsd-workflow-guard, gsd-worktree-path-guard.
Refs #2547
* fix(#2547): guard the String() coercion in normalizeKimiPayload too
Found by adversarial review of the first commit, then reproduced against
pristine next: `e?.old` closes the nullish dereference but leaves a second
route to the same crash-to-allow.
`{"toString": null}` is valid JSON, and coercing it throws
`TypeError: Cannot convert object to primitive value` — so an edit entry
that IS a well-formed object still crashes normalization, still lands in
the outer `catch { process.exit(0) }`, and still downgrades a should-BLOCK
call to a silent allow. Confirmed on both hard blocks:
{"tool_name":"Shell","tool_input":{
"command":"git add -f secret.env",
"edit":[{"old":{"toString":null},"new":"x"}]}} -> exit 0 (was)
{"tool_name":"StrReplaceFile","tool_input":{
"path":"<main-repo>/src/index.ts",
"edit":[{"old":{"toString":null},"new":"x"}]}} -> exit 0 (was)
Both exit 2 now.
The coercion is wrapped rather than replaced with a `typeof === 'string'`
test on purpose. Degrading only the non-coercible entry keeps
stringification identical for every value that CAN coerce — numbers,
arrays, plain objects — which matters because gsd-prompt-guard scans
new_string for injection patterns, and a `typeof` test would silently stop
scanning content that reaches that scan today (e.g. `new: ["ignore all
previous instructions"]` currently stringifies and is scanned). Verified:
zero behaviour change across string, number, bool, null, array-of-strings,
nested array, plain object and `__proto__`-keyed input; only the throwing
case changes, from crash to ''.
Regression cases are negative-controlled against the previous commit: the
four new coercion-trap tests fail with only the `e?.old` fix in place and
pass with this one.
Refs #2547
* chore(#2547): cover the String() coercion vector in the changeset
The release note described only the nullish-dereference route. Both routes
reach the same fail-open, so both belong in the changelog entry, along with
why the coercion is wrapped rather than type-tested.
Refs #2547
* chore(#2547): point the changeset fragment at the real PR number
The fragment has to exist before `gh pr create` runs, so it carried the
issue number as a placeholder. Corrected to 2595 now that the PR is open.
Refs #2547
* fix(#2547): make Kimi's `path` authoritative over a model-supplied `file_path`
normalizeKimiPayload copied Kimi's `path` into `file_path` only when
`file_path === undefined`, so any `file_path` the model chose to include won
outright. Every guard reads `file_path`; kimi-cli executes on `path`. The guard
therefore inspected one file while the write landed on another.
This bypass needs no crash. A payload pairing a cross-root `path` with a
spurious `file_path: ""` left gsd-worktree-path-guard reading an empty string
and exiting 0, while the identical write without the extra key blocked — the
same cross-root write the #260 block exists to catch. The shadowing also
preserved a non-string `file_path` (`[]`), which threw inside that guard's
path.isAbsolute() and reached its outer `catch { process.exit(0) }`: the same
crash-to-allow the rest of #2547 closes, reached through the guard's own read
rather than through normalization.
Reachability is not speculative. kimi-cli's soul/toolset.py json-parses the
model's raw tool arguments and passes that dict verbatim as tool_input to
PreToolUse, performing typed validation only later inside tool.call() — after
the hook has already decided. So the model controls extra keys in tool_input at
the moment the guard runs. kimi-cli's file tools carry no `file_path` field at
all (src/kimi_cli/tools/file/write.py, replace.py), so a `file_path` in a Kimi
payload is always model-supplied.
`path` now wins outright. Overwriting can only ever narrow what a guard inspects
to the path that will actually be written, so it cannot under-block.
Normalization returns early for non-Kimi tool names, so the native Claude Code
contract (file_path governs) is untouched.
Landed identically across all five inlined copies; the byte-identity assertion
in tests/kimi-guard-normalization-parity.test.cjs enforces that.
* test(#2547): cover the file_path-shadowing bypass in the #260 guard suite
Four cases, each exiting 0 (bypass) against the pre-fix guards: a spurious
empty-string file_path, an in-worktree decoy file_path, and non-string
file_path values (array and object) that additionally crashed
path.isAbsolute() into the outer catch.
Two controls that are not bypass cases and matter as much:
- an in-worktree write carrying a cross-root DECOY file_path must still exit
0. Pre-fix this blocked, because the decoy won; the guard now follows the
path kimi-cli executes on in both directions, so the fix narrows what is
inspected without over-blocking.
- a native Claude Edit (no `path` field) must still block on file_path alone.
normalizeKimiPayload returns early for non-Kimi tool names, and this pins
that the non-Kimi contract did not move. It passes both pre- and post-fix
by design.
Negative-controlled: run against the pre-fix hooks, the four bypass cases and
the decoy control fail, and the native-Claude control passes.
* test(#2547): back the totality claim with property tests over fc.anything()
This PR claims the fix "makes normalization total over the inputs JSON can
express" — a for-all guarantee — while the tests backing it are example-based,
each shape added reactively after a crash was found by hand (the String()
coercion trap was itself found by adversarial review after the first commit
shipped). Example-based tests cannot substantiate a for-all claim; they record
the counterexamples someone happened to think of.
Four properties over fc.anything(), which is exactly the JSON-expressible
domain the claim names:
(a) totality over any tool_input
(b) totality over any edit list — the crash surface both #2547 fixes targeted
(c) `path` always wins over any model-supplied `file_path` (the review blocker
invariant: a guard reading file_path can never be aimed at a file other
than the one kimi-cli writes)
(d) a non-Kimi tool_name passes through untouched — the native Claude contract
normalizeKimiPayload is inlined per hook with no runtime binding, so there is
nothing to require. The block is extracted from hook source and evaluated via
the SAME extraction contract kimi-guard-normalization-parity.test.cjs uses, so
a source edit that breaks one breaks both instead of silently testing a stale
block. An extraction floor test fails loudly if the extraction yields a no-op.
Non-vacuous, and checked rather than assumed: against pristine pre-#2547 `next`,
(a), (b) and (c) all FAIL and (d) passes. (a) needed the fix that makes it
meaningful — a bare fc.anything() for tool_input passed even against the live
defect, because arbitrary generation essentially never invents the `edit` key
the crash lives behind, so the generator is biased onto the keys normalization
actually reads and unioned back with unbiased input.
* chore(#2547): cover the shadowing vector in the changeset and regen goldens
Golden install-parity churn is hash-only, on exactly the five hook files this
round changed. gsd-phase-boundary.sh is deliberately unchanged.
* test(#2547): make the property test able to kill the coercion mutant
Review Major 1: the generative test added to stop the NEXT counterexample
could not kill the one it was written for. Reproduced the reviewer's matrix
independently — against the shipped generator, a mutant reverting `editText`
to the unguarded `String(v ?? '')` passed all four properties.
Cause, confirmed by measurement: the edit-array ENTRIES were bare
`fc.anything()`, which essentially never invents an `old`/`new` key, so
`e?.old` was always undefined and `String(undefined ?? '')` never coerced
anything. That is the same vacuity the file's own comment describes one level
up, reproduced one level down.
The review's prescribed fix — bias the entry onto `{old, new}` — is necessary
but NOT sufficient, and this is the part worth recording: measured over 20,000
draws, bare `fc.anything()` yields a non-coercible value 3 times (0.015%). At
`numRuns: 200` an `old` key holding a hostile value essentially never
co-occurs, and the mutant survives the entry bias too. Both levels need
biasing — the entry onto the keys normalization reads, and the VALUE onto the
shape that actually throws.
`{"toString": <non-function>}` is that shape and stays inside the
"JSON-expressible" domain the claim names (JSON.parse produces it verbatim);
`fc.anything({withNullPrototype: true})` would also kill the mutant but widens
the domain past what the PR asserts, so it is not used.
Verified: M1 now dies at every seed tried (1/7/42/99/4242/31337, failing
within 3-31 cases) while HEAD stays green at all of them.
Also closes three coverage gaps the review listed as nits — properties (e)
totality over any JSON value as the WHOLE payload, (f) the tool_output →
tool_response mapping (including that an existing tool_response is not
clobbered), and (g) an empty edit list reconstructing nothing.
Property (e) required a one-line fix in the normalizer itself: `JSON.parse
('null')` is null, and null/primitive payloads threw on the `data.tool_name`
read — falsifying the "total over the inputs JSON can express" claim. Harmless
in practice (the throw landed in the same fail-open catch as the exit 0 it now
takes deliberately), but the claim should be true as stated. Landed
byte-identically across all five copies.
Refs #2547
* fix(#2547): make old_string/new_string authoritative, not merely defaulted
Review Major 2: the shadowing class the prior round's BLOCKER closed for
`file_path` survived one field over. `old_string`/`new_string` were still
reconstructed only `if (input.<field> === undefined)`, so a model-supplied
value won.
The argument for making `path` authoritative applies verbatim here.
kimi-cli's StrReplaceFile schema is `path` + `edit` only
(src/kimi_cli/tools/file/replace.py @ 4a550ef) and carries no
`old_string`/`new_string` at all, so either key appearing in a Kimi payload is
always model-supplied — exactly like `file_path`.
Verified end-to-end against the reviewer's payload: a cross-root write
carrying `new_string: ""` alongside an injected `edit[].new` left
gsd-prompt-guard reading '' and returning at its `if (!content)` guard, so the
injection advisory never fired and the reconstructed content was never
scanned. `new_string: null` behaved identically. Negative-controlled: both
produce empty output against pre-fix source and fire the advisory after.
Chose unconditional reconstruction over the offered `typeof` alternative
deliberately. A type test closes `""`/`null` but leaves the interesting case
open — a benign NON-EMPTY decoy (`new_string: "chore: tidy"`) shadows just as
effectively and passes any type test. The new suite includes that case
specifically; it is what discriminates between the two candidate fixes.
Also pins the kimi-cli SHA in the authoritative-path comment, as requested —
it cited file names with no version while the issue pins 4a550ef.
Landed byte-identically across all five inlined copies; the parity test's
byte-identity assertion holds.
Refs #2547
* fix(#2547): close the non-string file_path crash-to-allow at every read site
Review Major 3: the crash-to-allow was closed only as a side effect of `path`
masking the bad value, while the changeset read as though it were closed
outright. Confirmed both of the review's reachability claims: `[]`/`{}`/`42`
are truthy, survive the `!rawFilePath` early-out, and throw inside
path.isAbsolute() into the outer `catch { process.exit(0) }`; and normalization
returns early for native Claude Code payloads (KIMI_TOOL_NAMES has no 'Edit'
entry), so `{"tool_name":"Edit","tool_input":{"file_path":[]}}` reached it
untouched — this guard's original #260 surface.
Reproduced on a real fixture: string cross-root path exits 2, the identical
payload with `[]` or `{}` exits 0.
Swept the class rather than the instance. Five more untyped read sites across
four other hooks, each one line from a type-strict or method-dependent call.
Census of what each can actually do:
gsd-worktree-path-guard.js:173 BLOCKS -> live bypass (the review's finding)
gsd-prompt-guard.js:128 scanner -> silenced the injection scan, the
same outcome as Major 2 by another
route; verified empirically
gsd-workflow-guard.js:206 advisory only (its exit-2 is the Bash
force-add path, which reads
`command`, not `file_path`)
gsd-read-guard.js:141 advisory only
gsd-read-injection-scanner.js:213 advisory only
gsd-windsurf-pre-write.js:75 ALREADY TYPED — the shape adopted here
All six now read typed. The workflow-guard site keeps its truthiness fallback
(`(typeof x === 'string' && x) || ...`) because a bare type test would let an
empty `file_path` shortcut the `path` fallback.
Also declares one swept hit NOT fixed: `gsd-workflow-guard.js:175` reads
`command` untyped on a genuinely blocking path. Same shape, but not
exploitable — unlike file_path/path there is no second field carrying the
executable value, so a non-string command cannot smuggle a real `git add -f`
past the block. Left alone rather than widen this PR into the Bash path.
The regression gate is a SOURCE-level invariant, not a behavioural one, and
that is deliberate: the fixed read and the crashing read are black-box
identical — both end at exit 0, one via the catch and one via the early-out.
A test asserting exit 0 on a non-string payload passes against the unfixed
code, which is the same false-green the review flagged in the existing
`['non-string file_path (array)', []]` cases. Repeating it one level up would
be no better. tests/kimi-guard-typed-payload-reads.test.cjs fails if any hook
regresses to an untyped read (negative-controlled: it reports all five
pre-fix sites with correct file:line).
The behavioural cases requested — non-string file_path with NO `path` key —
are added to worktree-safety.test.cjs and labelled honestly as documenting the
explicit fail-open rather than detecting a revert.
Also states the relative-path premise (review Minor 5) at the early-out that
depends on it: "always safe" holds only while every runtime reaching there
resolves relative paths against the tool CWD. Claude Code satisfies it by
requiring absolute paths; kimi-cli's resolution behaviour is NOT verified here
and is recorded as an unverified premise rather than an asserted bypass.
Refs #2547
* docs(#2547): correct the changeset's closed-claim and fold the misattributed note
Review Major 3 also flagged the fragment: it said the non-string vector "threw
inside that guard's path.isAbsolute() ... `path` now wins outright", which
reads as closed when it was closed only conditionally. Rewritten to state what
is now true — closed unconditionally at all six read sites — and extended with
the Major 2 finding.
Review Minor 6 (the #2547 scope note living in a `pr: 2518` fragment) turns out
to understate the problem. Rendering the changelog and re-parsing it shows the
note is not merely misattributed — it is DROPPED. serializeChangelog emits each
fragment as a single `- ` bullet, and parseChangelog terminates a bullet at the
first non-continuation line, so everything after a blank line is lost on
re-parse. Audited all 44 fragments: exactly one was lossy —
2304-kimi-guard-tool-name.md, losing 656 of 1730 characters, i.e. precisely
that second paragraph. Folding it into this PR's fragment fixes the
attribution and the silent loss together; all 44 now round-trip losslessly.
That same mechanism is why the remaining nit — reformat this fragment's
~2,000-character paragraph for readability — is NOT applied. A paragraph break
or a bullet list would silently truncate the entry at the first blank line
(verified for both). The single-paragraph form is load-bearing under the
current serializer, not an authoring preference. Worth its own issue; noted in
the PR thread rather than worked around here.
Refs #2547
* chore(#2547): regenerate golden install parity after rebase onto next
Rebased onto `next` @
|
||
|
|
7e8f6a6d7d |
enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (#2685)
* enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (W025) The artifact<->git check has existed since the beginning but was only ever pointed at .planning/research/SUMMARY.md (new-project.md:1145, new-milestone.md:425). Phase summaries -- the ones that actually claim "I created these files" -- were never checked. - extract verifySummaryCore from cmdVerifySummary: same checks, lifted out of the output() wrapper so callers consume {passed, checks, errors} directly instead of shelling out and re-parsing JSON; cmdVerifySummary is now a thin adapter over it - validate.health gains advisory W025 per phase SUMMARY with missing files Advisory only: appends to warnings[], never touches status escalation beyond the channel's own warning semantics, the repair set, or readVerificationStatus. Resolves both open questions from triage: (a) commits_exist is deliberately NOT surfaced -- its hash pattern matches any hex-shaped token in prose, too loose to show a user; (b) a phase carries N per-plan summaries plus a legacy bare SUMMARY.md, so all of them are checked via the repo-wide filter. * chore(#2572): add changeset fragment * feat(#2572): move the SUMMARY artifact check to phase completion Responds to the #2685 review. Three substantive changes. Seam (Blocker 2). The check now runs in cmdPhaseComplete, the seam the issue body cited (src/phase.cts:~1745), not validate.health. That channel does exist: cmdPhaseComplete declares warnings[], populates it from the UAT/VERIFICATION pre-scan, and emits it. The cycle objection raised against the earlier deviation holds for state.cts only -- verify.cts has no transitive import path to phase.cts, so phase.cts -> verify.cjs adds no cycle (verified over every src/*.cts). Moving it also retires the retroactive firing across all historical phases: this fires once, at completion, for the phase being completed. Extraction (Blocker 1). Pattern 2 now excludes [ and ] from its path class. The SUMMARY templates prescribe a YAML flow sequence (key-files.created: [a.ts, b.ts]) and the label matches case-insensitively, so the class previously captured the literal [ and produced a candidate that can never exist on disk -- firing on healthy projects built from the templates GSD itself ships. Stripping frontmatter was the other offered remedy; measured across all three shipped templates it is a no-op on top of the exclusion, so it is not carried. Consequence named in-code: the key-files block still is not read, which needs a real frontmatter parse. Also narrowed to the noise classes confirmed in review -- globs, bare hostnames, and paths resolving outside the project are skipped rather than reported, and the containment guard the old comment claimed now actually exists. Budget (Majors 1 and 3). verifySummaryCore takes a checkCommits option; phase completion passes false, so the discarded git cat-file probes are not spawned at all. It also passes Infinity, so every referenced file is reported instead of the first two -- a summary listing twelve files of which nine are missing now says nine, not zero. The verb keeps its historical 2-file default. Tests (Major 2). The vacuous fixtures are gone with the health block. The replacements use /-bearing paths that genuinely extract, and each fix was mutation-checked: un-anchoring pattern 2, dropping the glob, hostname or containment filter, forcing commit checking on, and re-capping at 2 each fail at least one test. * docs(#2572): describe the phase-completion SUMMARY artifact check The W025 text under /gsd-health is withdrawn with the health seam; the check is documented where it now runs, under `phase complete` in docs/CLI-TOOLS.md. Both the docs and the changeset previously overclaimed: they said a referenced file not on disk is warned about, while at most two candidates per SUMMARY were ever examined. The cap is gone at this seam, so the claim now holds -- and the text states the limits that remain, rather than leaving them to be discovered: the key-files frontmatter block is not read, commit hashes are not resolved, and globs, URLs, bare hostnames and out-of-project paths are skipped rather than reported. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
6932fb16d7 |
enhance(#1699): require read-and-cite provenance for in-repo discrete values (#2768)
* test(#1699): failing-first contract tests for in-repo value provenance
Nine assertions on the deployed gsd-phase-researcher contract: the discrete-value taxonomy, the same-session Read requirement, path-AND-line-range citation, grep-alone exclusion, the verbatim quote and paraphrase ban, the quote-is-the-checkable-artifact guard, [ASSUMED] routing for unquoted skeleton values, a no-regression guard on the pre-existing package name provenance rule, and a single-definition-site guard.
Eight of the nine fail against the unmodified agent on origin/next; the ninth is the no-regression invariant and passes in both states, which is the intended enhancement shape.
Added to tests/research-agent-profiles.test.cjs rather than a new file: that file already carries the allow-test-rule exemption <runtime-contract-is-the-product> research agent .md content is the governed surface, so no new allowlist entry and no change to the per-module test-file count.
* enhance(#1699): require read-and-cite provenance for in-repo discrete values
The claim-provenance system governed external facts (npm registry, official docs, Context7, package-name provenance). For an in-repo discrete value -- an enum, schema or type union, error code, status constant, or filesystem path -- [VERIFIED] could be earned from training memory or a bare codebase grep, which proves a string occurs, not that the definition was read.
A drifted value passes into RESEARCH.md, is lifted by the planner into PLAN.md's <interfaces> context block, and is trusted by the executor, where it fails at parse()/typecheck as a mid-execution deviation -- the most expensive place to discover it.
The rule lands beside its structural sibling, the package name provenance rule, since both say existence is not verification. The verbatim quote is named as the load-bearing artifact: a citation with no quote does not earn the tag, however precise the line range looks. That keeps the rule falsifiable against the file rather than a self-report, which is the Goodhart guard.
Scope note: the quote goes in RESEARCH.md beside the claim, NOT in an <interfaces> block. <interfaces> appears zero times in this agent on next -- it is planner-side, defined at gsd-core/references/planner-interface-context.md:15 as a PLAN.md structure. The issue text and triage both said <interfaces>; instructing the researcher to populate a block it does not emit would be undefined.
Defined once at the definition site; the source-hierarchy recap is deliberately untouched, since restating it is the paraphrase-drift mode META.RULE.brief-no-paraphrase names.
* test(#1699): regenerate agent size baseline and golden parity fixtures
Regenerated via npm run size:baseline and npm run gen:golden, never hand-edited. agent-size-baseline gsd-phase-researcher.md 40866 to 42020 (LARGE tier, cap 49152, 7132 bytes headroom remaining, per ADR-1610's per-file baseline guard). 18 of 19 golden fixtures updated; pi.json is unchanged because the pi runtime ships zero agents.
* chore(#1699): add changeset for the in-repo value citation rule
User-facing behavior change in gsd-phase-researcher, so a Changed fragment is required. pr: 2768.
* test(#1699): acknowledge the gsd-phase-researcher growth for the emitted-drift gate
CI test (ubuntu-latest, 22) failed on emitted-attribution.test.cjs: gsd-phase-researcher.md grew 1154 bytes (40866 -> 42020) without an acknowledgment. The differential emitted-attribution gate landed on next in
|
||
|
|
e276cc7f00 |
enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate check-glossary-refs asserts that every backticked tests/ token in CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719 section 3) is absent on a healthy next BY DESIGN — it appears only inside a PR that needs it, which is what makes touching it the alarm. It passed before only by accident of backtick pairing: CONTEXT.md's RULESET entries are themselves backtick-wrapped and contain backticks, so the token happened to fall outside a code span. Any edit that shifted the parity exposed it. A gate that passes by luck is not passing. The exemption is exact, not a prefix hole: a sibling missing tests/ path still fails, and a test locks that. * feat(#2778): make the size-ratchet failure name its own remedy The growth branch stated a requirement and withheld the means of satisfying it: no ack file named, no schema, no key format, and no do-not-regenerate line — so the likeliest guess was to hunt for a baseline that #2724 deleted. Observed live on #2543. All remediation now comes from one frozen REMEDIATION export whose example document is rendered from ACK_VERSION, so the taught schema cannot drift from the schema parseAck accepts. A round-trip test feeds the printed document back through parseAck. The report is now built as a typed IR (buildReport) that formatReport renders, so tests assert on structure rather than prose, per CONTRIBUTING.md's raw-text-matching rule. Two defects found and fixed inline while building: - diffEmitted's validation early-return omitted newFileCapExceeded while formatReport reads its length, so the branch that reports a failed git diff threw a TypeError instead of naming the problem. - Printing one complete ack document per failing branch made each read as the whole file, so pasting the second over the first silently lost an acknowledgment. One document now covers the whole report. Closes #2778 * chore(#2778): backfill changeset pr number to 2780 |
||
|
|
0997d4f443 |
fix(#2620): inject the reference DispatchLogger on the live dispatch seam when observability is enabled (#2621)
* fix(#2620): inject the reference DispatchLogger on the live dispatch seam when observability is enabled The Command Routing Hub defaulted to createNoOpLogger and no caller ever injected createDefaultLogger, so GSD_AUDIT=1 wrote nothing and failed dispatches emitted no structured JSON to stderr — contradicting ADR-0174 §5/§6, CONTEXT.md's Dispatch Observability Module contract, and docs/CONFIGURATION.md. Inject the reference logger at both live createHub() sites, gated on the existing opt-in signal (newly exported isAuditEnabled). When observability is off no logger is injected, so the Hub keeps its no-op fallback and default output stays byte-for-byte identical. Enabling stderr-on-error unconditionally adds a second line to the --json-errors envelope that callers parse as exactly one JSON line, so that is deferred to its own increment under #2619. * chore(#2620): add changeset for the dispatch logger wiring fix * test(#2620): cover the phase seam and drop try/finally from the adapter test Two review findings from the #2621 round-1 review. The fix wires the logger at BOTH live createHub() seams, but only cjs-command-router-adapter was exercised. Adds a fail-first regression test for src/phase-command-router.cts:258 — verified RED against a tree with that hunk reverted (1 fail, exact assertion) and GREEN with it restored — plus a negative pin that no trace file appears when GSD_AUDIT is unset. The negative case passes pre-fix and is a pin, not fail-first. CONTRIBUTING.md:344 forbids try/finally inside test bodies; the new adapter test used it. Converted to the Pattern-2 t.after() form, switched to the centralized createTempDir helper, and removed the now-unused os require. * chore(#2620): scope the changeset to the activation path that actually ships The fragment claimed config.audit.enabled activates the audit trail. It cannot: both seams call isAuditEnabled() with zero arguments, so the config branch in _isAuditEnabled is unreachable from production, and src/config-schema.cts registers no audit key at all — a user setting it would be silently dropped. That string ships in the user-facing CHANGELOG. Scoped to GSD_AUDIT=1, which is what actually works. The missing schema key stays a disclosed deferred sub-defect on #2620. Also adds the (#2620) issue backlink the other fragments carry. * docs(#2620): correct the fork-leaked issue reference in the wiring comments Four files cited this fix as #26, the issue number from the fork where the change was first written. Upstream #26 is an unrelated closed SDK issue, and next already uses #26 with that meaning in src/validate.cts:17,29,42 and src/config.cts:474, so these references pointed somewhere real and wrong rather than merely dangling. Baked into permanent doc comments, they reach users compiled via the ADR-457 build-at-publish path. The changeset and tests/phase-command-router.test.cjs already cited #2620; this brings the remaining four files into line. Comment-only, no behaviour change. build:lib produces no generated drift. The rename is scoped to these four files so the pre-existing SDK #26 references in validate.cts, config.cts, health-validation.test.cjs and config.test.cjs are deliberately left untouched. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
16e59d0db5 |
fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs
Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @
|
||
|
|
3274db2757 |
fix(#2526): remove gsd-ui-auditor's uncallable Playwright-MCP block (#2594)
* fix(#2526): drop gsd-ui-auditor's uncallable Playwright-MCP block The agent declares `tools: Read, Write, Bash, Grep, Glob, Skill` — no `mcp__*` grant of any kind — while its body presented a `<playwright_mcp_approach>` block as the *preferred* capture path. That branch was unreachable by construction: the availability check had a fixed answer, the three `mcp__playwright__*` calls could never dispatch, and the "when Playwright-MCP is NOT available" fallback was the only branch that ever ran — 39 lines of instruction loaded on every /gsd-ui-review spawn that also invited the model to claim a capture path it could not take. Remove the dead block, leaving the CLI screenshot path as the sole documented approach. Guard the class in tests/mcp-tool-inheritance.test.cjs, which already owns agent MCP-grant parity: the new block generalizes the #1284 researcher check from two agents and one dispatch table to every agents/*.md and its whole body — no agent may document an `mcp__<server>__*` namespace absent from its own `tools:` declaration. Frontmatter is read through the canonical parser (gsd-core/bin/lib/frontmatter.cjs) rather than a hand-rolled scan, so inline CSV, block sequences, flow arrays, quoted scalars and full-line comments are handled by construction; inline comments inside a scalar survive that parser, so they are stripped explicitly. The check is server-level by design, ignores prose metavariables like `mcp__X__*`, matches hyphenated server ids, and carries a discovery guard plus negative controls for every documented boundary so it cannot decay into a vacuous pass. The session-level Playwright-MCP pass in gsd-core/workflows/ui-review.md is deliberately untouched — workflow files carry no fixed allowlist, so their availability check is genuinely runtime-detected and honest. Fixes #2526 * chore(#2526): set changeset fragment pr to 2594 The fragment shipped with the documented `pr: 0` placeholder because the PR number does not exist until the PR is opened, and scripts/changeset/parse.cjs rejects `pr <= 0`. Now that the PR is open, set the real number so changeset-lint passes. * test(#2526): cover the two-char server-id boundary of the metavariable exclusion The length-1 "prose metavariable" exclusion was tested at length=1 and at real ids (>=3 chars), but never at length=2 — the limit+1 boundary where a server id starts being recognized. Review finding on #2594: `mcp__ab__foo` in a body with no grant must flag `['ab']`. * fix(#2526): treat a bare mcp__* grant as covering every server `grantedServers()` stripped `mcp__*` to the empty string and dropped it via `if (server)`, so an allowlist that grants every MCP server read as granting none — and the guard then fired against a body the grant plainly covered. That is the one input shape that inverts the check, turning it against a correct agent rather than merely missing a bad one. A `/^mcp__\*+$/` token now sets a GRANT_ALL sentinel that short-circuits `ungrantedServers()`. The sentinel `*` is outside REFERENCE_RE's character class, so no body reference can collide with it. A bare `mcp__` with no wildcard stays a typo rather than a grant and keeps failing closed. No agent uses the `mcp__*` spelling today, so this was latent rather than live. Two negative controls pin both halves. * fix(#2526): scan the frontmatter description for MCP references too `ungrantedServers()` scanned `stripFrontmatter(content)` only, so an `mcp__foo__bar` reference in the `description` field escaped the check. That field ships with the agent and the dispatcher reads it, which makes a dead reference there exactly as dead as one in the body. Only `description` is added to the scanned surface, never the whole frontmatter: `tools:` is the grant list itself, so scanning it would let every allowlist satisfy itself and turn the guard vacuous. A negative control pins that boundary alongside the new positive case. All 34 per-agent tests still pass with the wider surface, so no live agent verdict changes — this was latent. * test(#2526): give multi-character placeholders a convention the checker knows The metavariable exclusion is `length === 1`, so the natural placeholders `mcp__SRV__*` and `mcp__SERVER__*` were flagged as real references — and the failure message then offered an author two remedies ("grant the namespace or drop the block") that both misread what they wrote. Adopts the angle-bracket half of the suggested fix: `mcp__<SERVER>__*` is the sanctioned multi-character placeholder, exempt by construction because `<` is outside the reference pattern's character class. This pins an existing property rather than adding a special case. Declines the all-caps half. An all-caps exemption would be a false NEGATIVE for any real server spelled in caps, and a guard that misses a dead reference fails in exactly the direction this check exists to prevent. The bare-caps form keeps firing; the message now names the convention as a third remedy. Also corrects "grants neither" in that message, which was wrong for any count other than two. * test(#2526): pin the zero-length server id, completing the boundary triple `mcp____foo` yields `[]`, but for a different reason than the length-1 case: it is unrepresentable by `/mcp__([A-Za-z0-9_-]+?)__/g` since `+?` requires at least one character, so the pattern skips it before the metavariable exclusion is ever consulted. Pinning limit-1 completes the 0/1/2 boundary rule on its own terms and records which mechanism owns the case. * docs(#2526): correct every drifted AGENTS.md Tools row, not just the one The review asked for the one-line `gsd-ui-auditor` correction (missing `Skill`). Sweeping the defect class first — every `**Tools**` row in docs/AGENTS.md against its agent's `tools:` frontmatter — found it was 26 of 34 rows, so the one-line framing was the reviewer's premise rather than the population. Breakdown of the 26: * 21 omitted `Skill`, 6 omitted `Edit` (overlapping) — under-promises, the same drift class as #2526 but in the harmless direction. * 8 wrote `mcp (context7)` as shorthand while frontmatter granted up to 8 servers (firecrawl, exa, tavily, ref, jina, perplexity, both context7s). * 1 was actively wrong: gsd-debug-session-manager documented `Task`, a tool name that no longer exists — the #2526 shape at the doc layer, naming a capability that cannot dispatch. Every row is now the frontmatter `tools:` value verbatim, which is also what makes the parity guard in the following commit non-brittle. The diff is 26 insertions / 26 deletions, all Tools rows. * test(#2526): guard AGENTS.md Tools rows against agent frontmatter The 26 corrected rows in the previous commit were free to drift because nothing asserted the role card and the frontmatter agreed — the same reason the #2526 block itself survived. Correcting them without an invariant just resets the clock. Lands in agent-classification-parity.test.cjs rather than a new file: that suite already owns docs/AGENTS.md as a contract surface, already carries the `allow-test-rule` exemption for treating the doc as the product, and file count is the unit of CI overhead (docs/TESTING-SUITES.md). Compares the row to the frontmatter value VERBATIM, not as a set — a set comparison would keep accepting the "mcp (context7)" shorthand that hid eight grants behind one, which is the under-documentation half of the drift. Carries the same discovery guard #2526's own check uses: a section with a granted `tools:` but no **Tools** row fails loudly, so deleting a row cannot silently retire its assertion. Both halves are negative-controlled — against the pre-fix doc it fails naming 26 rows (gsd-ui-auditor:339 among them), and with a row deleted it fails on the missing-row assertion. * chore(#2526): note the AGENTS.md drift correction in the changeset The role cards are user-visible, and 26 of them documented a tool set the agent did not have. Type, `pr: 2594`, and the trailing `(#2526)` are unchanged. * fix(#2526): use a CRLF-safe split in the AGENTS.md Tools-row guard `lint-tests` (npm run lint:ci) rejected `rawAgentsMd.split('\n')` under the repo's local/no-crlf-fragile-split rule: Windows autocrlf yields CRLF, so a trailing \r rides into the parsed line. Switched to `/\r?\n/`. Caught by CI on the round-3 push before the response comment went out. * docs(#2526): correct the drift tallies stated in c5a9607e Re-derived the census programmatically from the pre-fix doc instead of by eye. The 26-of-34 headline was right; the breakdown was not. Skill omitted 21 -> 22 Edit omitted 6 -> 7 "mcp (context7)" shorthand 8 rows -> 7 rows The 8 was conflating two things: 8 rows omitted MCP grants entirely, but only 7 of them used the "mcp (context7)" shorthand — gsd-executor listed no MCP at all. Also names the one `Agent` omission (gsd-debug-session-manager, the row that still read `Task`). c5a9607e's message keeps the wrong numbers rather than rewriting a pushed branch mid-review; the test comment and changeset are the durable statements and both are corrected here. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1c1af70a4b |
refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator Removes the 19 committed path->hash manifests, the two per-file size baselines, tests/golden-install-parity.test.cjs, and scripts/gen-golden-install-parity-zcode.cjs. These were pure functions of the source tree (ADR-2719); the differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the sole gate for emitted-artifact propagation. tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs are unchanged (ADR-2719 section 7 exception). Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's existence guard, .gitattributes, package.json scripts, the emitted-provenance totality guard's IO, the differential check's baseline acquisition, CI wiring to publish/restore the baseline artifact, and docs. * refactor(#2724): make the differential attribution check self-sufficient Three fixes required to delete the golden fixtures without breaking CI: - scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs from the three rules that named it. #2759's missingRuleTestFiles guard hard-throws at module load if a rule names a test file absent from disk, which would break the changes job on every PR the moment the fixture-deletion commit landed. - tests/helpers/emitted-provenance.cjs: loadManifests() read the committed golden fixture directory. With that directory deleted at every future ref, this would throw at module load forever, taking the Phase 2 totality guard down with it. Rebuilt from real installer spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest), the same shape emitted-runtime.cjs's currentManifests() already uses. - tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs: the real-tree test's baseline acquisition swaps from baselineManifestsAtRef(base) (git show at a ref that no longer carries fixtures) to resolveBaseline()'s documented precedence: env, then the on-disk cache, then an in-job build. The build fallback (buildBaselineAtRef, new) checks out base into a throwaway git worktree and runs the new scripts/gen-emitted-baseline.cjs there -- no npm ci needed, since bin/install.js and the test helper shells are Node-builtins-only. That script also publishes the baseline artifact from CI's push-to-next job (wired in a follow-up commit). * refactor(#2724): retire the merge-driver bridge and per-file size baselines The Phase 1 bridge (#2721) is retired now that the artifacts it guarded are deleted: scripts/git-merge-regen-driver.cjs, its test, and the 'setup:merge-driver' npm script are removed, and the .gitattributes merge=gsd-regen/linguist-generated block for the three deleted-path globs is dropped. tests/fixtures/install-tree/*.json keeps its normal merge behavior, unchanged (ADR-2719 section 7). scripts/update-size-baseline.cjs and its test are removed: their sole purpose was regenerating tests/workflow-size-baseline.json and tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm script and its step in 'regen:derived' go with it. The per-file baseline describe blocks in tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs are removed for the same reason; the independent loose-tier hard caps are untouched. The differential attribution check's size ratchet (tests/emitted-diff.cjs, already shipped in #2723) is the replacement anti-creep mechanism. 'npm run gen:golden' is replaced by 'npm run gen:install-tree', which keeps regenerating tests/fixtures/install-tree/*.json (the one artifact family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's error messages point at the new command name. tests/golden-parity-single-source.test.cjs's anti-divergence guard (#2266) is retargeted from the two deleted golden-parity consumers to their two replacements (tests/helpers/emitted-runtime.cjs and tests/helpers/emitted-provenance.cjs), which import buildParityManifest the same way — the divergence risk the guard exists for is unchanged. Also wires CI: a new publish-emitted-baseline job runs scripts/gen-emitted-baseline.cjs after a push to next and caches the result keyed on the sha; the test and test-full jobs restore that cache on pull_request events, keyed on the PR's base sha, and export GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree test to pick up. * docs(#2724): flip ADR-2719 to Accepted and update contributor docs Status: Proposed -> Accepted. Regenerated docs/adr/README.md index. CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET. EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET. AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry) no longer point at the deleted golden-install-parity fixtures, size baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver / git-merge-regen-driver.cjs bridge. Editing shipped content now requires zero manual fixture regeneration, documented against the differential attribution check instead of the deleted commands. * docs(#2724): add changeset for removed golden-parity commands * fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs token in CONTEXT.md resolves to a real file. The RULESET. EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks, which the checker reads as a live reference, not historical prose. * test(#2724): retarget ci-test-scope tests off the deleted golden test tests/ci-test-scope.test.cjs asserted specific RULES entries select tests/golden-install-parity.test.cjs, and that every rule selecting it also selects both emitted gates. Both premises broke when the golden test was deleted (#2724): the deleted filename never re-appears in targeted_tests, and there was no longer a third file for the gates to travel alongside. Retargeted the two selection describe blocks to assert tests/emitted-provenance.test.cjs directly (the drift guard the golden gate's rules were retargeted to), and simplified the third block to assert the two emitted gates always travel together, without reference to the golden filename. * docs(#2724): repoint two contributor how-to guides at the differential check Both guides told contributors to regenerate a baseline against tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed at the differential attribution check (tests/emitted-attribution.test.cjs, ADR-2719), which needs no manual regeneration step. * fix(#2724): repair phase6-capstone-conformance's deleted-baseline read An independent orthogonal review caught a real regression this branch introduced into a test file the branch's diff never touched: tests/phase6-capstone-conformance.test.cjs read tests/workflow-size-baseline.json (deleted earlier in this branch) with no fallback, so the whole suite would throw ENOENT the moment this branch landed. The test's actual intent — prove the host-loop workflow files are real, tracked, non-empty docs — is preserved by asserting the live byte count via the same shared counter (scripts/workflow-size.cjs) the size guards already use, instead of a committed snapshot. Also, from the same review: a stale doc comment in scripts/workflow-size.cjs still named the deleted scripts/update-size-baseline.cjs as a consumer, and buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left two fs.rmSync calls unguarded against masking the primary result/error, inconsistent with the try/catch already wrapping the git cleanup beside them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef explaining why it (and its siblings) are kept despite having no production caller post-cutover — they still answer real questions about refs that predate the cutover. * fix(#2724): repair three real regressions found by remote verification 1. tests/emitted-provenance.test.cjs's two hostile-input tests (non-object manifest, unreadable fixture) drove loadManifests(tmp) and monkeypatched fs.readFileSync, both premised on the deleted fixture-directory read this branch already replaced with real installer spawns -- the negative assertions silently stopped firing. loadManifests() now accepts injected {families, install, build, clean} (defaulting to production values), giving the tests a real seam to drive a bad build result and a build failure through the ACTUAL loader instead of a reimplementation, and added coverage that clean() still runs on both paths. 2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE' steps hardcoded shell: bash, which is wrong on windows-latest (native pwsh) and on test-full's macos-latest legs (native zsh per that job's own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning .test.cjs) caught it. Replaced the inline bash script with scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a bare 'node <path>' command line has no shell-specific syntax, so it runs correctly under bash, zsh, and pwsh without a shell override. tests/phase6-capstone-conformance.test.cjs's deleted-baseline read (caught by the same remote run, at a commit prior to this one) was already fixed in d0c3b1242 and is not touched here; verified still passing after these changes. * fix(#2724): revive ADR-1610's new-file size cap inside the differential An isolated review caught a real regression: deleting tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP (ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor) with no successor. tests/helpers/emitted-diff.cjs's size ratchet already 'continue's past any file absent from sizeBaseline -- exactly the files this cap exists to bound -- so a brand-new workflow file sized 32,769-40,960 bytes passed CI clean and shipped, then risked silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted and never referenced anywhere in this branch. Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline, name) signal the growth check already computes -- 'new' is exactly 'present in sizeCurrent, absent from sizeBaseline'. Not ack-able, matching the tier hard caps it sits beside: the fix is extraction, not an acknowledgment entry. Documented, disclosed narrowing: the pure differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering (tests/workflow-size-budget.test.cjs's classification), so a legitimately large new file must extract rather than tier in, one release earlier than an existing file would need to. ADR-1610 itself is left unamended -- this restores its decision rather than re-litigating it. Also fixes a stale comment plus a redundant real 19-installer-spawn assertion left over from the pre-injection-seam version of tests/emitted-provenance.test.cjs's build-failure test, and annotates 3 of 4 stale golden-fixture citations in docs/reference/host-integration-capability-matrix.md as superseded (the 4th is an accurate historical PR narrative, left alone). * fix(#2724): repair three red CI defects on the golden-fixture cutover Windows-only provenance false attribution (defect A): the `hooks-built` provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are Windows-only installer output (ensureCodexHooksJsonSessionStart / ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so the self-attribution resolved to a path that exists on no platform. Only windows-latest ever emits the key, so this only failed there. Fixed by special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated rule) — a dedicated rule would match zero paths, and therefore report as a dead rule, on every non-Windows lane of the same totality guard. `sources` already supported per-match functions; `transforms` is extended to support the same shape so the attribution can vary by match within one rule. Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef` ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but that script is new in this PR and therefore absent at any base ref that predates it — every call failed closed with "Cannot find module". Fixed by running the PR checkout's own generator against the worktree via a new `--dir` parameter, decoupling "which copy of the script runs" from "which tree it measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override, threaded down to `runMinimalInstall`'s new `installScript` override). This is not just a bootstrap fix: a differential needs ONE measurement schema applied to both sides, or the two stop being comparable the moment that schema evolves — running each side's own copy would silently reintroduce that risk. Verified locally end-to-end against real origin/next: resolves a valid {version, sha, manifests, sizes} artifact with the correct sha and no leaked worktree. Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let docs-lint evaluate the fragment for the first time; it already passes (docs/TESTING-SUITES.md and friends already document the removed scripts). Also fixed while in this file: an eslint no-unused-vars warning surfaced by the changed lint run (unused `cleanup` import in tests/emitted-provenance.test.cjs). Added regression coverage for both A and B: a cross-platform spot-check that drives the real hooks-built rule against `.cmd` keys directly (not through a real Windows install), and a real-tree test that drives buildBaselineAtRef against a base ref verified (via git cat-file) to lack the generator, both skipping honestly rather than false-passing when their precondition does not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test Two isolated-review findings on PR #2767: - `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the wrapped `hooks/<name>.js` script, asserting a byte-provenance link that does not exist — traced against buildCodexHookWindowsShimIR (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that same file) flows into the .cmd bytes, never its content. Point `sources` at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention used elsewhere in the table. Since `sources` is checked before `transforms` in the differential, the wrong mapping silently excused any .cmd byte movement caused by editing the wrapped .js file. - The `buildBaselineAtRef` regression test skipped unless a resolvable base ref still lacked scripts/gen-emitted-baseline.cjs — true only until this PR merges, after which every base ref carries the file and the test skips forever with zero ongoing coverage. Rebuilt hermetically: synthesize the missing-generator condition in-place via git plumbing (a throwaway commit, child of HEAD, with just that one file removed from a scratch index), never touching the real working tree, HEAD, or index, and never depending on ambient history or remotes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path The runner container mounts the repo at a path owned by a different uid than the process running the suite, so git's dubious-ownership protection refuses every git operation there. GitHub Actions never hits this because actions/checkout registers the workspace as safe automatically; this runner's container does not. buildBaselineAtRef is the production build-fallback the sole remaining emitted gate depends on (resolveBaseline's in-job-build leg), not just a test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs) that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's worktree add/remove/prune, and the hermetic regression test added in the prior commit — funnels through, plus gen-emitted-baseline.cjs's own rev-parse (now reusing that same wrapper instead of a second execFileSync, so the fix has one source of truth). Each call declares -c safe.directory=<the exact directory it already operates on>, never the * wildcard. Audited every other helper on this surface (emitted-diff.cjs, emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them shell out to git at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d04592de58 |
fix(#2758): select the emitted differential wherever golden-parity runs (#2759)
* fix(#2758): select the emitted differential wherever golden-parity runs Add tests/emitted-provenance.test.cjs and tests/emitted-attribution.test.cjs to every scripts/ci-test-scope.cjs rule that already selects tests/golden-install-parity.test.cjs, so the ADR-2719 dual-run differential travels with the golden on the targeted CI lane instead of being selected by no rule at all. Add an independent module-load totality guard (missingRuleTestFiles) that throws when any rule names a test file absent from disk -- the guard that would have caught the post-Phase-4-cutover hole. It immediately surfaced three pre-existing phantom entries left behind by consolidation epic #1969 (bug-3588/bug-10/bug-3683 filenames folded into other suites months ago but never removed from the rule table); fixed in the same change rather than deferred. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2758): trim unused exports and normalize the path require style Code-review (Standards axis) flagged two judgement-call smells: exporting classify/isInertCi with no caller (Speculative Generality), and requiring path with a node: prefix while the file's other core requires do not (inconsistent style within one file). Both addressed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44707c2c5e |
fix(#2757): attribute transform-driven emitted changes and reclassify agents-verbatim (#2760)
* fix(#2757): let derived provenance rules attribute a transform change The Phase 2 provenance table (#2722) could only explain a moved emitted path via its `sources`, so a `kind: 'derived'` artifact whose TRANSFORM code changed (not its source) was unattributable by construction: 16 emitted agents/*.toml moved by PR #2566's runtime-artifact-conversion.cts change, with zero agents/*.md in the diff. Adds an optional per-rule `transforms: string[]`. A moved emitted path now attributes via its source OR its transform, reusing sourceSatisfiedBy so exact/prefix matching stays identical for both. Declared narrowly on agents-toml-derived and agents-verbatim (the two families verified to route through the same conversion pipeline) as src/runtime-artifact-conversion.cts, src/install-effort-resolver.cts, and src/model-catalog.cts — never bin/install.js, which would be a blanket escape hatch spanning every installer concern. Also corrects agents-verbatim from kind:'identity' to kind:'derived': measured against origin/next's own fixtures, the same agents/<name>.md hashes differently per runtime (e.g. codex vs claude), which a true verbatim copy cannot do. Verified empirically by installing every runtime and diffing against the raw repo source — none reproduce it byte-for-byte; every one rewrites frontmatter and the hardcoded .claude/ self-reference, and claude additionally gets an effort: line injected. A new assertNoIdentityTransforms invariant rejects an identity rule that declares a non-empty transforms list. Closes #2757 * docs(#2757): flag PR #2566's transform-role files for later verification src/agent-tools-contract.cts (added by #2566, absent on next) and src/agent-install-check.cts (modified +112/-1 by #2566, read-only today) were reviewed for prospective inclusion in AGENT_TRANSFORM_SRCS. Excluded for now: the nonexistent file would fail this fix's own "every declared transform path exists" hygiene test, and the modified file's future role cannot be verified without the unmerged PR's diff. Documents the reasoning, the interim ack-file safety net, and the verification method to apply once these files stabilize, so the next PR touching them has a pre-scoped one-line fix rather than a silent gap. Related to #2757 |
||
|
|
9cbb48afa1 |
fix(#2753): scan every key occurrence when asserting a documented default (#2756)
* fix(#2753): scan every key occurrence when asserting a documented default The settings-default assertion located each key with indexOf and checked only a 400-char window after the FIRST match. Its stated contract is "the workflow documents the default for this key"; what it actually asserted was "the first mention of this key is followed by the default" - an ordering assumption that was never part of the contract and that breaks the moment a workflow names a key in prose before its settings-table entry. PR #2558 does exactly that: a shared language directive puts response_language at index 6 while the table documents null at 7185, so the gate failed a document that was correct, and the failure message showed the prose window rather than the cause. Extracted findDocumentedDefault, which scans every occurrence and reports how many it examined. Widening the window to the whole file was rejected - it would pass on any unrelated occurrence of the token - as was parsing the settings table, which would couple the check to table markup. The negative case still fails: a key that no occurrence documents is a failure, now with the occurrence count so a genuine miss stays distinguishable from this false negative. * fix(#2753): guard the empty needle and de-vacuum the newline test Isolated review found a real hang: String#indexOf('', pos) clamps to str.length rather than returning -1, so an empty key made the scan loop stabilize at the end of the document and spin forever. Unreachable from SPEC_FIELDS today, but the docstring claimed termination while reasoning only about self-overlap. Guarded, with the clamping behavior named so the guard is not tidied away later. The newline test asserted only that CRLF and LF agree, which a constant stub satisfies. It now pins the absolute verdict on both, plus a document neither style can rescue. Added the property test the review noted was missing: windowSize is a budget limit, so the verdict is asserted to be exactly "key + gap + default fits the window" over disjoint alphabets. It would have caught the empty-key hang on its own. * fix(#2753): assert examined-occurrence count, not the document's total The no-regression test asserted occurrences === 2 because the fixture mentions the key twice. The scan short-circuits on the first documenting window, so exactly one occurrence is examined - the assertion was describing the fixture rather than the function. Documented the semantic properly instead of just correcting the number: occurrences is how many were EXAMINED before deciding, so on success it is the 1-based position of the match and on failure the document's full count. The asymmetry is deliberate - the failure path is where the number must be trustworthy, since "examined N, none documented it" is what separates a genuine miss from the first-occurrence false negative this function removes. |
||
|
|
0f60266042 |
fix(#2723): reconcile emitted manifest families as a set, not a count (#2750)
* fix(#2723): reconcile emitted manifest families as a set, not a shared count EXPECTED_MANIFEST_COUNT was a single literal 19 asserted against both the baseline (built at the base ref) and the current tree (built at PR HEAD). Those sides legitimately differ by one family whenever a PR adds or removes a runtime, so no value satisfied both: 19 rejected the current side, 20 rejected the baseline side. Every runtime-adding PR was hard-blocked. Replace the shared literal with three independent signals - the derived family set, the recorded fixture set, and the families present at the base ref - reconciled as sets in both directions. A family may appear or vanish only when the diff plausibly touches the runtime registry, and the failure names the family rather than a count. An absolute floor catches the uniformly shrunken universe a same-count self-check passes vacuously. Found by tracing #2005 (Qoder runtime) through the gate during the ADR-2719 dual-run window. * fix(#2723): read the baseline family set from the ref, not HEAD's registry Review found three defects in the first cut. Blocker: baselineManifestsAtRef enumerated MANIFEST_FAMILIES, which is imported at module load and therefore describes PR HEAD. A runtime REMOVED by the PR is already absent from that list, so the base ref was never asked for it, the baseline silently omitted a family that genuinely existed, and the dropped-family check could never fire in production - while its unit tests passed, because they inject the baseline directly. Enumerate from the ref with git ls-tree instead. Also: narrow the registry-signal set to the two surfaces that actually define the family set, since every extra path widens what excuses an unattributed delta; drop the ack bypass, which was a one-sided escape hatch making removals easier to wave through than additions; and gate the derived/fixtures inputs so malformed values return a verdict rather than an unhandled TypeError. * fix(#2723): filter prototype-shaped family names read from git output baselineFamilyNamesAtRef derives object keys from git ls-tree output rather than a trusted constant, so a fixture committed as __proto__.json would turn the manifests[name] assignment into a prototype write. Compared inline rather than through a Set, which is the form the prototype-pollution analysis recognizes. * fix(#2723): require an exact capability path depth and make the ref test hermetic The remote runner went red on both linux lanes with two real defects. The capability signal matched by prefix+suffix, so 'capabilities/capability.json' (no runtime segment) and 'capabilities/a/b/capability.json' (wrong depth) both attributed a family change and would have excused an unattributed delta. Anchored to an exact single-segment pattern. The ref-derivation test reached for this repo's root commit, which is not stable: the remote runner shallow-clones, so rev-list --max-parents=0 returns the grafted boundary carrying every fixture, and this repo has two root commits locally anyway. It now builds its own git repo containing a family absent from the current registry - the real discriminator, and one the root-commit version could never assert. Lint then caught a third: the test called t.after() without declaring t. * fix(#2723): stop asserting ref enumeration against the ambient checkout The remote runner returned [] for the repo's own HEAD while every hermetic temp-repo assertion in the same test passed. That is this function's documented behavior when git cannot read the ref - the runner works from a shallow clone under a bind-mounted workdir - so the assertion was testing the checkout rather than the code. Dropped it. The temp repo already proves the property that matters, and proves it more strongly: it contains a family absent from the current registry, which a registry-derived implementation could never report. The ambient path stays covered by the real-tree test, which skips explicitly when no base ref is resolvable. A git failure is not silently permissive downstream: baselineManifestsAtRef returns null on an empty family set and the real-tree test asserts the baseline is non-empty. |
||
|
|
9138271b5f |
test(#2723): differential emitted-attribution check, dual-run beside the golden (#2737)
* test(#2723): differential emitted-attribution check, dual-run beside the golden Phase 3 of #2719. The conservation law itself, running BESIDE golden-install-parity.test.cjs -- both green, fixtures untouched. Every emitted path whose hash moved between next HEAD and PR HEAD must be attributable, through the Phase 2 table, to a path the PR actually changed. Unattributable deltas fail with the paths NAMED. The only way through is a committed acknowledgment, never a flag -- a contributor facing a red gate sets a flag, which is what UPDATE_GOLDEN=1 is today. The central decision is that the law is a PURE function (no fs, git, installer, or clock), with I/O confined to a separate resolver. The naive one-big-integration-test shape would need ~38 installer spawns per assertion, so #2723's four failing-first criteria would not in practice have been written -- which is exactly how a phase ships promised-but-not-built. Pure, they are millisecond table tests, and the Stryker gate can actually bite. Buckets are conserved: every moved path lands in exactly one of attributed | unattributable | acked, property-tested at 400 runs. A path the provenance table cannot resolve surfaces as an error, never a silent skip. Asymmetries that are deliberate, each with a test: - an ADDED emitted key is a ripple too, not just a modified one - synthesized paths are exempt; code-derived ones are NOT (Phase 2 refused to mark them exempt precisely because exempt means permanently blind) - shrinkage needs no ack; growth does. Gating shrinkage would punish exactly what the size ratchet wants - a STALE ack is a hard failure -- an ack outliving its ripple pre-clears the next one on that path - a failed `git diff` is an explicit error, never an empty changedPaths set; reading it as "nothing changed" would make everything unattributable and produce a failure storm that reads like a real finding - prefix sources are SEGMENT-aware, so `agents/` does not attribute `agentsfoo/x.md` Baseline is cached, not committed, keyed on the next sha. A stale key is refused rather than used: absence fails loudly and gets fixed, whereas staleness produces a confident wrong answer. An explicitly pointed-at GSD_EMITTED_BASELINE that is stale is a hard stop; a stale cache falls through to the in-job build. No baseline-unavailable path returns -- ADR-2719 section 6 names that trap, since in node:test a bare return is a PASS. Both the conservation property and the staleness gate were mutation-verified (injecting a swallowed key fails 9 tests; disabling the staleness comparison fails 5). Fixtures, generators, the merge driver and the ADR status are untouched -- those are Phase 4 (#2724). Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): compute stale acks once, after the size pass Self-review defect found while the reviewers were running. `staleAcks` was computed twice: once between the hash pass and the size pass, then again after. Only the second value was returned, so the first was dead code -- and the dead one was placed where it would have been WRONG. An acknowledgment can be consumed by either a hash move or a size growth. Computing staleness before the size pass reports a legitimate growth ack as stale, which is a false failure that pushes a contributor to delete the very ack that is doing its job. Now computed once, after both passes, with a regression test. Verified by mutation: restoring the early computation fails 2 tests. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): wire the attribution check to the real tree, not just synthetic input An isolated reviewer caught that the first cut was INTERFACE-ONLY: nothing read the ack file from disk, nothing shelled git, nothing built real manifests. Every test was true of hand-built inputs and none of the repo, so the acceptance criterion "both this check and golden-install-parity green on the same tree" was trivially true rather than meaningfully true. That is the promised-but-not-built failure this epic keeps finding in its predecessors, recurring one phase later for the wiring itself. Taken, not argued. Adds tests/helpers/emitted-runtime.cjs -- the only module that touches git, disk, or the installer -- and an integration test that runs the same pure law against reality: - CURRENT side: 19 real installer spawns via runMinimalInstall + buildParityManifest, the same machinery the golden harness uses. - BASELINE side: `git show origin/next:<fixture>`. That is next's RECORDED emitted state and it costs nothing. Deliberately NOT the working-tree fixtures, which are whatever this PR's author regenerated -- comparing against those would be vacuous. Phase 4 deletes the fixtures and swaps in resolveBaseline's cache path, already implemented and tested. - changed paths from real `git diff --name-only origin/next...HEAD`, with the git subprocess bounded at 30s per CLAUDE.md's unbounded-subprocess rule. - the real tests/emitted-drift-ack.json (absent is legal; present-but-empty or unparseable throws rather than being read as absent). Verified it can actually fail: an uncommitted edit to a shipped workflow moves emitted output but never appears in the committed diff, and the check names all 18 affected emitted paths with the message format ADR-2719 §1 specifies. Restores clean. Also from review: - readAckFile now has a real test exercising the SUT across absent / valid / empty / unparseable / unreadable. The previous test asserted fs behaviour rather than SUT behaviour, because no SUT ack-reading path existed yet. - formatReport's sampleLimit gains true limit-1/limit/limit+1 coverage at 19/20/21. A test was previously NAMED "(limit+1)" while testing no numeric limit at all, which is worse than no coverage because it reads as covered. Windows uses an explicit t.skip (install output is platform-specific there, mirroring the golden harness) -- never a bare return, which node:test scores as a PASS. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): cover the claude-local manifest family in the real-tree check Isolated adversarial review, MAJOR. The real-tree wiring enumerated Object.keys(RUNTIME_META) -- 18 entries -- while the emitted manifest set has 19 families. The 19th is claude-local: claude is the reference host and the only runtime with a distinct LOCAL "legacy flat-commands" layout (commands/gsd-*.md + agents/gsd-*.md at project scope), which golden-install-parity.test.cjs guards with a hand-coded test outside its RUNTIME_META loop (#2086). The family was dropped from BOTH sides, so the test's own self-check (current.length === baseline.length) passed vacuously at 18 === 18. A PR changing Claude's local-scope output would have failed the golden while this check reported ok -- and that disagreement is precisely what the dual-run window is designed to surface as a provenance-table hole. A wiring omission masquerading as one is the worst available failure here, because it would have been read as evidence about Phase 2 rather than a bug in Phase 3. Fixed by deriving MANIFEST_FAMILIES explicitly (18 global + claude-local at local scope) instead of inferring the set from RUNTIME_META. The self-check is also repaired: it now asserts both sides against the INDEPENDENT EXPECTED_MANIFEST_COUNT from the Phase 2 table, and asserts claude-local specifically. Comparing the two sides to each other can never catch a family missing from both -- the assertion has to come from outside. Verified by mutation: removing claude-local again fails the test. Also from the same review: - sourceSatisfiedBy returns the matched source string, so an empty-string source would return '' and the caller's `if (hit)` would silently discard a real match. Unreachable today (every rule source is a non-empty template) but a footgun for the next rule author; now `!== null`. - the purity fixture used a single-element changedPaths array, so an in-place sort would have been invisible. Now three elements in unsorted order. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2723): resolve the base ref tolerantly instead of hard-requiring origin/next The first matrix run failed on both linux lanes: differential attribution over the real tree cannot resolve origin/next (Command failed: git rev-parse origin/next) Not a flake, and not an environment excuse -- a real defect in this diff. The gsd-test runner shallow-clones and merges base+head, so no origin/* remote- tracking refs exist in the container. My own fail-loud path fired correctly; what was wrong was hard-depending on that ref existing. GitHub Actions has the same shape by default, which is exactly why changeset-required.yml carries an explicit `git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"`. Now resolved through an ordered candidate list -- GSD_EMITTED_BASE (explicit lane override), then origin/$GITHUB_BASE_REF and $GITHUB_BASE_REF, then origin/next and next -- de-duplicated, each verified with `rev-parse --verify <ref>^{commit}`. When NO candidate resolves the test takes an explicit t.skip() naming every ref it tried and stating that the gate did not run here. That is the ADR-2719 section 6 distinction: t.skip is REPORTED as skipped, whereas a bare return is scored as a PASS. Hard-failing was the other option and is wrong -- it would make the suite permanently red wherever a base ref cannot exist by construction, which is a statement about the checkout, not a propagation finding. The candidate ordering is pinned by a unit test rather than left implicit, since the ordering IS the fix. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f6822ccba |
test(#2722): emitted-artifact provenance table with a totality guard (#2735)
* test(#2722): emitted-artifact provenance table with a totality guard Adds the declarative emitted-path -> source-path table that ADR-2719 §2 specifies, plus the totality guard that keeps it honest. Phase 2 of #2719. Every emitted path across all 19 committed golden-parity manifests (8,524 paths) must match exactly one rule. Zero matches, two matches, and a rule matching nothing are all hard failures, so a new emitted family fails the build loudly instead of passing through unattributed. The measured surface is larger than #2722 estimated from claude.json alone (26 top-level families across 19 runtimes, not 13), which is itself what the totality guard exists to surface. It resolves to 19 rules. Building the table caught three false attributions that were total but resolved to repo files that do not exist -- Copilot's `<name>.agent.md` rename, Kimi's code-literal `agents/gsd.{yaml,md}` root agent, and Copilot's `hooks/gsd-session.json` registration. The "every attributed source exists" test is therefore a first-class gate, not a nicety. Notable correctness decisions: - Emitted shapes are hard-coded; deriving them from the installer would make the guard tautological (it would follow any installer change silently). Only source paths read a first-party descriptor, and only where the descriptor is the sole declaration (hostBehaviors.nativePlugin.source). - Emitted skills attribute to commands/gsd/*.md, NOT the repo skills/ dir -- that directory is generated from commands/gsd by gen-plugin-skills.cjs, so attributing to it would be false attribution that still passes totality. - Attribution is keyed on (rel, runtime): plugins/gsd-core.js has different sources for opencode and kilo. - Rule order carries no semantics (property-tested), since exactly-one matching is enforced rather than first-match-wins. Nothing here reads a git diff, builds a live manifest, or touches a fixture; the differential check, drift-ack file and size ratchet are #2723. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * docs(#2722): record the delivered provenance table in the CONTEXT.md glossary The `### Emitted Artifact Provenance` entry landed in #2721 describing the table as future work. Phase 2 delivers it, so the glossary now records what actually exists and the invariants #2723 must preserve: - where the table lives, its rule count, and that it is total over all 8,524 emitted paths across the 19 manifests - dead-rule detection, so table rot is loud in both directions - the corrected surface measurement (26 families, not the 13 estimated from claude.json alone) - the two invariants #2723 inherits: shapes hard-coded (deriving them would make the guard tautological), and attribution keyed on (rel, runtime) - the skills/ false-attribution trap, and that totality does NOT catch a wrong-source rule — the source-existence assertion is what does Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2722): close five review findings on the provenance table Two orthogonal review passes plus an isolated adversarial reviewer returned findings at minor..major. All fixed; no blockers were raised. Standards axis (CONTEXT.md:456, RULESET.TESTS.guard-toplevel-readFileSync): - module-level loadManifests() threw at require time before any test() registered, turning a missing fixture dir into an opaque crash instead of one named failure. Now a memoized lazy accessor. - matchRules and assertTotality each carried their own copy of the matching loop; assertTotality now calls matchRules. That is the #2266 divergence class, and two copies could let the guard and the attributor disagree. - named the corpus stride constant; dropped an inline require. Spec axis: - the CONTEXT.md glossary carried a "26 families" figure that is not reproducible from the code and that no test pinned -- a hand-maintained number in permanent canon, i.e. exactly the silent drift this epic exists to end. All volatile counts are now removed from the glossary, with the reason stated inline: the guard recomputes them every run, so they belong in a failure message, not in prose. No test was added to pin the count, because that would rebuild the brittle committed number we are deleting. Isolated adversarial review: - `.+` tail captures let a `..` segment reach a constructed source path that resolves outside the repo. Not live-exploitable (fixtures are committed and the only consumer is an existsSync probe) but Phase 3 feeds these strings into a diff-consuming check, so assertSafeRelPath now fails closed once, in matchRules, rather than per-rule. - attributeEmittedPath's ambiguous-match branch was never exercised; only assertTotality's parallel path was. Now tested directly. - sampleLimit's truncation branch had no limit-1/limit/limit+1 coverage. - the fast-check property could not fail for the reason it was named for. That last one took two attempts and is the one worth reading. The property hand-rolled its shuffled side from the per-rule matchOne primitive, which is order-independent by construction, so it held for reasons unrelated to the shipped matchRules. Routing it through the real matchRules was still not enough: on an unambiguous table, first-match-wins and collect-all return identical results for every path (measured: 0 of 190 corpus paths differ). Order can only matter where more than one rule matches, so the property now also asserts that an intentionally ambiguous table reports BOTH hits as a set under every permutation. Verified by mutation -- injecting a `break` into matchRules makes it fail, and restoring makes it pass. Enabling all of the above: matchRules and attributeEmittedPath now take an injectable rules table, so tests can drive the real code path instead of re-implementing it by hand. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
80778e2674 |
fix(#1881): report an unreadable ROADMAP instead of reading it as absent (#2729)
* test(#1882): stage one file per commit in the base-ref ancestry fixture CI failed on ubuntu-24 inside this test's setup loop, before any code under test ran: at commit 32 of 60 the index referenced a blob whose object write had not landed -- "invalid object ... for 'base-31.txt' / Error building trees". The loop staged with `git add .`, which re-stages every file already in the tree. Across 60 iterations that rehashes O(n squared) blobs -- roughly 1,800 stagings and 60 full index rewrites to add 60 one-line files -- and that churn is what the object store failed under. Each commit only ever adds a single new file, so staging that one path is equivalent and removes the redundant work entirely. Verified the loop still builds the intended history: 61 commits, git fsck clean. The fixture already carries a note from an earlier fix in this epic recording that it passed on ubuntu-22 and windows-24 and failed on ubuntu-24 for the same commit. That was a different stage -- fetch versus diff -- but the same lane and the same brittleness, so this is the second time this fixture's cost has surfaced as a red build rather than as a test failure. Not caused by this PR's change, which touches two configuration lists and cannot reach a scratch git repository in tmpdir. Fixed here rather than deferred, because the run surfaced it. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1881): prove an unreadable ROADMAP is indistinguishable from an absent one Failing-first. Encodes the issue's runtime repro: an unreadable ROADMAP.md makes getRoadmapPhaseInternal return the same null it returns for "phase not found", and getMilestoneInfo return the same {v1.0, milestone} it returns for a project with no roadmap at all -- so a permission or I/O fault reads as a brand-new project. Half these cases exist to hold the opposite line. getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT is converted to a synthetic Error carrying no errno, and that lands in the SAME catch as a real EACCES. Reporting unconditionally there would flag every project without a ROADMAP.md -- every brand-new project -- as corrupt. The absent case, the errno-less error, a non-string errno, unparseable content and a genuinely missing phase are all pinned silent. One case guards a decision rather than behaviour: an unreadable STATE.md alone must stay silent, because the inner catch that swallows it is deliberate and documented under the #2245 audit as an optional enhancement falling back to ROADMAP-only heuristics. Two more pin the invariant ADR-1411 names explicitly -- neither function may throw, because src/state.cts removed its own defensive try/catch on the strength of that guarantee. Assertions are on the frozen reason enum and the emission counter, never on diagnostic prose. Faults are injected by overriding the platformReadSync seam and restoring in t.after(), never chmod 0o000, which root bypasses. Adds the ROADMAP_UNREADABLE reason to the shared vocabulary as scaffolding; no call site emits it yet, which is what makes these tests red. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1881): report an unreadable ROADMAP instead of reading it as absent getRoadmapPhaseInternal returned null for a read failure exactly as it does for "phase not found", and getMilestoneInfo returned {v1.0, milestone} exactly as it does for a project with no roadmap -- so a permission or I/O fault presented as a brand-new project and workflows synthesised a blank phase or skipped requirement extraction with no signal. Both return values are preserved exactly, per ADR-1411's amendment: continuity is correct, the silence was the defect. Each catch now reports through the shared unusable-input seam that shipped with #1882 rather than a second copy of the same mechanism. The discriminator is the errno, and it is load-bearing in the silent direction. getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT is converted into a synthetic Error with no code that lands in the same catch as a real EACCES. Reporting unconditionally there would flag every project without a ROADMAP.md -- every brand-new project -- as corrupt. A genuine read fault always carries an errno; absence never does. The parse is regex over text and cannot throw, so nothing else reaches these catches. Neither function gains a throw. ADR-1411 names this explicitly: src/state.cts removed its defensive try/catch around getMilestoneInfo under the #2245 audit because it never throws, and two tests pin that. The inner STATE.md catch stays untouched and silent -- its fallback to ROADMAP-only heuristics is a deliberate, documented optional-enhancement path, not a fault. Where the fix belongs was the design question. platformReadSync does not leak: it keeps absent and unusable as two channels, exactly as an abstraction should. Both callers re-collapsed that distinction, so the fix is caller-side and the projection seam -- with roughly ninety other dependents -- is untouched. Closes #1881 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1881): admit the roadmap reason to the locked vocabulary The seam documents adding a reason as three coordinated changes -- the enum entry, the emitting call site, and the test that locks Object.keys(...).sort(). This PR made the first two and the lock caught the third, which is the whole point of pinning the key set rather than asserting each value exists. The roadmap suite no longer re-locks the full set. Two complete locks would mean two files to update every time a later phase adds a reason, and #1883 and #1884 are both going to. The canonical lock stays in the seam's own suite; the roadmap suite asserts only the value it introduces. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1881): resolve the roadmap path inside the try, not outside it Naming the file in the diagnostic required the resolved path in the catch, and the obvious way to get it was to hoist `path.join(planningDir(cwd), 'ROADMAP.md')` above the try. planningDir throws a plain Error for an invalid GSD_WORKSTREAM or GSD_PROJECT segment -- one containing a slash, backslash or `..` -- so hoisting it let that throw escape uncaught. That broke the exact invariant ADR-1411 names as this file's hazard: src/state.cts removed its defensive try/catch around getMilestoneInfo under the #2245 audit because that function never throws. Of its callers only archivePhaseDirectories wraps it; cmdInitExecutePhase, cmdInitNewMilestone, cmdInitMilestoneOp, cmdInitManager, cmdInitProgress, cmdProgressRender and cmdStats all call it bare, so a workstream name with a slash in it crashed the CLI outright instead of degrading. The previous commit asserted "neither function gains a throw -- two tests pin that". That was false. Both tests inject faults through platformReadSync only and never through planningDir, so neither could have exercised the path that broke. The guarantee was claimed, not demonstrated. The path is now declared before the try and resolved inside it, so the catch can still name the file when there is one, and a path that never resolved reports nothing and returns the sentinel unchanged. The two test names are narrowed to what they actually prove -- that a failing READ does not throw -- and a new case injects the planningDir failure directly, which is what would have caught this. getRoadmapPhaseInternal carried the same hazard, resolving the path outside its try since before this branch. It is fixed the same way rather than left: ADR-227 is explicit that throwing breaks pipeline continuity, this read path already degrades to null for every other failure, and a PR whose purpose is hardening this invariant is the wrong place to leave the sibling crashing. Behaviour otherwise unchanged and re-verified: healthy lookups, EACCES reporting on both functions, absent-roadmap silence, and the errno discriminator all unaffected. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1881): backfill changeset pr number to 2729 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
60bc6ddd6e |
fix(#2568): commit the debug session doc on the manager-driven terminal path (#2731)
* test(#2568): failing-first contract for the debug manager's doc commit agents/gsd-debug-session-manager.md contains zero occurrences of 'commit', so on the manager-driven path — the normal /gsd-debug flow — nothing consults commit_docs and session docs are left untracked. The step exists only in gsd-debugger.md, which does not reach the end of a multi-cycle session. Two layers. The spec placement is asserted against the shipped agent text because that text is what the orchestrator executes; the commit_docs gate the fix relies on is exercised behaviorally through the CLI in temp git projects, because 'the CLI no-ops when disabled' is the assumption that makes calling it unconditionally correct — asserting it in prose would be assuming the load-bearing part. The most important case is the negative: CONTINUE_REQUIRED is non-terminal and must NOT commit. A fix that satisfied the positive cases by committing unconditionally would strand a half-finished session looking done, which is worse than the bug. RED expected on the five spec assertions; the three CLI gate tests pin existing behavior and pass both sides. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): commit the debug session doc on the manager-driven terminal path agents/gsd-debug-session-manager.md contained zero occurrences of "commit", so on the normal /gsd-debug flow nothing ever consulted commit_docs and session docs were left untracked — reproduced by the reporter across three consecutive real sessions. An obligation left behind during a responsibility move. gsd-debugger.md commits the doc at :1214, but that runs only when the debugger carries a fix to completion inside a single spawn. The multi-cycle manager took over checkpointing, fix application, archival to resolved/, and the terminal summary — every step that finishes a session — without the commit step. The debugger's copy still exists and still works on its own path, so "is this handled anywhere" answered yes. The manager now commits before both terminal shapes, and explicitly NOT on CONTINUE_REQUIRED. That exclusion is the load-bearing part: the manager's own contract already forbids fabricating a terminal summary on that path, and committing there would be the same lie in git form — a half-finished session stranded looking done. CHECKPOINT REACHED likewise does not commit. Two facts kept the fix small. The manager already carries the gsd_run preamble (:96, used at :97 for resolve-model), so no new plumbing. And cmdCommit (src/commands.cts:809-814) already gates on commit_docs and returns skipped_commit_docs_false when disabled — so the agent calls it unconditionally and correctness follows from the CLI rather than from a second copy of the config check that could drift. Calling git commit directly was rejected for exactly that reason: it would bypass a user's explicit commit_docs:false. In-session fix code is staged by specific file, never git add -A, which would sweep unrelated working-tree changes into a debug commit. Body-only edit — frontmatter untouched, so the research-profiles/AGENTS.md ripple does not apply. gsd-debugger.md's own commit step is preserved; the single-spawn path still ends there, and a double commit is harmless since the second finds nothing to stage. Fixes #2568 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): bind the commit paths, define gsd_run in-block, and make the fix commit idempotent Three defects in my own first cut, all found by the isolated adversarial pass. 1. BLOCKER — {debug_dir} was a dangling substitution. I copied it verbatim from the issue's suggested patch without checking scope. It is a real variable in gsd-core/workflows/debug.md, where the ORCHESTRATOR receives it from init JSON and uses it to build debug_file_path — but this agent never receives it: <session_parameters> declares only slug, debug_file_path, symptoms_prefilled, tdd_mode, goal and specialist_dispatch_enabled. An unbound token makes --files resolve to a nonexistent path; cmdCommit skips missing explicit files, staging stays empty, and the doc silently never commits — reproducing #2568 through a different broken path. Same class as #2684's dangling placeholders. Now spelled literally as .planning/debug/resolved/{slug}.md, matching gsd-debugger.md:1214 and this file's own prose. 2. MAJOR — gsd_run was undefined in the commit block's shell. Shell state does not persist across tool invocations, and the Step 2 preamble is ~230 lines and an entire spawn-and-loop earlier. gsd-debugger.md redeclares the full preamble immediately before each of its own call sites; the commit block now does the same, byte-identical to this file's existing definition. 3. MAJOR — the in-session fix commit was not idempotent. gsd-debugger.md's archive_session already commits the fix on the confirmed-checkpoint path, which is the standard find_and_fix flow, so `git add X && git commit` would hit an empty diff, exit non-zero, and abort the step before the summary was returned. Guarded with `git diff --cached --quiet || git commit`. The doc-commit half was already safe — query commit treats an empty diff as nothing_to_commit and exits 0 — and that asymmetry is now stated rather than assumed. Three tests added for exactly these, since the reviewer correctly noted the suite would have caught none of them: every token substituted into a commit command must be a declared session parameter; the preamble must sit in the same block as the call with no step boundary between; and the fix commit must carry the staged-content guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): keep the single canonical gsd_run preamble The previous commit added a second preamble to the commit block, on the isolated review's theory that shell state does not persist across tool invocations. That theory is reasonable in general and wrong for this corpus: the repo enforces "each agent .md using gsd_run contains exactly ONE canonical preamble, before the first gsd_run call". Adding a second broke that invariant and five suites with it — the B-agents preamble check, runtime-launcher-parity, and three slash- namespace guards. The single Step 2 preamble already precedes the commit call, so coverage was never actually missing. Reverted to one, and the test now asserts the real property — exactly one preamble, positioned before the call that needs it — rather than the locality I had wrongly encoded. Recorded because the direction of the mistake matters: the review was right that the question needed asking and wrong about the answer, and I shipped the wrong answer without checking the invariant that already governs it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): use the canonical /gsd:debug slash form in the new prose My explanatory sentence wrote `/gsd-debug`, the retired syntax. Three slash-namespace guards caught it: the #3443 invariant, the #1975 folded bug-2543 check, and the retired-syntax scan over Claude-facing source. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2568): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a613caaeef |
enhance(#2721): regenerating merge driver, regen:derived, and a name for the emitted-artifact family (#2730)
* test(#2721): failing-first suite for the gsd-regen driver and CONTEXT.md parity Tests precede the implementation per the TDD gate. The driver module does not exist yet, so tests/git-merge-regen-driver.test.cjs fails at require time; the contributor-standards parity assertions fail against next as it stands today, where the standards doc names two CONTEXT.md headings that have never existed. Refs #2721 * feat(#2721): add the gsd-regen merge driver and regen:derived The golden parity manifests and the two size baselines are pure functions of the source tree, so their only correct merge is "recompute" -- something git's ours/theirs interface cannot express. 140 of 143 conflicted-file instances across the open PR queue are these files. The driver deliberately does NOT regenerate. Four probes established that at merge-driver time neither the working tree nor the index reflects the merge: both hold the ours side, a file added by theirs does not exist yet, and MERGE_HEAD is unwritten. Git also invokes the driver once per conflicted path (20 here). A regenerating driver would therefore read the ours-side tree and emit a plausible-but-wrong hash manifest -- worse than a conflict, because a conflict is visible. So it accepts %A, runs zero subprocesses, records the resolved paths, and prints one notice pointing at npm run regen:derived. Staleness stays caught where it already was, by golden-install-parity in CI. Every failure path degrades toward today's behaviour (a normal conflict). install-tree is deliberately excluded per ADR-2719 section 7. Also folded in, per the no-defer rule: workflow-size.cjs claimed .md files have no eol=lf in .gitattributes; git check-attr shows eol: lf, set by .gitattributes line 2 since #1088. Refs #2721 * docs(#2721): document regen:derived and the gsd-regen merge driver Adds the how-to a contributor actually reaches for when the generated parity manifests or size baselines conflict, in both places they would look: the merge-conflict path in CONTRIBUTING.md and the full guide in TESTING-SUITES.md, including what the driver deliberately does not do (it does not clear GitHub's CONFLICTING label, and it does not regenerate mid-merge). Also scopes the new contributor-standards parity assertion to the doc's own CONTEXT.md section. Its first run flagged `## Decision`, `## Consequences` and `## Standards followed`, which the doc attributes to an ADR body and a PR body rather than to CONTEXT.md -- a doc-wide extractor would have demanded CONTEXT.md grow headings that do not belong to it. Refs #2721 * fix(#2721): stop passing %P to the merge driver — shell injection The isolated adversarial review found, and I independently reproduced, local arbitrary command execution. Git does not invoke a merge driver with an argv array. It substitutes %O %A %B %L %P textually into the configured string and runs the whole thing through a shell, and $(...) executes inside POSIX double quotes -- so quoting the placeholder does not neutralise it. %O/%A/%B are git-generated temp names and %L is an integer, but %P is the file's own path, chosen freely by any contributor. A branch renaming a covered fixture to evil$(touch PWNED_SENTINEL).json executed that command on the machine of every maintainer who merged it, and the merge still reported success. Fix removes the input rather than filtering it: %P is no longer registered, so the driver receives no attacker-controlled argument at all. The marker records a count instead of path names. A metacharacter filter would have been a guess about shell grammar; passing nothing is a property. Re-ran the identical exploit against the fixed command: nothing executed, conflict still resolved. Two regressions guard it -- a platform-independent assertion that the registered command carries no %P, and a real merge driven by the actual planInstall output with a $(...) filename. Also from review: CLI dispatch had no coverage at all (CONTRIBUTING's "CLI and command routing" matrix), which is why runInstall/runStatus now take {repoRoot} -- hardcoding REPO_ROOT was what made them untestable. Renamed planResolution to resolveAndRecord since the plan* prefix promised purity it did not have. Reconciled the eleven-vs-twelve generator count across CONTEXT.md, CONTRIBUTING.md and the changeset. Refs #2721 * test(#2721): scope safe.directory for the check-attr helper The 66f4d85a run failed 11 assertions, all in the .gitattributes scoping block, with "fatal: detected dubious ownership in repository at '/work'". The test container checks the repo out at a path its user does not own, so git refuses check-attr outright. Everything else passed (27,185). `check-attr` is a pure read of .gitattributes -- no hooks, no filters -- so the exemption is scoped to that one invocation. It is deliberately NOT applied to the driver's own production `git config` calls, which run in the user's own clone and should keep the protection. Refs #2721 * test(#2721): delete the stale assertion that the driver command carries %P The plex2 run on bdfd0856 left exactly two failures, both this test: it still asserted the pre-fix command string, i.e. the vulnerable behaviour. Deleted rather than relaxed, per RULESET.TESTS.delete-bad-tests -- its useful half is already covered, in both directions, by registeredDriverCommandNeverPassesThePlaceholderForTheFilePath. Refs #2721 * test(#2721): drive the end-to-end merges from the real planInstall output The e2e helper hand-rolled its own driver registration, and still carried %P. That meant the five real-git tests were not exercising the production command string at all -- planInstall could drift and they would keep passing. They now register exactly what a contributor gets from npm run setup:merge-driver. Refs #2721 * chore(#2721): backfill changeset pr number to 2730 |
||
|
|
e48eb44003 |
fix(#1856): give the executor-worktree refusal a handoff instead of a dead end (#2727)
* test(#1856): failing-first contract for the orchestrator cwd-drift guard handoff The #48 guard correctly refuses to execute waves from an agent worktree, but the refusal is a dead end: that worktree can hold committed fixes AND uncommitted work, and "re-run from the orchestrator's own worktree" silently means abandoning them. The reporter was left choosing between continuing from a blocked worktree and losing the work. The guard is shell embedded in execute-phase.md, so these tests extract the block by a stable marker and EXECUTE it against real git fixtures — the shipped text is the runtime contract. Covers the stranded-commit and dirty-tree report, the integration commands, both agent- namespaces, commit-count boundaries 0/1/2, and the constraints the guard's own comment records: it must NOT fire on an ordinary branch, on 'agentic-refactor', or on a legitimate feature worktree under .claude/worktrees/, and must degrade cleanly with no resolvable base or a detached HEAD. RED expected: the marker does not exist, so extraction fails and every case errors. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#1856): give the executor-worktree refusal a handoff instead of a dead end The #48 cwd-drift guard correctly refuses to execute waves from an agent worktree — its comment records why ("this is how a wrong-base merge nearly shipped ~1000 files"), and that refusal is untouched here. The defect is that it was a dead end. At the moment it fires, the worktree can hold committed product work, uncommitted product and planning changes, and the live gap-planning context. Telling the user to "re-run from the orchestrator's own worktree" silently means abandoning all of it, because the orchestrator worktree cannot see commits that live only on the agent branch. The reporter was left choosing between continuing from a blocked worktree and losing five commits plus uncommitted work. The refusal now reports what is actually stranded — the commit count and log against the resolved base, and the uncommitted files — followed by the concrete integration sequence (commit here, switch to an orchestrator-safe checkout, merge or cherry-pick, re-run) and a verify command. Nothing is claimed that is not there: a clean worktree with no commits ahead prints the plain refusal with no empty sections. Every added command is diagnostic and `|| true`-guarded, so a failure degrades to the original refusal rather than crashing before the message prints. Verified: an unresolvable base still refuses cleanly. Deliberately NOT done: auto-merging or auto-cherry-picking the agent branch. That is precisely the operation #48 exists to stop the orchestrator performing from a drifted cwd, at the moment it has least confidence about which tree is which. Reporting beats acting here. The guard block carries a `gsd:guard=orchestrator-cwd-drift` marker so the new contract test can extract and EXECUTE the shipped shell against real git fixtures rather than asserting on its characters. Fixes #1856 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#1856): report true counts, add the changeset, and document the seam Review findings, all from the isolated adversarial pass: - The dirty-file list was capped at 20 with no indication, so a worktree with 27 uncommitted files reported 20 — under-informing the user about exactly what is stranded, which is the entire point of this change. Both lists now count BEFORE truncating and print "… and N more". Verified with 25 commits / 27 dirty files. - The has-commits condition was written out twice and could drift on a future edit. Collapsed to a single _WT_HAS_COMMITS flag. - The changeset fragment existed but was untracked, so it was in neither commit on this branch and the PR gate would have failed against real history. - CONTEXT.md:122 documents this exact seam ("the orchestrator runs a cwd-drift guard at execute_waves entry…") and was not extended. Now records the handoff report, that the refusal condition and exit code are unchanged, and that every added command is diagnostic and || true-guarded. Verified NOT a defect, correcting the review's premise: the guard block does break when its line endings are CRLF, but .gitattributes:2 is `* text=auto eol=lf`, which OVERRIDES core.autocrlf and forces LF on checkout on every platform including Windows — so the shipped file is LF there too, and the installer copies it through Node without translating endings. The reproduction (mine and the reviewer's) required injecting CRLF by hand. It is also not fixable from inside the script: a \r breaks the shell parse at the block's first line, before any #1856 code runs. Neither introduced nor amplified by this change. Also noted and left as-is by design: the review flagged that #1856's "offer an explicit recovery option" could be read as requiring an interactive/automated integration rather than printed instructions. Deliberate — see the commit that added the block: auto-merging is the exact operation #48 exists to prevent the orchestrator performing from a drifted cwd. Called out in the PR body for a maintainer decision rather than silently chosen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#1856): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
90ba0ef10b |
docs(#2720): submit ADR-2719 — emitted-artifact attribution design contract (#2726)
Replaces the committed golden-install-parity hash manifests and per-file size baselines with a computed conservation law: every emitted path whose hash moves must be attributable, via a declarative provenance table, to a path the PR actually changed. Supersedes ADR-2264 Decision §2-§4 and its Amendment; ADR-2264 Phase 1 (the single-source buildParityManifest and exclusion constants) is retained and depended upon. Satisfies ADR-2264 AC1 rather than rewording it away, per that ADR's own 2026-07-17 audit. Docs-only. Both sides of the supersession edited together; ADR index regenerated with gen-adr-index.cjs --write. Closes #2720 Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
09477f925e |
fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading The Workflow backend emitted every agent() call with no model, so model_overrides / model_policy / model_profile were silently inert on that path while the inline path honored them (ADR-1411). Neither existing suite contained the string 'model' at all. The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor') rather than hardcoding either, so it asserts backend parity rather than a fixed string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids reaching the code generator, the #2285 composed seam, CLI config-defaulting, and a fast-check round-trip property. RED expected: no model key is emitted anywhere, and --executor-model does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): thread the resolved executor model into the Workflow backend The Workflow backend emitted every agent() call with no model at all, so model_overrides / model_policy / model_profile_overrides / model_profile were silently inert on that path while the inline path honored all of them. The model was not dropped at the last step — it was absent from the whole seam: agentOptions() took no model, EmitInput had no field to carry one, and ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could not forward one. The generated script asserted the parity it broke. VERIFY-FIRST, which #2686 flags as the question that decides the fix: the Workflow tool's agent() DOES accept a per-call model. Its documented signature is agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? }) so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled out. ADR-1143:24's option enumeration omitting `model` is an incomplete enumeration, not a decision to exclude it. - agentOptions(p, executorModel) emits `model` only when it is a non-empty string that is not "inherit" (#2517: an empty model 404s on runtimes without native tier aliases). A non-string is a malformed config: omit, never throw. - executorModel threaded through EmitInput and ResolveWaveDispatchInput. - The CLI resolves gsd-executor from project config by DEFAULT rather than requiring a flag, reading the same source the inline path reads. An orchestrator that never learns about a new flag would otherwise silently keep the old bug. --executor-model exists only to pin/override. - ADR-1411 provenance: the generated header now states which model was applied, or that none resolved and why. A fallback must be a visible value. Compatibility: when nothing resolves, the emitted options object is byte-identical to before, so every existing caller and assertion is unaffected. Behavior change (Hyrum's Law): opted-in users move from session inheritance to the catalog-resolved executor model. Adding a `model` key also changes agent() opts, which invalidates the cached prefix of any in-flight resumeFromRunId run — a one-time re-execution. Both disclosed in the changeset. Fixes #2686 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): reject script-breaking model ids and share the emit predicate The isolated adversarial review found a BLOCKER in my own provenance comment, proven by execution (the emitted script exited 42 from an injected statement). U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment in EVERY engine — the ES2019 change legalized them inside string LITERALS only. So quoteString (JSON.stringify) is sufficient for the `model: "..."` object literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in a model id closed the comment and made the rest of the line live top-level code. The value is reachable from `.planning/config.json` (model_overrides / model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any non-claude runtime — attacker-influenceable in a cloned repo. `emitWorkflowScript` now rejects a string executorModel carrying any character in UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to phaseDir/runId, which is proof the codebase knew this hazard. Rejection is ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an emit failure to the inline backend WITH that reason, so the degradation is visible. A non-string stays on the existing defensive path (omit, never throw) — that is malformed config, not an injection attempt. Also from the reviews: - The predicate deciding "is this model emittable" was duplicated between the emission and the comment asserting it. Extracted to emittableModel() so a generated comment can never claim something the generator did not do — the exact failure class #2686 was filed for. - That predicate now trims and lower-cases before comparing, closing a real #2517-class gap: " " and "INHERIT" were previously emitted verbatim. - The adversarial test was pass-always against this very vulnerability — it asserted only that JSON.stringify appeared. Replaced with the real contract (rejection) plus an execution-level check that no LineTerminator survives into the comment. A raw U+2028 had also been committed into that test's fixture array where a tab was intended; both are now explicit \u escapes. - optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated model contained — silently not testing what it claimed. Now brace- and string-aware. Stale-test corrections in tests/fix-2285-*: three assertions froze the exact options literal `{ agentType: "gsd-executor" }`. The object legitimately gained an optional additive `model` key, so they now assert the invariant they exist to protect (agentType present, isolation absent) rather than a frozen literal. The CLI-vs-pure equality test pins --executor-model on both sides; otherwise it compared a config-resolved CLI run against a pure call given no model. CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): fix the options extractor and model the rejection path Two defects in my own test helper, caught by the full matrix: - optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The emitted shape is agent("brief", { ... }), so that never matched and the helper returned an empty array, making every assertion over it vacuously true. It now anchors on agent( and takes the first balanced, string-aware {...} after it. - The fast-check property predated the security fix and asserted ok:true for any generated string. Strings carrying an unscriptable character are now rejected, so the property models the real three-way contract: unscriptable -> ok:false; trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the trimmed value. Verified locally against the built module: omit values clean, both plans carry the model on the parity path, property passes 500 runs at seed 42. Test file re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are explicit \u escapes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): scope no-control-regex on the mirrored unscriptable-char class The class is the point of the assertion — those bytes are exactly what must be rejected — so the rule is disabled at that line rather than the class weakened. UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence the mirror. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2686): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1c93df04db |
fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows (#2713)
* test(#2711): derive the omit-rule guarded set from the corpus instead of a hand list The GUARDED array was a Goodhart metric: it reported green across 15 non-compliant workflows for no better reason than that nobody had added them to it. The guard now derives its set — every workflow emitting a model="{…}" dispatch site must state the omit-on-inherit/empty rule — and asserts the derivation is non-empty so a broken scan fails rather than passes. Rule detection stays a PROPERTY check, not a template match: plan-phase.md and execute-phase.md state it in different words and both are correct. RED expected on 15 workflows: audit-milestone, code-review, code-review-fix, debug, discuss-phase-assumptions, docs-update, map-codebase, new-milestone, new-project, quick, secure-phase, ui-phase, ui-review, validate-phase, verify-work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows 15 of the 19 model=-dispatching workflows carried no omit-on-inherit/empty guidance — 43 unguarded dispatch sites. Each would emit model="" whenever the bound *_model resolved empty, which is the DEFAULT state on non-Claude runtimes: the installer writes resolve_model_ids:"omit" into ~/.gsd/defaults.json for every one of them (references/model-profiles.md:101), and resolveModelInternal returns "" for that case (src/model-resolver.cts:383-386) and "inherit" for opus-tier agents and the inherit profile (:395). Both 404 on runtimes without native tier aliases — the failure #2517 documented and fixed in one file. Each file now carries a `<!-- #2517 model-omit-on-inherit -->` blockquote naming its own bound placeholders and linking the canonical statement in references/model-profile-resolution.md, mirroring the `<!-- #2508 runtime-aware-dispatch -->` block already present in all 15. The rule text lives in the reference; the workflows carry a pointer plus the one-line instruction, so the next revision edits one file rather than fifteen. plan-phase.md and execute-phase.md are deliberately untouched — they already state the rule in their own wording, and the guard checks the property rather than a template string. No dispatch site is edited and no placeholder renamed: the #2684 binding guard reports the same 19 files / 60 placeholders / 0 findings before and after, which is the independence proof that this change is additive prose only. There is no Hyrum's-Law routing change to disclose. Placement is span-aware. An initial pass anchored to the #2508 marker, but in six files that marker sits INSIDE the Agent(prompt="…") string, so the new paragraph's literal model= landed in a dispatch call span and tripped the #2284 fail-closed Hermes projection guard (bin/install.js:3704), refusing the install. Blocks are now anchored before the opening Agent( of the span owning the first dispatch, and verified to fall inside no span. gen:golden exits 0 across all 19 runtimes. Fixes #2711 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): cite the issue number in the changeset body and tidy block placement Review findings from the two orthogonal passes: - The changeset body ended (#0). Repo convention across every prior fragment (e.g. #2617/#2693, #2608, #2605) is that the trailing (#NNN) is the ISSUE number, known at authoring time; only the frontmatter pr: field carries the 0 placeholder pending backfill. (#0) would have rendered a dead link in the published release notes. - new-milestone.md glued the inserted block directly under the preceding paragraph with no blank line, inconsistent with the other 14 insertions. - The derived-guard non-vacuity floor was >=17 against an actual derived count of 19, tolerating a silent two-file regression. Tightened to >=19. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2711): reword the omit block so it survives Hermes projection, and exempt quick.md by size The first block wording regressed two suites on the full matrix (4 failures on both linux-node22 and linux-node24). gen:golden passing was not sufficient evidence — it exercises the installer's own fail-closed guard, which is narrower than the dedicated tests. 1. tests/fix-2284-hermes-agent-delegate-task-projection.test.cjs asserts that the INSTALLED code-review-fix.md contains no `model=` anywhere outside a string literal — masked whole-file, not merely inside call spans. The block's backticked `model=` survived the mask. The assertion is right: on Hermes the projection strips the parameter because delegate_task has no per-call model at all, so instructing the orchestrator to "omit the model= parameter" is advice about a parameter that does not exist there. The block now says "the `model` parameter" and carries no bare `model=` token. 2. tests/prompt-injection-scan.security.test.cjs flagged quick.md at 50,164 normalized chars against a 50,000 prompt-stuffing threshold. quick.md sits just under the line on next, so any insertion trips it — the situation review.md is already documented for in SIZE_ONLY_WORKFLOWS ("sat at 49,971 chars — 29 below the threshold — so it was going to trip on whatever was added to it next"). quick.md joins it with the same justification. This is a size-finding exemption only: the file is still fully injection scanned, and every other security check still runs on it. Because the canonical block can no longer carry a literal `model=`, the guard's detector now accepts the `<!-- #2517 model-omit-on-inherit -->` marker as the canonical signal, falling back to the inline-prose property for the four files that predate it (plan-phase, execute-phase, scan, ship — all four match the legacy branch). That is strictly stronger than word-proximity matching, and it keeps the guard a property check rather than a template match. Verified: derived guard 19/19 with 0 missing; the #2684 binding guard unchanged at 19 files / 60 placeholders / 0 findings; no inserted block contains a bare model= token; the masked-projection assertion passes for code-review-fix.md; gen:golden exits 0 across all 19 runtimes; lint:ci exits 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2711): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6fb832a93d |
fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve (#2710)
* test(#2684): failing-first guard for unbound model= dispatch placeholders Extends the #2517 omit-on-inherit bucket with a behavioral binding guard: every model="{X}" in a workflow must name a field that workflow actually binds — an init-payload key (queried for real), a shell assignment, or a declared parse field. scan.md ({resolved_model}) and ship.md ({balanced_model}) substitute names nothing emits, so the orchestrator invents the value (ADR-1411). Also pins the shipped reference that seeded the placeholder and instructs the #2517-forbidden model="inherit". RED expected on scan.md, ship.md, and references/model-profile-resolution.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve scan.md:85 passed model="{resolved_model}" and ship.md:486 passed model="{balanced_model}" — names neither workflow's init payload emits. init.map-codebase emits mapper_model; init.phase-op emits no model field at all. With no source for the substitution the orchestrator invents a value, so model_overrides/model_policy are silently inert at both sites — the invisible partial application ADR-1411 prohibits. - scan.md: declare the init fields in a Parse JSON line and substitute {mapper_model}, matching its sibling map-codebase.md. - ship.md: ref.agent is only known at runtime, so resolve it per hook via query resolve-model and dispatch with {HOOK_AGENT_MODEL}, following the same NAME=$(...) → model="{NAME}" convention every other shell-resolved dispatch in the corpus uses (code-review.md, secure-phase.md, ui-phase.md). - Both sites now carry the #2517 rule: omit model= entirely when the resolved value is "inherit" or empty. A bare rename would have traded a dangling placeholder for model="", which 404s on non-Claude runtimes — an agent type absent from the profile table resolves to the empty string, which is exactly ship.md's case. - references/model-profile-resolution.md was the seam: it shipped the copy-pasteable {resolved_model} snippet scan.md inherited, used the stale Task( spelling, and instructed passing model="inherit" outright. Rewritten to teach the real binding convention and the omit rule. Two defects surfaced while fixing this and fixed inline rather than deferred: 1. ref.agent originates in a capability manifest, which may be third-party. Resolving it at runtime made this the first place that value reaches a shell command, so ship.md now validates its shape before interpolating and skips the hook when it fails — matching code-review.md's existing defense-in-depth pattern. Covered by a test that runs the shipped regex against real agent names and injection payloads. 2. The reference doc's omit example first placed a literal model= inside an Agent(...) comment, which tripped the #2284 fail-closed Hermes projection guard (bin/install.js:3704) and refused the install outright. Moved out of the call span; gen:golden is green across all 19 runtimes. Guard tests extend the existing #2517 bucket: every model="{X}" in a workflow must name a field that workflow binds — an init-payload key queried for real, a shell assignment, or a declared parse field. Behavior change (Hyrum's Law): the scan mapper now runs on the catalog-resolved model rather than the session model. The omit path is unchanged. Disclosed in the changeset body. Fixes #2684 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2684): validate capability-supplied ref.agent in-context, not in the shell The first cut of the ref.agent guard was placed after the injection point it was meant to close. It instructed the orchestrator to substitute the untrusted manifest value into a shell assignment and test it there: HOOK_AGENT="<the ref.agent value>" if [[ "$HOOK_AGENT" =~ ^[A-Za-z0-9][A-Za-z0-9._-]*$ ]]; then … Substitution happens before bash parses anything, so a manifest supplying `x"; touch /tmp/pwned; echo "` yields three statements and runs the middle one unconditionally — the regex fires afterwards and protects nothing. ship.md now requires the check to run in-context, the same way the workflow already reads activeHooks ("do NOT pipe it through a shell parser"), and to skip the hook outright on a mismatch. Only a value that has already matched ^[A-Za-z0-9][A-Za-z0-9._-]*$ ever reaches a command line. The guard test asserts the ordering — the in-context requirement and the absence of any raw shell assignment — not just that the pattern rejects metacharacters, since a pattern alone was exactly what gave false assurance here. Also corrects the empty-string explanation in references/model-profile- resolution.md. model_profile:"inherit" resolves to the literal "inherit", not "" (model-resolver.cts:395), and an unknown agent takes the empty-string path via resolve_model_ids:"omit" rather than by absence alone — the doc claimed all three produced "". The workflow instructions were already correct (omit on "inherit" OR empty); only the rationale in the citable reference was wrong. Both found by the isolated adversarial review pass for #2684. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2684): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a7f521c84e |
docs(#2699): record the generalized subagent watchdog as out-of-scope (#2706)
Records the No-go disposition for #2699 in the rejected-capability knowledge base consulted by /triage-review on every future run. GSD does not take a generalized, event-driven, heartbeat-augmented watchdog layer. The artifact-aware spot-check it generalizes already ships for the executor, and the planner gap that motivated the proposal has a scoped fix already diagnosed on #2650 that explicitly excludes this rearchitecture. The entry carries a "What this does NOT cover" section because its keyword surface (watchdog, stall, timeout, heartbeat, orphan, recovery) overlaps request types this decision deliberately does not deny -- notably extending the existing spot-check pattern to another spawn site, which remains the sanctioned incremental path. Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
27c2279a39 |
fix(#2617): project verification next_command onto the runtime's command surface (#2700)
* fix(#2617): project verification next_command onto the runtime's command surface `src/verification.cts` stored and synthesized hard-coded `/gsd:…` command strings with no runtime context, and `phase complete` relayed that raw field straight into its verification-blocked error. On a Codex project the suggested next step was `/gsd:execute-phase`, a surface Codex does not install — it installs `$gsd-execute-phase`. The colon form is wrong twice over: `runtime-slash.cts` documents that "the colon form is never emitted", so EVERY runtime — not just Codex — was being handed a deprecated shape. Fixed at the one routing seam rather than per caller: - The routing table now stores BARE command names (`execute-phase`), never a prefixed literal. A prefixed literal in the table is what leaked. - A single `projectNextCommand(bare, runtime, tail)` helper runs every return path through `formatGsdSlash`, preserving the argument tail (`01 --gaps`) untouched. An empty command stays empty, so "no next step" never becomes a bare prefix. - `readVerificationStatus` accepts `opts.runtime`; `cmdVerificationStatus` and `phase complete` pass `resolveRuntime(cwd)`. The default is `claude`, which yields the canonical `/gsd-` hyphen form. All four routed states are covered: missing, unknown, gaps_found, stale. `init.cts` keeps its own projector deliberately. It already formats correctly, and its command CONTENT differs from the router's on purpose (it appends the phase number to `execute-phase`, and routes `human_needed` to `verify-work`). Consolidating them would silently change `init`'s user-visible output, which this issue did not ask for — so the divergence is left intact and the new tests instead pin the property that matters on both surfaces: no raw colon form escapes. Failing-first record: `origin/next:src/verification.cts` carried the four `/gsd:` literals (lines 101, 108, 382, 392), and 11 existing assertions in tests/verification-status.test.cjs asserted the colon form. Those 11 are corrected in this commit — they passed before the fix and fail after it, which is precisely the regression this closes. Tests are folded into the module's primary suite rather than added as a third file (`lint-test-file-count` caps the `verification` module at two, and consolidating is its documented remedy — growing the allowlist is not). The `phase complete` assertion reads `res.error`, not `res.stderr`: `runGsdTools` exposes a clean non-zero exit's stderr as `error`, and reading the wrong field yields '' and makes the whole check vacuous — which is how this user-visible path stayed untested. Closes #2617 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): scope the new hooks to their describes; cover gaps_found through the CLI Two findings from the orthogonal review of the first commit, both in the tests this change added. 1. The folded block's `beforeEach`/`afterEach` were declared at MODULE scope. node:test applies module-scope hooks to every test in the file, so hooks added for the #2617 suites also wrapped the ~40 pre-existing tests in verification-status.test.cjs — making an unrelated block a single point of failure for them (currently benign, but a throwing hook would have failed suites it has nothing to do with). They now install inside their own describes via a small `useProjectionPhaseDir()` helper, with a comment recording why. 2. The live-CLI `phase complete` test exercised only the `missing` state, so a regression in any other routed branch would have shown up in the router's return object but not in the text a user actually reads. Added a `gaps_found` case per runtime, asserting the projected `plan-phase <N> --gaps` reaches the blocked-completion error. Whole file verified green: 48 tests, 48 pass — the ~40 pre-existing ones included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): correct the last colon-form assertion in phase.test.cjs The remote run surfaced one more stale assertion outside tests/verification-status.test.cjs: the `phase complete` canonical-gate suite matched the blocked-completion message against `/\/gsd:verify-work 0?1/`. That project fixture configures no runtime, so it takes the `claude` default, which now yields the canonical `/gsd-verify-work 01` hyphen form. The colon form this asserted is exactly the deprecated shape #2617 removes — `runtime-slash.cts` documents that "the colon form is never emitted". Like the eleven corrected in the first commit, this assertion passed before the fix and fails after it, which is the regression record rather than a test being loosened: the surrounding assertions (failure reason, `stale` wording, and that neither ROADMAP.md nor STATE.md was mutated) are untouched. Verified against the real CLI: the emitted message is now "Phase 1 verification is incomplete: Verification is stale. Re-run verify-work before transition. Next: /gsd-verify-work 01". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2617): collapse the two verification projectors into one seam The orthogonal review found that `init.cts` carried a second, independently maintained `verificationNextCommand()` that had drifted from the router's table in CONTENT, not just formatting: state router (before) init.cts missing execute-phase execute-phase <N> unknown execute-phase execute-phase <N> human_needed "" (no command) verify-work <N> The `human_needed` row is the sharp one: two GSD surfaces disagreed about whether a next command existed at all, and the router's own next_action told the user to "re-run the verify step until status is passed" while naming no command to run. init's answers were the useful ones, so the router adopts them and init now delegates to it — satisfying the issue's "keep one verification-routing seam" direction. `verificationNextCommand()` is deleted. Appending the phase number surfaced a trap the old bare commands hid. `extractPhaseToken` also returns project-code forms (`PROJ-07`), which are indistinguishable by shape from an ordinary directory name — `gsd-651-parent` yields `gsd-651` — so deriving the argument blindly emits `execute-phase gsd-651`. The number is therefore appended only when it is unambiguously numeric, or when the caller supplies it explicitly. `init` does supply it: its `phaseDir` is unresolved in several branches, where the router could not derive one at all. dir `01-example` -> $gsd-execute-phase 01, $gsd-verify-work 01 dir `gsd-651-parent` -> $gsd-execute-phase, $gsd-verify-work Suites verified green against the built lib: verification-status 50/50, phase 268/268, init 143/143, init-manager 40/40. `npm run lint:ci` clean. User-visible change beyond the reported bug, as agreed: `query verification.status` and `phase complete` now append the phase number for missing/unknown, and emit `verify-work <N>` for human_needed where they previously emitted nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2617): backfill changeset PR number (#2700) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1008aabd31 |
fix(#2615): document the effortSurface axis in the host-integration matrix (#2698)
* fix(#2615): document the effortSurface axis in the host-integration matrix #2481 added `effortSurface` as the ninth negotiated `hostIntegration` axis and wrote documentation-sourced values into 18 descriptors, but never touched `docs/reference/host-integration-capability-matrix.md`. The matrix that ADR-1239 designates the cited source of truth had zero occurrences of the axis: no entry in the axes legend, and no row in any of the per-runtime tables. `src/host-integration.cts` states "every value is documented or explicitly 'undocumented'" — for this axis that was false for every runtime. Adds the legend entry (the `argv` / `none` / `undocumented` vocabulary, plus why there is deliberately no config-file member) and an `effortSurface` row to all 19 per-runtime tables. Every citation is carried over from #2481's own commit message, where the values were sourced: - claude argv -- `claude --help` documents `--effort <level>` - opencode argv -- `opencode run --help` documents `--variant` - codex argv -- `model_reasoning_effort` is a config.toml key, not a dedicated flag, so the generic `-c key=value` override is the only argv route (still argv) - 15 hosts undocumented -- their docs state no reasoning setting kimi-code is the nineteenth section (added by #2603 after #2481) and is the one runtime with no declared value. Its row and a Documentation-gaps entry record why rather than inventing one: Kimi Code documents `/effort` (alias `/thinking`), but only as an INTERACTIVE slash command — `-m, --model` is the only model-adjacent argv. Neither vocabulary member is accurate (`none` would deny a mechanism the host has, `argv` would claim one it does not expose), so closing that gap needs a vocabulary decision, which is a negotiation change and not a documentation one. The absent value already degrades closed exactly as the sentinel does. The regression test derives its runtime list from the registry rather than hardcoding it, so a runtime added later fails until its matrix row exists — the ratchet whose absence let #2481 add an axis with nothing catching the missing docs. Closes #2615 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * docs(#2615): honest citations for the undocumented rows; fix the four stale 8-axis lists Two findings from the orthogonal review of the first commit. 1. The 15 `undocumented` rows shared byte-identical text — "searched the runtime's official docs (see Sources consulted above)" — which is weaker than this file's own convention ("no authoritative doc — searched: <url>") and, worse, implies a per-host targeted search that did not happen: each section's Sources-consulted list was gathered for OTHER axes and contains no CLI-reference or reasoning-effort source. The rows now say plainly what the finding is — an ABSENCE established by #2481's cross-host survey — and cite that survey rather than implying a URL was checked per host. 2. Four normative docs still described "the eight negotiated axes" and omitted effortSurface entirely. The worst of them is docs/how-to/add-or-update-a-host-integration.md — the maintainer's own guide for onboarding a host, whose Step 2 axis table would have a maintainer reproduce exactly the gap #2615 exists to close. Also fixed: docs/reference/host-integration-interface.md (which calls itself the normative reference and had no effortSurface row at all), docs/how-to/author-a-host-plugin.md, docs/registries/README.md ("**exactly** the eight … axes keys"), and CONTEXT.md's matching EoS-registry sentence. Deliberately NOT changed, because they are historical records rather than current contract: docs/whats-new-1.7.0.md and docs/FEATURES.md's 1.7.0 entry (effortSurface shipped in 1.8.0 via #2481 — rewriting them would falsify the release history), ADR-1239's pre-amendment body (already superseded by its own "Amendment (2026-07-21): effortSurface axis (#2481)"), and ADR-1016's "original eight axes", which refers to a different axis set entirely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2615): backfill changeset PR number (#2698) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
28e486faf7 |
fix(#2608): fail closed when git add fails during commit staging (#2693)
* fix(#2608): fail closed when `git add` fails during commit staging `cmdCommit` ignored `git add` failures. #2523 had already stopped a failed path entering the commit pathspec, but skipping it silently left two bad outcomes, both reproduced against the pre-fix build: - SOME paths fail -> `{"committed":true}`. `git commit` still ran and PARTIALLY committed the subset that happened to stage, under a message describing the full requested scope. - EVERY path fails -> `{"reason":"nothing_to_commit"}`, which is not what happened and points the operator nowhere. In both cases git's original `add` stderr was discarded, so the user saw a downstream `commit_failed` / pathspec error naming an innocent file — the symptom reported in the issue from a linked worktree whose git directory was outside the managed writable root. Staging failures are now collected and the command fails closed BEFORE `git commit` runs, returning the issue's specified shape: { committed: false, hash: null, reason: "staging_failed", file: "<first failing path>", error: "<original git add stderr>", failures: [ { file, error, timed_out }, ... ] } A timeout is distinguished as `staging_timeout` (issue AC5) using the projection's SIGTERM+ETIMEDOUT signal — the same idiom worktree-safety.cts uses. The check is placed ahead of the `nothing_to_commit` branch so an all-paths-failed run reports the staging cause rather than an empty changeset. Unchanged: successful staging still commits exactly the declared scope and leaves unrelated staged files alone; an explicitly-named file that does not exist is still skipped rather than staged as a deletion (#2014/#2523), and a request where every named file is missing still reports `nothing_to_commit` — no `git add` ran, so there is no staging failure to report. Regression tests inject the failure by monkeypatching `execGit` on the projection module (per CLAUDE.md, over `chmod 0o000`, which does not fault under root and would make the tests vacuous), driven in a `node -e` child because `output()` writes via `fs.writeSync(1, …)` and cannot be captured in-process. Pre-fix, 6 of the 10 assertions fail; post-fix all pass. Closes #2608 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): roll back the index, guard the sibling surfaces, document the new reasons Six findings from the orthogonal review of the first commit, all fixed here. 1. A `staging_failed` return left the index PARTIALLY STAGED. The paths that did stage stayed in the index with no commit made and no cleanup, so the next bare `git commit` would sweep them up — the same silent partial commit this fix exists to prevent, deferred one step. (Pre-fix the partial state at least got consumed by the incorrect commit.) The staging failure path now resets the paths it staged, matching cmdPrSubrepo's established rollback-then-error convention. The reset is scoped to what THIS call staged — paths the caller had already staged are captured up front and excluded, so a caller's own work is never destroyed — and is best-effort, since an unwritable index (the very failure being reported) cannot be reset either. 2. `cmdCommitToSubrepo` still had the identical defect: a failed `git add` was dropped silently and the function committed the subset that happened to stage, discarding git's stderr. It now fails closed per sub-repo with the same staging_failed/staging_timeout reasons and the same scoped rollback. 3. The `git rm --cached --ignore-unmatch` branch (default mode, for a planning file that no longer exists on disk) still discarded its result. It mutates the index exactly like `git add`, and `--ignore-unmatch` already makes "no such path" a success, so a non-zero exit there is a real I/O failure — now routed through the same staging-failure path. 4. `agents/gsd-executor.md` documented the commit envelope as an exhaustive three-shape enum and pattern-matched only `nothing_to_commit | commit_failed`. It is the sole consumer doc for this surface, so the new reasons are added with explicit guidance not to retry (a retry hits the same unwritable index), and the "one of three shapes" framing is corrected. 5. The default (non---files) staging path and `--amend` are now covered by tests. Both were already guarded by the first commit but unexercised. 6. The changeset framed the fix as `--files`-only; it applies to default and sub-repo commits too, and now mentions the rollback. Regenerated the agent size baseline and the 18 golden install-parity fixtures for the gsd-executor.md edit. 16 assertions across both surfaces verified against the built lib. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2608): update the #2523 out-of-repo contract to the new staging_failed reason The remote test run surfaced this: `#2523: out-of-repo --files path is rejected by git` asserted `reason: 'nothing_to_commit'`, and now gets `staging_failed`. This is a deliberate contract improvement, not a papered-over failure. The old reason existed only because a failed `git add` was skipped and the resulting empty `stagedPaths` fell through to the empty-changeset branch. But "nothing to commit" is not what happened — the caller named a file and git refused it — and that misreport is exactly the class of defect #2608 closes. The result now carries the offending path and git's own message ("… is outside repository at …"), which is strictly more actionable for the same condition. #2523's two substantive invariants are untouched and still asserted: no commit is created, and the index is left clean. Two assertions are ADDED (the path is named, git's message is preserved) so the richer contract is pinned rather than merely allowed. Per CONTRIBUTING, a stale-test correction rides its own commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): compact the executor doc addition to stay under the agent LARGE cap The remote test run failed: `gsd-executor.md is 49217 bytes — exceeds the LARGE hard cap of 49152`. The file was already at 48596 (556 bytes of headroom) and the new commit-envelope documentation pushed it 65 bytes over. The cap is a red line, not a budget to raise, so the addition is compacted rather than the cap moved: four lines instead of eight, keeping the load-bearing facts — the two new reasons, that nothing was committed and the index was rolled back, that `file` + `error` should be surfaced, and that retrying is wrong because a retry hits the same cause. Dropped only the restatement of the linked-worktree example (already in the changeset and PR) and the `failures[]` field (a superset of `file`/`error`, discoverable from the payload). Net addition is now 276 bytes; the file sits at 48872 with 280 bytes of headroom. Extracting the agent's shared boilerplate to references/ would buy much more, but that is a restructuring of the executor agent and does not belong in a commit-staging bugfix. Agent size baseline and the golden install-parity fixtures regenerated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2608): backfill changeset PR number (#2693) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c44241a0b |
fix(#2605): make dropped local-server reviewer lanes loud instead of silent (#2689)
* fix(#2605): make dropped lm_studio / llama_cpp reviewer lanes loud, not silent The `lm_studio` and `llama_cpp` reviewer legs in `gsd-core/workflows/review.md` carried the same empty-output defect as the claude/gemini legs (#2494, fixed in #2592) in a worse variant: when the local endpoint was unreachable or returned empty content, they wrote NOTHING to `{run_dir}/gsd-review-<leg>.md`. There was no `[ ! -s … ]` stub at all, so the file never existed, `write_reviews` omitted that reviewer's section, and the outcome was indistinguishable from the reviewer never having been selected. Two diagnostic holes are closed, because an OpenAI-compatible server fails in two ways that leave evidence in different places: - Transport failure (endpoint unreachable): curl writes to stderr and exits non-zero. Both legs used `curl -s`, which suppresses curl's ERROR text as well as the progress meter, and then discarded stderr to `/dev/null` — so there was nothing to capture even in principle. Now `-sS` with stderr to a `.err` sidecar, matching the claude/gemini/codex legs. - Application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in the response BODY, so stderr is empty and only the body is evidence. The stub appends the raw response. The `llama_cpp` leg additionally piped curl straight into `jq`, throwing the body away before anything could inspect it; the response is now captured to a variable first, as the `lm_studio` leg already did. The existing `>&2` warning is kept and now points at the stub file. Failing-first verified by extracting both shipped blocks and running them under a real bash against a stubbed curl: pre-fix, all three failure modes produce NO review file and no `.err` sidecar for both legs; post-fix, each produces a diagnosable stub, and a successful review still passes through untouched. Closes #2605 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2605): close the remaining silent-drop paths the review surfaced Four defects found by the orthogonal review of the first commit, all fixed here rather than deferred (they are the same defect class the issue is about, and three of them sit inside the code that commit touched). 1. CodeRabbit leg was still unguarded. `coderabbit review --prompt-only 2>/dev/null > file` with no `[ ! -s … ]` stub — the last leg still shaped like pre-#2494 code. A missing or unauthenticated binary left a zero-byte file that write_reviews rendered as "ran cleanly, nothing to report". It now captures stderr to a `.err` sidecar and emits the same diagnosable stub as every other leg. This is the leg the next issue in the #2494 -> #2592 -> #2605 series would have been about. 2. The budget-skip path dropped the lane just as silently. When `prepare_trimmed_prompt_for_reviewer` fails, `*_SKIP=1` bypasses the entire block — guard included — so no file was written and the only trace was a stderr warning nothing persists. All three local-server legs now write a "review skipped: prompt budget too small" stub on that path. 3. Whitespace-only replies evaded the guard. `[ ! -s … ]` counts BYTES, and command substitution strips trailing newlines but not spaces, so a reply of `" "` was written out and passed as a "successful" but vacuous review — the same indistinguishable-from-success outcome the guard exists to prevent. A `case` glob now normalizes whitespace-only content to empty. 4. `echo "$VAR"` swallowed option-like content. bash's builtin `echo` treats a value of exactly `-n`/`-e`/`-E` as a flag and writes 0 bytes, which would trip the empty guard and DISCARD a genuine reply. Switched to `printf '%s\n'`, the idiom the OpenCode leg in this same file already uses for this reason. Also brings the Ollama leg to parity while it is in hand: it always emitted a stub so it never silently vanished, but it was the least diagnosable of the three local-server legs — bare `-s`, stderr to /dev/null, and the response piped straight into jq so the error body was discarded unread. Verified by extracting all four shipped blocks and running them under a real bash against stubbed CLIs: 22 cases (7 per local-server leg x 3, plus CodeRabbit) all produce the contracted output, and a successful review still passes through untouched on every leg. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2605): regenerate golden install-parity fixtures for the review.md edit The remote test run on the prior commit failed with 19 "golden parity — <runtime>" mismatches. review.md is installed into every runtime's tree, so editing it changes its content hash in all 19 golden fixtures. Regenerated with `npm run gen:golden`; the diff is exactly one line per fixture — the gsd-core/workflows/review.md hash — and nothing else. This is the second ripple of a workflow edit, alongside tests/workflow-size-baseline.json which the first commit already updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2605): backfill changeset PR number (#2689) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3eb1cede26 |
fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent Failing-first. Encodes the issue's runtime repro: a trailing comma in .planning/config.json currently yields source:builtin-defaults with degraded:false - byte-identical to the file not existing - and the user's entire configuration is silently discarded. Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather than diagnostic prose, per the ADR-1411 amendment's test-methodology clause and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000 (root bypasses mode bits). Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): distinguish a corrupt config from an absent one loadConfigResolved wrapped the read, the JSON.parse and the entire config build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell through to the same defaults and the branches returned degraded:false - actively asserting health over discarded configuration. A single trailing comma in .planning/config.json silently replaced the user's whole config, reporting source:builtin-defaults degraded:false, byte-identical to having no config file at all. ConfigResolution now carries a machine-readable reason. Genuine absence keeps degraded:false / not_configured; a file that exists but cannot be used sets degraded:true with config_unparseable or config_unreadable. The same split applies to the root config and to ~/.gsd/defaults.json. Control flow is deliberately unchanged. preflight_check reports cyclomatic 141 / cognitive 196 and 93 dependents on this function, with the guidance that small edits beat one big one, so faults are CAPTURED at the existing read sites and stamped onto the returns rather than the try/catch being restructured. Also carries the ADR-1411 amendment's wiring clause: loadConfig returns .config alone to ~51 call sites and would never see the new field, so an unusable file emits a deduplicated stderr diagnostic keyed on resolved path plus errno. Without it the reason would be an unreachable field and the user whose config was discarded would still get no signal - the actual defect. Registers the config-loader seam in lint-resolution-provenance, which until now guarded only agent-skills. Caller audit: ConfigResolution.degraded has exactly one consumer outside this module, cmdAgentSkills (src/init.cts:2259), which destructures {config, source, degraded} - adding a field does not break it. Its --json IR now reports degraded:true for a corrupt config, which is the intended fix and the one observable behavior change. Closes #1880 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): degrade when any config on the path is unusable, not just the last Two defects found by isolated adversarial review of the first cut. BLOCKER: the success-path return did not consult configFault. A corrupt ROOT config whose workstream override happened to parse returned degraded:false / reason:resolved - the root's settings silently dropped, which is the exact failure this issue closes, reappearing for any project using workstreams. The stderr diagnostic fired, so the out-of-band half worked while the in-band half reported a clean resolve; a --json consumer saw health. MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE - so an empty workstream file inheriting a non-empty root reported resolved despite carrying no settings. Emptiness is now judged on the file actually read, snapshotted before normalizeLegacyKeys mutates it. Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880 degraded contract; adds a fast-check property asserting a PRESENT file is never reported not_configured whatever its bytes (CONTRIBUTING.md parser rule); and asserts the literal enum values so the provenance lint's configured_empty/not_configured markers check real assertions rather than incidental prose in test titles. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): reject valid JSON that is not a config object at the read seam The fast-check property added in the previous commit failed on both node lanes: a config.json containing 0, "str", [], null or true is valid JSON, so it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer catch reported not_configured - a PRESENT file reported as absent, which is precisely the collapse this issue exists to close. The property asserts a present file is never not_configured, and it caught it. _readConfigFile now validates shape, not just parseability (ADR-227: check the semantic shape at a trust boundary, not merely the type). A non-object JSON document is an unusable config, reported config_unparseable. Adds named regression cases for each non-object form alongside the property, so the class is documented and not only randomly sampled. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1880): backfill changeset pr number (pr:0 -> 2688) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2452): record a fetch-time shallow failure instead of crashing This guard failed CI on ubuntu-24 while passing on ubuntu-22 and windows-24 for the same commit, and passed on other PRs. Not a flake and not caused by the change under test - a real fragility in the test. runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it assumed the failure mode is always 'fetch succeeds, diff reports no merge base'. At a shallow boundary that lands short of the merge base, git can instead fail during the FETCH ('unable to parse commit' - the boundary commit's parent is not available). Which stage git fails at is version and transport dependent, so on some runners the error escaped runnerDiff and crashed the test rather than being recorded as the ok:false the assertions expect. Both stages mean the same thing for what this guard protects: a shallow base ref cannot resolve the three-dot diff. Also drops two assert.match calls against git's stderr prose. 'no merge base' and 'unable to parse commit' are the same condition reported at different stages, and CONTRIBUTING prohibits raw text matching on subprocess output. The typed outcome (ok === false) is the contract; the tests now assert that plus the presence of a cause. Found while investigating the red lane on #2688; fixed here per the no-defer rule rather than filed. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c07734216f |
fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19 installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped with no citation and no evidence quote. Sourcing every axis independently against Kimi Code CLI's own docs (the issue's explicit requirement — `kimi` and `kimi-code` are distinct products) showed three values had been inherited from the Python `kimi` descriptor rather than sourced: - `embeddingMode` imperative -> declarative. Kimi Code plugins are a `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic API (docs/en/customization/plugins.md) — the same shape as `codex`. - `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md). The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not. - `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth bound is published, so the fail-closed sentinel applies over a guessed integer. `dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is registered with the host and `resolveDispatchType` maps every role onto coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance under Documentation gaps instead. Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/ backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499), and the install adapter is not selected by `embeddingMode` (install.js:543 always uses the imperative adapter). The one visible effect is the curated profile pin, which moves programmatic-cli -> declarative-cli. Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit member that has been in the closed vocabulary since kimi-code shipped. Same defect class and countermeasure as #2598: pin the corrected values and require the matrix to agree with the descriptor, because a descriptor/matrix disagreement is how the gap survived. Closes #2603 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed Surfaced by the orthogonal review of this change. `negotiateHostCapabilities` emits a sentinel-specific warning for every dispatch sub-axis carrying the documented `undocumented` value — namedDispatch, nested, background, subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell through to the numeric guard and reported `host dispatch.maxDepth is missing or not a number — treating as 0`. That message is indistinguishable from a genuinely malformed descriptor, so a correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this PR's kimi-code correction adds a seventh, which is why it is fixed here rather than left in place. The numeric guard keeps firing for genuinely malformed values; both paths still degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests, including the boundary case that the sentinel carve-out must not swallow a real malformed value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2603): backfill changeset PR number (#2687) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a3853472de |
fix(#2598): declare OpenCode subagent dispatch synchronous, not background (#2682)
* fix(#2598): declare OpenCode subagent dispatch synchronous, not background capabilities/opencode/capability.json advertised dispatch.background: true and dispatch.backgroundDispatch: true. negotiateHostCapabilities and every degradationFor / shouldFlattenDispatch consumer trusts these per-field values, so declaring a capability the host lacks OVERSTATES it — the opposite of the fail-closed posture the negotiation is built for. The issue's own citations needed checking before acting: the host-integration matrix (ADR-1239's designated deployment source-of-truth) documented `true` with NEWER evidence than the issue cited, and explicitly marked the issue's sst/opencode#5887 reference as a stale snapshot superseded by #2087. git log confirms #2087 deliberately flipped these from false to true, citing OpenCode v1.15.0/v1.17 as "background subagents enabled by default in all modes". Applying the issue as filed would, on that evidence, have REGRESSED a deliberate update. So the claim was verified against current upstream rather than either document. `packages/opencode/src/effect/runtime-flags.ts` on `dev` today reads: experimentalBackgroundSubagents: enabledByExperimental("OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS") `enabledByExperimental` falls back to the `experimental` flag and `bool()` defaults to false — the parameter is hidden from the model unless an operator opts in by env var. Upstream #29638 is still OPEN and confirms the session loop `tasks.pop()`s one subtask at a time. #2087's "default-on in all modes" reading does not hold against current dev. The issue's CONCLUSION is therefore right even though part of its evidence was superseded: concurrent dispatch cannot be relied on, so both fields are false. The matrix rows are corrected with the verified citation rather than reverted to the old #5887 quote, so the record shows why the value is false TODAY rather than re-asserting evidence that was legitimately superseded. Neighbouring sub-fields are untouched and pinned by test: namedDispatch, subagentToolkit, and isolation:'orchestrator-worktree' (which works via `opencode run --dir` at the OS process level and is unaffected — #2584 does not depend on this value either way). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2598): re-pin the dispatch contract tests to synchronous OpenCode dispatch gsd-test on the descriptor change came back FAILED (5 unique, both node versions). The failures were not incidental — they were deliberate contract-pin tests encoding #2087's decision, one named literally "background UPGRADE": tests/host-integration-descriptors.test.cjs - EXPECTED_FLATTEN[opencode] === false (background-eligible) - the derived background-eligible set pin tests/opencode-imperative-reference.test.cjs - "descriptor declares background dispatch true/true (v1.15/v1.17 upgrade)" - "background UPGRADE changes shouldFlattenDispatch: false now" So this is a recorded decision being reversed, not drift being corrected, and it is reversed on evidence: current upstream `dev` gates the capability behind OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS (default false) and upstream #29638 (OPEN) confirms the session loop still handles one subtask at a time. The issue is filed by the maintainer and explicitly directs "update golden-parity / validator fixtures as needed", which sanctions re-pinning. Behavioral consequence, verified: shouldFlattenDispatch(opencode) now returns TRUE, so GSD serializes opencode dispatch instead of trusting concurrency it cannot get. That is the correct fail-closed direction and is safe today — no shipped GSD flow drives OpenCode background waves (per the issue), and isolation:'orchestrator-worktree' is unaffected because it works at the OS process level via `opencode run --dir`, not via the native subagent. Each re-pinned test now asserts the retracted contract in the opposite direction — feeding the #2087 axes back in must still yield "would not flatten" — so a silent re-flip of either field is caught rather than merely un-asserted. lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2598): backfill changeset pr number (#2682) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5633bb32f |
enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures * feat(#2671): brand raw vs calibrated token types * refactor(#2671): hoist type-compile into a before() hook Two review responses: - The fixture compile ran in the describe() body, so it executed at collection time even when the block was filtered out, and a failed precondition collapsed eight independent assertions into one opaque describe-level failure. A before() hook is this repo's documented idiom and preserves per-test granularity. - parseTokensFlag now records WHY it returns an unbranded number: it validates the magnitude of --tokens, but the basis is decided by --calibrated, so branding here would be wrong for half its callers. The assertion belongs to cmdEstimateCheck, its only caller. * test(#2671): pin each brand diagnostic to its OFFENDING marker Adversarial review demonstrated that asserting only exactly-one-diagnostic- at-code-N is not airtight. Repairing a fixture's brand violation while injecting an unrelated error of the same code (a string passed as the budget argument) still yielded exactly one TS2345, so the fixture would have reported green while no longer testing its regression at all. Each bad-* fixture now routes its violating value through a const named OFFENDING, and the test asserts the diagnostic's start offset falls inside that node — located through the AST, so it survives reformatting and never pattern-matches source text. Replaying the proof-of-concept against the new assertion rejects it: the diagnostic lands on the budget literal, not the marker. Also corrects a doc comment that claimed the program type-checks all of src/; it covers phase-estimation.cts and its transitive dependencies. * chore(#2671): backfill changeset PR number (#2676) |
||
|
|
0d08c32048 |
fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable Every emitted script was rejected. Four invalid constructs, the first fatal on its own, so the Workflow backend could never dispatch a wave: 1. no `export const meta = {…}` first statement -> whole script rejected 2. resumeFromRunId("<id>") -> "resumeFromRunId is not defined". It is a Workflow TOOL INPUT parameter, not a script function. The run id still reaches the caller via summary.resumeRunId, to pass as that input. 3. budget(<n>) -> "budget is not a function". `budget` is a read-only object { total, spent(), remaining() } fed by the caller's token directive; a script cannot set it. Recorded as intent in a comment. 4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions". Now parallel([() => agent(…), …]) — passing agent() results directly also started every agent eagerly, before parallel() could bound concurrency. The single-plan stage had its own branch with the same parallel() defect; both branches are now one array-emitting path. Waves also emit phase() calls whose titles match meta.phases exactly, so progress groups correctly. Two secondary defects kept the script from ever being REACHED — which is why this shipped undetected: 5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no scriptable way" to introspect it and told callers to omit the flag, so gate 5 returned agent_sdk_version_unknown on every automated run while `capability state` still reported active:true. True for bash, false for Node: the router now reads the installed @anthropic-ai/claude-agent-sdk version, walking node_modules up the tree and reading package.json directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the SDK's exports map does not expose ./package.json. Precedence: explicit flag > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an unresolvable version still declines to inline. A too-old SDK now reports the truthful agent_sdk_version_below_floor instead of unknown. 6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any invocation without --runtime reported runtime_not_claude on an ordinary Claude project. Now delegates to runtime-slash.resolveRuntime. The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` snippet is removed rather than repaired: it was also shell-dependent — zsh does not word-split unquoted parameter expansions, so it collapsed to a single argv element, argValue() never matched, and the run failed into the same agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto- resolution removes the need for the construct entirely. Verified with the issue's own repro: no flags now reaches the version gate; an SDK above the floor yields backend:"workflow" with a script that parses as a real ES module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids Findings from the isolated review, all fixed. HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was already RED because of it. The registry embeds the fragment text INLINE, so the shipped/installed copy still taught the exact broken contract this PR fixes: the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown" guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of checking its exit code, so I recorded a red chain as green — checking $? now.) HIGH — three existing tests asserted the OLD broken shape and would have failed CI; none was touched by the first commit: tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…") tests/claude-orchestration.test.cjs — .includes('budget(') tests/claude-orchestration-command-router.test.cjs — .includes('budget(') Each now asserts the corrected contract: the id/pool reaches the caller via summary, and neither construct is ever CALLED. Two sibling assertions had also gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the new explanatory COMMENT contains that substring, not because anything is wired. Rewritten to assert the real property. MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked within a wave, but nothing checked wave ids across waves. That was harmless before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a matching meta.phases entry and the tool matches titles by exact string — two waves sharing an id would collapse into one progress group and misattribute the second wave's agents to the first. Rejected at validation, with tests either side of the boundary. MEDIUM — the fragment contradicted itself (its "Manifest construction" header still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never told the orchestrator to pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to a tool-invocation input, an implementer following only the fragment would have silently regressed phase-resume to a no-op. Both fixed. MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and docs/explanation/claude-orchestration-capability.md documented `resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output — teaching the bug as the feature. Updated to the real contract, including the required meta block and the thunk-array parallel() form. (The changeset is `Fixed`, so the docs gate exempts this; it is corrected because it is wrong, not because a gate demanded it.) LOW — the router's top-of-file comment still described the divergent `--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the fix; and inserting resolveInstalledAgentSdkVersion had orphaned resolveDetectionArgs' JSDoc above the wrong function. Both repaired. lint:ci now exits 0 (verified by exit code, not by reading output). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2590): backfill changeset pr number (#2681) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c3958018dd |
docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern ADR-1411 reasons only about a resolution miss. It is silent on input that is present but not usable, which is how five engine read paths (#1879) could fold an unusable input into the value meaning 'genuinely absent' without contradicting an Accepted ADR. Read together, ADR-1411 and ADR-227 converge and do not license throwing as the cluster's answer: ADR-227 requires malformed input to be coerced rather than propagated and carves out only genuinely-fatal fields, while ADR-1411 already permits a fallback provided it is 'a visible value, not a silent substitution'. The defect in these five sites is therefore not that they fall back but that they fall back invisibly. Records the pattern that follows: every current return value is preserved, and the cause is made visible in-band where the result already carries a provenance envelope, or out-of-band via a deduplicated stderr diagnostic where it returns a bare value it cannot extend. Throwing stays confined to ADR-227's genuinely-fatal carve-out, decided per call. Also names the per-applier caller audit and the lint-resolution-provenance registry gap. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2674): prove the warning-state reset misses the unknown-key dedup set The two existing cases in this suite only pass because each picks a key name no other case reuses, so neither can observe whether the reset the beforeEach calls actually runs. Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after _resetRuntimeWarningCacheForTests(). It is not - the helper clears only _warnedConfigKeys despite documenting itself as resetting per-process warning state. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2674): reset the unknown-key dedup set with the runtime warning cache _resetRuntimeWarningCacheForTests documents itself as resetting per-process warning state but cleared only _warnedConfigKeys, leaving _warnedUnknownConfigKeys populated across cases. The suite that exists to test that set - 'loadConfig - unknown-key warning dedup' - calls the helper in beforeEach expecting exactly this, so the reset was a silent no-op for it; both cases passed only because each picked a key name the other never reused. Any later case reusing a key would have had its warning suppressed by leaked state. Found while amending ADR-1411, which names this dedup guard as the pattern five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the fix would have propagated the footgun to each of them. Folded in here per CLAUDE.md's no-defer rule rather than filed. RED verified on 3c4895841 (test only, no fix): linux-node22 reported 'FAIL tests/config-loader.test.cjs - the documented per-process warning-state reset must clear the unknown-key dedup set too'. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): document src/ in the changeset-lint trigger list CONTRIBUTING.md presented the Changeset Required trigger list as bin/, gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is the most-edited user-facing path in the repo and the omission sends any contributor who touches it into a CI failure the doc says cannot happen. Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so running it bare locally reports success without evaluating the branch. This PR hit exactly that: a local run said ok_fragment_present and CI failed fail_missing_fragment on the same diff. Found while opening this PR; folded in per the no-defer rule. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): restore the round-2 review corrections to the amendment These edits were made in response to the second isolated review pass but never staged: later commits used targeted `git add <file>` for the test and the source fix, so the two markdown files stayed dirty and shipped nothing. The branch carried the round-1 text, including the ADR-227 misquote the reviewer raised as a blocker. Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in was never implemented, so citing it as the precedent was wrong), the dedup key, #1882 folded into the out-of-band mechanism instead of a fourth mechanism-less category, the narrowed caller-audit rationale, and the test-methodology clause. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9ade602219 |
docs(#2585): single lockfile-driven bootstrap path in CONTRIBUTING.md (#2675)
Getting Started opened with `npm install`, contradicting both the `npm ci` quick start fifteen lines below it in the same file and docs/contributing/bootstrap.md, which states `npm ci` is required so installs are reproducible, lockfile-driven, and fail fast when package-lock.json is out of sync. A contributor taking the shortest path — the first copyable block — got the wrong bootstrap contract. Rather than correct `npm install` in place and leave two near-identical blocks, the duplicate "Bootstrap your environment" quick start is folded into Getting Started so CONTRIBUTING.md carries exactly ONE fresh-checkout path, in the canonical order from bootstrap.md (nvm use -> npm run check:env -> npm ci), and names bootstrap.md as the source of truth for everything else (fnm/asdf/mise, the environment validator, daily commands, troubleshooting). That satisfies the issue's "keep bootstrap.md as the source of truth rather than introducing another variant" — three competing variants become one pointer plus one sequence. No regression test: the change is prose in a contributor guide, not a runtime contract, so a readFileSync+includes assertion would be exactly the source-grep test scripts/lint-no-source-grep.cjs rejects. Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
920a5f3f06 |
fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep (#2673)
* fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep The reviewer/workflow config lookups resolved scalars and object fields with a `gsd_run query <cmd> … | jq … 2>/dev/null || <default>` shape. On any machine without jq (the default on Windows/Git-Bash) the jq stage fails with exit 127, the failure is swallowed by 2>/dev/null + the trailing || default, and the variable comes back EMPTY — the configured per-lane model/host/budget is silently dropped and the lane falls back to CLI defaults with no diagnostic. gsd-tools ships native flags that do the same job with no external dep: config-get <key> --raw (strips JSON quotes off a scalar) resolve-model <id> --pick model (descends an object) resolve-execution … --pick <f> (same) verification.status … --pick status Replaced every jq-piped config/model/verify lookup across review.md (×23), plan-phase.md, ship.md, debug.md (incl. the redundant boolean coercion — --raw returns true/false as bare tokens natively), autonomous.md (×2), ai-integration-phase.md (×4), and eval-review.md. The legitimate structured-JSON jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile) are untouched — only the jq-replaceable lookups moved to the native flags. Adds tests/fix-2589-config-get-no-jq.test.cjs: a source-invariant guard asserting no audited workflow pipes config-get/resolve-model/resolve-execution/verification.status to jq (fails-first on the pre-fix text, passes after). * test(#2589): update autonomous-converge jq assertion to --pick; regen golden fixtures Two test consequences of the workflow-doc edits in the prior commit: 1. tests/autonomous-converge.test.cjs pinned the OLD jq-dependent shape (`verification.status … | jq -r '.status//empty'`) as the canonical routing contract. The test's INTENT is correct (route human validation through canonical verification.status) but it over-specified the MECHANISM (the jq pipe). Updated the assertion to match the new native --pick status shape; the contract being guarded (canonical verification.status read before the human_needed branch) is unchanged. 2. The golden-install-parity fixtures (19 runtimes) record a content hash of every installed workflow .md; the 7 edited workflows changed those hashes. Regenerated via `npm run gen:golden` (the test's own failure message instructs this). Only the 7 edited workflow hashes changed in each fixture. * fix(#2589): declare jq a prerequisite for the lanes that still need it; repair test file Three defects in the first cut of the #2589 fix: 1. tests/autonomous-converge.test.cjs was a JavaScript syntax error. The regex literal /...2>\/dev/null .../ left the second slash unescaped, terminating the literal early and parsing `null` as regex flags: SyntaxError: Invalid regular expression flags The whole file failed to load, so every assertion in it — including the #1522 and #1526 guards — silently stopped running. Replaced with the string-compare form already used at line 202 for the sibling shell-snippet assertion. 2. lint:ci failed. tests/fix-2589-config-get-no-jq.test.cjs buckets into the capped `config` production module via its `config-get-...` effective prefix, making it a novel offender against the 2-file cap. The test is about workflow documents, not the config module, so it is renamed to fix-2589-workflow-jq-dependency.test.cjs (free prefix) rather than growing the allowlist with a module that does not actually need a 5th test file. 3. The fix deleted the repo's only jq-prerequisite declaration. review.md:244 ("install jq if missing") was the anchor plan-review-convergence.md cites by line number, and it went away with the jq pipes — while the ollama, lm_studio, llama_cpp, opencode, and agy lanes still hard-require jq to parse HTTP /v1/chat/completions responses, opencode's JSONL event stream, and agy's conversation cache. On a jq-less host those five lanes swallow exit 127 into empty output: the same silent-degradation class #2589 exists to close. detect_clis now probes jq alongside the other prerequisites and emits jq:available / jq:missing, and the five dependent lanes are treated as undetected when it is absent, with an install hint. The six lanes that do not need jq stay selectable. plan-review-convergence.md now cites the section by name instead of a line number that moves. Regression guards added to the renamed test file: review.md must keep the jq probe and must name all five dependent lanes, and no workflow may cite review.md by line number. Workflow-size baseline and the 19 install-parity goldens regenerated for the review.md / plan-review-convergence.md edits. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2589): decide the ship verification gate on a single verification.status read Isolated review finding (medium). Pre-fix, ship.md captured verification.status ONCE into $VERIFICATION and picked status / next_action / next_command off that cached JSON with three jq calls. --pick takes a single dot-path field, so the mechanical conversion issued three separate queries up front: three node spawns that each re-read the phase VERIFICATION.md and re-derive the commit-time vs mtime staleness comparison, on every ship — including the common passing path that never uses the two message fields. It also meant the gate's verdict and the message shown to the user were derived from three reads with no guarantee they observed the same state. The gate now reads `status` once and decides. The two message-only fields are read on the blocking path only, after PHASE_VERIFICATION_INCOMPLETE is already determined — so the passing path costs one query instead of three, and a concurrent write between reads can no longer make the gate and its message disagree, because the block/allow decision no longer depends on them. Adding a multi-field --pick to gsd-tools would have collapsed this to one query, but that changes the flag's output contract and belongs in its own change. Regression guard in tests/fix-2589-workflow-jq-dependency.test.cjs: ship.md must read verification.status exactly three times total, the block decision must follow the status read, and next_action / next_command must both appear after the blocking prose so they cannot drift back onto the passing path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * docs(#2589): document the jq prerequisite for the five reviewer lanes that need it /gsd-review's ollama, lm_studio, llama_cpp, opencode, and agy lanes parse JSON GSD does not produce (OpenAI-compatible /v1/chat/completions responses, OpenCode's JSONL event stream, Antigravity's conversation cache), so they require jq on PATH. Nothing in docs/ said so. Records which five lanes need it, which six do not, that reading configured models/hosts/budgets no longer requires jq at all, and what /gsd-review now does when jq is absent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2589): backfill changeset pr number (#2673) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |