cfdfdf0b4c857d6dfe64e7981bbc4f894b5832ec
167 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7f13ee5373 |
enhance(#2793): ADR-2782 — reviewer lane becomes a declared capability surface (#2809)
* docs(#2793): add ADR-2782 — reviewer lane capability surface Design lock for epic #2782. Declares a reviewer lane as capability data rather than a core patch across three unrelated surfaces. Key decisions: - D2 transport discriminator (spawn | openai-http) — a survey of all twelve lanes found three that are HTTP endpoints with no binary, which invalidated the single-invoke-shape draft. - D4 the reviewer body is optional and absent-safe at every layer. - D5 a fourth executable-surface disclosure class covering the lane binary or host AND its egress payload classes. - D6 handler is a closed first-party enum, upholding ADR-1016; the consequence — third-party lanes are data-only — is stated plainly. - D7 probe kinds wider than existence, and every probe bounded. Amends ADR-857, ADR-894, ADR-1016, ADR-1244. Also records the D7/D8-extended-by-ADR-1244 marker on ADR-857 that ADR-1244 D8 promised but never added. Closes #2793 * docs(#2793): address orthogonal review findings on ADR-2782 Two blockers from the isolated adversarial pass: - D5 disclosed the spawn binary but not its args, reopening the #1459 bug class already fixed for MCP servers (binary python3 + args -c <program>). args are now disclosed and signature-bound. - hostConfigKey resolves from .planning/config.json, which is mutable after consent with no integrity check, so a lane consented against localhost could be silently redirected to a remote host by an ordinary PR. The resolved host is now consent-bound and re-verified on the invocation path; a mismatch blocks the lane. Majors and spec gaps: - D4 gains an explicit-selection carve-out. Absent-safe governs discovery, never a lane the user named; the current selector records that as info, which Phase 1 now corrects. - D4 gains a warning delivery channel. - D6 enumerates the handler closed-enum members; a closed enum whose membership is left to the implementing phase is not closed. - D6 records aider and plandex as concrete lanes the vocabulary cannot express, rather than claiming sufficiency it did not verify. - D2 gains evidenceClass, requiresBinaries, promptBudgetKey for per-lane divergence that was only prose, and motivates the one-member outputChannel enum. - reviewer.requires renamed requiresBinaries — it collided with the envelope requires (capability deps) at a different nesting depth. - Antigravity two-level timeout: Context cited it then dropped it; now explicitly delegated to the handler. - D9 gains a per-key ownership table, including three keys that stay central because they are policy across lanes, not lane properties. - Phase table maps every decision D1-D9 to a delivering phase; D6 handler modules and D5 invocation-time re-verification were previously unclaimed. - American English per house style. |
||
|
|
1e3c995e6f |
fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it Every input to `diffEmitted` is base-relative -- `baseline` vs `current`, `changedPaths` from `git diff base...HEAD` -- except the ack set, which was read absolutely, from the working tree only. A differential machine consulting a non-differential input. So `staleAcks` asks exactly one question, "did a delta consume you?", and that cannot distinguish an ack that never explained anything (an authoring mistake) from one whose ripple is now absorbed into the base (the ack's SUCCESS condition). After merge an ack is in the second state but reports as the first. The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request events only, so a push to `next` falls through to origin/next -- the very commit under test. Both sides build identical content, no deltas remain, and every live ack is reported stale. PR #2768 acked a deliberate 40866 -> 42020 byte growth, was green on its own lane, and reddened `next` the moment it merged. It also reds every PR branching off the poisoned base, and since publish-emitted-baseline is gated on the test job, it blocked baseline publication too. Give the ack the base side it was missing. `diffEmitted` now takes `baseAck` -- the same document at the base ref, via `readAckFileAtRef`. An entry already present there is SPENT: it may no longer consume a delta and is never reported stale, only surfaced as `spentAcks` for tidying. An entry new or reworded in this diff stays live, and if nothing consumes it that genuinely fails, with blame on the author who just wrote it. This closes a hazard the IMPLEMENTATION named but could not prevent -- a leftover ack silently pre-clearing the next ripple on its path. (ADR-2719 §3 asserted only that TOUCHING the file is the alarm; its residual-risk list never covered pre-clearing, and §3 now carries an amendment.) Verified against the two-PR laundering sequence -- land an innocuous ack, then change the artifact -- which passed silently before and now fails on both the hash pass and the size ratchet. Three things the design has to get right, each of which was wrong first: - A read failure on the base document THROWS; only absence-at-the-ref returns null. Returning null on error LOOKS armed (every entry stays live) but a live entry's defining power is that it CONSUMES a delta, so null is armed on the staleness axis and DISARMED on consumption -- silently the whole pre-#2789 gate. `git show` cannot tell absence from fault, so absence is established with `ls-tree`. - Re-arming a spent ack costs actual PROSE. Internal whitespace and the zero-width family collapse, and `runtime` is not compared: a doubled space, an invisible character, or a decorative field would otherwise re-arm an ack whose justification still describes the previous ripple, showing a reviewer nothing. - `baseAck` is REQUIRED once an ack declares entries -- omission is an error, not a silent "inherit nothing" -- so a dropped argument fails loudly instead of quietly restoring this bug with the suite green. Because a corrupt document ON THE BASE is expensive (the loud base-side failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs blocks one from landing. It is standalone rather than importing parseAck -- scripts/ ships in the npm package and tests/ does not -- so a parity test runs both surfaces over one corpus and fails on divergence; it caught one immediately, a `null` document, now classed as policy rather than schema. Deadlock is separately foreclosed: a tree carrying no ack never reads the base, so the PR that DELETES a corrupt file still lands. `readAckFileAtRef` takes an injected git runner so all four branches are tested deterministically; it never executes in the remote runner, where the real-tree test skips for want of a base ref. It also refuses an option-shaped ref, since execFileSync's array form stops shell metacharacters but not git's own option parsing. Rejected: skipping the differential when base == HEAD. It treats the symptom, costs real coverage on the push-to-next lane, and does nothing about the downstream PRs the same flaw was reddening. Deletes the now-spent tests/emitted-drift-ack.json, and updates the CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a LIVE entry is, and a spent one is inert. Closes #2789 * chore(#2789): backfill changeset PR number |
||
|
|
e276cc7f00 |
enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate check-glossary-refs asserts that every backticked tests/ token in CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719 section 3) is absent on a healthy next BY DESIGN — it appears only inside a PR that needs it, which is what makes touching it the alarm. It passed before only by accident of backtick pairing: CONTEXT.md's RULESET entries are themselves backtick-wrapped and contain backticks, so the token happened to fall outside a code span. Any edit that shifted the parity exposed it. A gate that passes by luck is not passing. The exemption is exact, not a prefix hole: a sibling missing tests/ path still fails, and a test locks that. * feat(#2778): make the size-ratchet failure name its own remedy The growth branch stated a requirement and withheld the means of satisfying it: no ack file named, no schema, no key format, and no do-not-regenerate line — so the likeliest guess was to hunt for a baseline that #2724 deleted. Observed live on #2543. All remediation now comes from one frozen REMEDIATION export whose example document is rendered from ACK_VERSION, so the taught schema cannot drift from the schema parseAck accepts. A round-trip test feeds the printed document back through parseAck. The report is now built as a typed IR (buildReport) that formatReport renders, so tests assert on structure rather than prose, per CONTRIBUTING.md's raw-text-matching rule. Two defects found and fixed inline while building: - diffEmitted's validation early-return omitted newFileCapExceeded while formatReport reads its length, so the branch that reports a failed git diff threw a TypeError instead of naming the problem. - Printing one complete ack document per failing branch made each read as the whole file, so pasting the second over the first silently lost an acknowledgment. One document now covers the whole report. Closes #2778 * chore(#2778): backfill changeset pr number to 2780 |
||
|
|
16e59d0db5 |
fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs
Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @
|
||
|
|
1c1af70a4b |
refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator Removes the 19 committed path->hash manifests, the two per-file size baselines, tests/golden-install-parity.test.cjs, and scripts/gen-golden-install-parity-zcode.cjs. These were pure functions of the source tree (ADR-2719); the differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the sole gate for emitted-artifact propagation. tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs are unchanged (ADR-2719 section 7 exception). Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's existence guard, .gitattributes, package.json scripts, the emitted-provenance totality guard's IO, the differential check's baseline acquisition, CI wiring to publish/restore the baseline artifact, and docs. * refactor(#2724): make the differential attribution check self-sufficient Three fixes required to delete the golden fixtures without breaking CI: - scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs from the three rules that named it. #2759's missingRuleTestFiles guard hard-throws at module load if a rule names a test file absent from disk, which would break the changes job on every PR the moment the fixture-deletion commit landed. - tests/helpers/emitted-provenance.cjs: loadManifests() read the committed golden fixture directory. With that directory deleted at every future ref, this would throw at module load forever, taking the Phase 2 totality guard down with it. Rebuilt from real installer spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest), the same shape emitted-runtime.cjs's currentManifests() already uses. - tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs: the real-tree test's baseline acquisition swaps from baselineManifestsAtRef(base) (git show at a ref that no longer carries fixtures) to resolveBaseline()'s documented precedence: env, then the on-disk cache, then an in-job build. The build fallback (buildBaselineAtRef, new) checks out base into a throwaway git worktree and runs the new scripts/gen-emitted-baseline.cjs there -- no npm ci needed, since bin/install.js and the test helper shells are Node-builtins-only. That script also publishes the baseline artifact from CI's push-to-next job (wired in a follow-up commit). * refactor(#2724): retire the merge-driver bridge and per-file size baselines The Phase 1 bridge (#2721) is retired now that the artifacts it guarded are deleted: scripts/git-merge-regen-driver.cjs, its test, and the 'setup:merge-driver' npm script are removed, and the .gitattributes merge=gsd-regen/linguist-generated block for the three deleted-path globs is dropped. tests/fixtures/install-tree/*.json keeps its normal merge behavior, unchanged (ADR-2719 section 7). scripts/update-size-baseline.cjs and its test are removed: their sole purpose was regenerating tests/workflow-size-baseline.json and tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm script and its step in 'regen:derived' go with it. The per-file baseline describe blocks in tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs are removed for the same reason; the independent loose-tier hard caps are untouched. The differential attribution check's size ratchet (tests/emitted-diff.cjs, already shipped in #2723) is the replacement anti-creep mechanism. 'npm run gen:golden' is replaced by 'npm run gen:install-tree', which keeps regenerating tests/fixtures/install-tree/*.json (the one artifact family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's error messages point at the new command name. tests/golden-parity-single-source.test.cjs's anti-divergence guard (#2266) is retargeted from the two deleted golden-parity consumers to their two replacements (tests/helpers/emitted-runtime.cjs and tests/helpers/emitted-provenance.cjs), which import buildParityManifest the same way — the divergence risk the guard exists for is unchanged. Also wires CI: a new publish-emitted-baseline job runs scripts/gen-emitted-baseline.cjs after a push to next and caches the result keyed on the sha; the test and test-full jobs restore that cache on pull_request events, keyed on the PR's base sha, and export GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree test to pick up. * docs(#2724): flip ADR-2719 to Accepted and update contributor docs Status: Proposed -> Accepted. Regenerated docs/adr/README.md index. CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET. EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET. AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry) no longer point at the deleted golden-install-parity fixtures, size baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver / git-merge-regen-driver.cjs bridge. Editing shipped content now requires zero manual fixture regeneration, documented against the differential attribution check instead of the deleted commands. * docs(#2724): add changeset for removed golden-parity commands * fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs token in CONTEXT.md resolves to a real file. The RULESET. EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks, which the checker reads as a live reference, not historical prose. * test(#2724): retarget ci-test-scope tests off the deleted golden test tests/ci-test-scope.test.cjs asserted specific RULES entries select tests/golden-install-parity.test.cjs, and that every rule selecting it also selects both emitted gates. Both premises broke when the golden test was deleted (#2724): the deleted filename never re-appears in targeted_tests, and there was no longer a third file for the gates to travel alongside. Retargeted the two selection describe blocks to assert tests/emitted-provenance.test.cjs directly (the drift guard the golden gate's rules were retargeted to), and simplified the third block to assert the two emitted gates always travel together, without reference to the golden filename. * docs(#2724): repoint two contributor how-to guides at the differential check Both guides told contributors to regenerate a baseline against tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed at the differential attribution check (tests/emitted-attribution.test.cjs, ADR-2719), which needs no manual regeneration step. * fix(#2724): repair phase6-capstone-conformance's deleted-baseline read An independent orthogonal review caught a real regression this branch introduced into a test file the branch's diff never touched: tests/phase6-capstone-conformance.test.cjs read tests/workflow-size-baseline.json (deleted earlier in this branch) with no fallback, so the whole suite would throw ENOENT the moment this branch landed. The test's actual intent — prove the host-loop workflow files are real, tracked, non-empty docs — is preserved by asserting the live byte count via the same shared counter (scripts/workflow-size.cjs) the size guards already use, instead of a committed snapshot. Also, from the same review: a stale doc comment in scripts/workflow-size.cjs still named the deleted scripts/update-size-baseline.cjs as a consumer, and buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left two fs.rmSync calls unguarded against masking the primary result/error, inconsistent with the try/catch already wrapping the git cleanup beside them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef explaining why it (and its siblings) are kept despite having no production caller post-cutover — they still answer real questions about refs that predate the cutover. * fix(#2724): repair three real regressions found by remote verification 1. tests/emitted-provenance.test.cjs's two hostile-input tests (non-object manifest, unreadable fixture) drove loadManifests(tmp) and monkeypatched fs.readFileSync, both premised on the deleted fixture-directory read this branch already replaced with real installer spawns -- the negative assertions silently stopped firing. loadManifests() now accepts injected {families, install, build, clean} (defaulting to production values), giving the tests a real seam to drive a bad build result and a build failure through the ACTUAL loader instead of a reimplementation, and added coverage that clean() still runs on both paths. 2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE' steps hardcoded shell: bash, which is wrong on windows-latest (native pwsh) and on test-full's macos-latest legs (native zsh per that job's own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning .test.cjs) caught it. Replaced the inline bash script with scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a bare 'node <path>' command line has no shell-specific syntax, so it runs correctly under bash, zsh, and pwsh without a shell override. tests/phase6-capstone-conformance.test.cjs's deleted-baseline read (caught by the same remote run, at a commit prior to this one) was already fixed in d0c3b1242 and is not touched here; verified still passing after these changes. * fix(#2724): revive ADR-1610's new-file size cap inside the differential An isolated review caught a real regression: deleting tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP (ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor) with no successor. tests/helpers/emitted-diff.cjs's size ratchet already 'continue's past any file absent from sizeBaseline -- exactly the files this cap exists to bound -- so a brand-new workflow file sized 32,769-40,960 bytes passed CI clean and shipped, then risked silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted and never referenced anywhere in this branch. Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline, name) signal the growth check already computes -- 'new' is exactly 'present in sizeCurrent, absent from sizeBaseline'. Not ack-able, matching the tier hard caps it sits beside: the fix is extraction, not an acknowledgment entry. Documented, disclosed narrowing: the pure differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering (tests/workflow-size-budget.test.cjs's classification), so a legitimately large new file must extract rather than tier in, one release earlier than an existing file would need to. ADR-1610 itself is left unamended -- this restores its decision rather than re-litigating it. Also fixes a stale comment plus a redundant real 19-installer-spawn assertion left over from the pre-injection-seam version of tests/emitted-provenance.test.cjs's build-failure test, and annotates 3 of 4 stale golden-fixture citations in docs/reference/host-integration-capability-matrix.md as superseded (the 4th is an accurate historical PR narrative, left alone). * fix(#2724): repair three red CI defects on the golden-fixture cutover Windows-only provenance false attribution (defect A): the `hooks-built` provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are Windows-only installer output (ensureCodexHooksJsonSessionStart / ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so the self-attribution resolved to a path that exists on no platform. Only windows-latest ever emits the key, so this only failed there. Fixed by special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated rule) — a dedicated rule would match zero paths, and therefore report as a dead rule, on every non-Windows lane of the same totality guard. `sources` already supported per-match functions; `transforms` is extended to support the same shape so the attribution can vary by match within one rule. Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef` ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but that script is new in this PR and therefore absent at any base ref that predates it — every call failed closed with "Cannot find module". Fixed by running the PR checkout's own generator against the worktree via a new `--dir` parameter, decoupling "which copy of the script runs" from "which tree it measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override, threaded down to `runMinimalInstall`'s new `installScript` override). This is not just a bootstrap fix: a differential needs ONE measurement schema applied to both sides, or the two stop being comparable the moment that schema evolves — running each side's own copy would silently reintroduce that risk. Verified locally end-to-end against real origin/next: resolves a valid {version, sha, manifests, sizes} artifact with the correct sha and no leaked worktree. Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let docs-lint evaluate the fragment for the first time; it already passes (docs/TESTING-SUITES.md and friends already document the removed scripts). Also fixed while in this file: an eslint no-unused-vars warning surfaced by the changed lint run (unused `cleanup` import in tests/emitted-provenance.test.cjs). Added regression coverage for both A and B: a cross-platform spot-check that drives the real hooks-built rule against `.cmd` keys directly (not through a real Windows install), and a real-tree test that drives buildBaselineAtRef against a base ref verified (via git cat-file) to lack the generator, both skipping honestly rather than false-passing when their precondition does not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test Two isolated-review findings on PR #2767: - `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the wrapped `hooks/<name>.js` script, asserting a byte-provenance link that does not exist — traced against buildCodexHookWindowsShimIR (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that same file) flows into the .cmd bytes, never its content. Point `sources` at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention used elsewhere in the table. Since `sources` is checked before `transforms` in the differential, the wrong mapping silently excused any .cmd byte movement caused by editing the wrapped .js file. - The `buildBaselineAtRef` regression test skipped unless a resolvable base ref still lacked scripts/gen-emitted-baseline.cjs — true only until this PR merges, after which every base ref carries the file and the test skips forever with zero ongoing coverage. Rebuilt hermetically: synthesize the missing-generator condition in-place via git plumbing (a throwaway commit, child of HEAD, with just that one file removed from a scratch index), never touching the real working tree, HEAD, or index, and never depending on ambient history or remotes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path The runner container mounts the repo at a path owned by a different uid than the process running the suite, so git's dubious-ownership protection refuses every git operation there. GitHub Actions never hits this because actions/checkout registers the workspace as safe automatically; this runner's container does not. buildBaselineAtRef is the production build-fallback the sole remaining emitted gate depends on (resolveBaseline's in-job-build leg), not just a test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs) that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's worktree add/remove/prune, and the hermetic regression test added in the prior commit — funnels through, plus gen-emitted-baseline.cjs's own rev-parse (now reusing that same wrapper instead of a second execFileSync, so the fix has one source of truth). Each call declares -c safe.directory=<the exact directory it already operates on>, never the * wildcard. Audited every other helper on this surface (emitted-diff.cjs, emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them shell out to git at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
90ba0ef10b |
docs(#2720): submit ADR-2719 — emitted-artifact attribution design contract (#2726)
Replaces the committed golden-install-parity hash manifests and per-file size baselines with a computed conservation law: every emitted path whose hash moves must be attributable, via a declarative provenance table, to a path the PR actually changed. Supersedes ADR-2264 Decision §2-§4 and its Amendment; ADR-2264 Phase 1 (the single-source buildParityManifest and exclusion constants) is retained and depended upon. Satisfies ADR-2264 AC1 rather than rewording it away, per that ADR's own 2026-07-17 audit. Docs-only. Both sides of the supersession edited together; ADR index regenerated with gen-adr-index.cjs --write. Closes #2720 Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a5633bb32f |
enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures * feat(#2671): brand raw vs calibrated token types * refactor(#2671): hoist type-compile into a before() hook Two review responses: - The fixture compile ran in the describe() body, so it executed at collection time even when the block was filtered out, and a failed precondition collapsed eight independent assertions into one opaque describe-level failure. A before() hook is this repo's documented idiom and preserves per-test granularity. - parseTokensFlag now records WHY it returns an unbranded number: it validates the magnitude of --tokens, but the basis is decided by --calibrated, so branding here would be wrong for half its callers. The assertion belongs to cmdEstimateCheck, its only caller. * test(#2671): pin each brand diagnostic to its OFFENDING marker Adversarial review demonstrated that asserting only exactly-one-diagnostic- at-code-N is not airtight. Repairing a fixture's brand violation while injecting an unrelated error of the same code (a string passed as the budget argument) still yielded exactly one TS2345, so the fixture would have reported green while no longer testing its regression at all. Each bad-* fixture now routes its violating value through a const named OFFENDING, and the test asserts the diagnostic's start offset falls inside that node — located through the AST, so it survives reformatting and never pattern-matches source text. Replaying the proof-of-concept against the new assertion rejects it: the diagnostic lands on the budget literal, not the marker. Also corrects a doc comment that claimed the program type-checks all of src/; it covers phase-estimation.cts and its transitive dependencies. * chore(#2671): backfill changeset PR number (#2676) |
||
|
|
c3958018dd |
docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern ADR-1411 reasons only about a resolution miss. It is silent on input that is present but not usable, which is how five engine read paths (#1879) could fold an unusable input into the value meaning 'genuinely absent' without contradicting an Accepted ADR. Read together, ADR-1411 and ADR-227 converge and do not license throwing as the cluster's answer: ADR-227 requires malformed input to be coerced rather than propagated and carves out only genuinely-fatal fields, while ADR-1411 already permits a fallback provided it is 'a visible value, not a silent substitution'. The defect in these five sites is therefore not that they fall back but that they fall back invisibly. Records the pattern that follows: every current return value is preserved, and the cause is made visible in-band where the result already carries a provenance envelope, or out-of-band via a deduplicated stderr diagnostic where it returns a bare value it cannot extend. Throwing stays confined to ADR-227's genuinely-fatal carve-out, decided per call. Also names the per-applier caller audit and the lint-resolution-provenance registry gap. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2674): prove the warning-state reset misses the unknown-key dedup set The two existing cases in this suite only pass because each picks a key name no other case reuses, so neither can observe whether the reset the beforeEach calls actually runs. Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after _resetRuntimeWarningCacheForTests(). It is not - the helper clears only _warnedConfigKeys despite documenting itself as resetting per-process warning state. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2674): reset the unknown-key dedup set with the runtime warning cache _resetRuntimeWarningCacheForTests documents itself as resetting per-process warning state but cleared only _warnedConfigKeys, leaving _warnedUnknownConfigKeys populated across cases. The suite that exists to test that set - 'loadConfig - unknown-key warning dedup' - calls the helper in beforeEach expecting exactly this, so the reset was a silent no-op for it; both cases passed only because each picked a key name the other never reused. Any later case reusing a key would have had its warning suppressed by leaked state. Found while amending ADR-1411, which names this dedup guard as the pattern five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the fix would have propagated the footgun to each of them. Folded in here per CLAUDE.md's no-defer rule rather than filed. RED verified on 3c4895841 (test only, no fix): linux-node22 reported 'FAIL tests/config-loader.test.cjs - the documented per-process warning-state reset must clear the unknown-key dedup set too'. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): document src/ in the changeset-lint trigger list CONTRIBUTING.md presented the Changeset Required trigger list as bin/, gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is the most-edited user-facing path in the repo and the omission sends any contributor who touches it into a CI failure the doc says cannot happen. Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so running it bare locally reports success without evaluating the branch. This PR hit exactly that: a local run said ok_fragment_present and CI failed fail_missing_fragment on the same diff. Found while opening this PR; folded in per the no-defer rule. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): restore the round-2 review corrections to the amendment These edits were made in response to the second isolated review pass but never staged: later commits used targeted `git add <file>` for the test and the source fix, so the two markdown files stayed dirty and shipped nothing. The branch carried the round-1 text, including the ADR-227 misquote the reviewer raised as a blocker. Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in was never implemented, so citing it as the precedent was wrong), the dedup key, #1882 folded into the out-of-band mechanism instead of a fourth mechanism-less category, the narrowed caller-audit rationale, and the test-methodology clause. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
1a9ae7601b |
docs(#2629): adr — phase effort estimation & calibration design lock (#2636)
* docs(#2629): adr — phase effort estimation & calibration design lock * docs(#2629): annotate phase-0 status and link adr cross-reference * docs(#2629): derive estimate confidence from sample count, not self-rating |
||
|
|
0f46fa366f | docs(#2606): ratify ADR-612 getMilestoneFromPhaseId bracket return-form (vN.0) (#2607) | ||
|
|
be3bf97eff |
docs(#2584): ADR-1239 Codex-binding amendment + dispatch.isolation capability (Phase 0) (#2600)
* docs(#2584): add ADR-1239 Codex-binding amendment + dispatch.isolation capability * chore(#2584): backfill changeset PR number (#2600) |
||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
67a9243cf1 |
chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants
The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.
Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:
- scripts/gen-adr-index.cjs generates the index between markers and validates
the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.
Correct the lifecycle metadata the gate surfaced, without flipping any status:
- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
capability system shipped and epic #857 is closed. Ratification is a
maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: capture stderr via spawnSync; record ADR-0010 draft supersession
Two fixes surfaced by the first gsd-test run and by regenerating the index:
- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
through the thrown error on non-zero exit. The `--write` path exits 0 while
reporting outstanding violations on stderr, so the helper always saw ''.
spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
"earlier draft superseded by ADR-0011" while the file itself still said
Proposed. Deriving the index from the files would have dropped that
assertion and resurrected a superseded draft as a live decision, so it is
recorded at its source, with the reciprocal Supersedes on ADR-0011.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174
src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").
ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (
|
||
|
|
cf004df678 |
refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in runCommand's default case after capability/overlay dispatch, before the unknown-command error. Migrates 'state' as the pilot: removes the hardcoded case 'state': arm; state now dispatches default -> dispatchHostCommand -> routeStateCommand, byte-identical to the old path (proven by the new state-command-cutover equivalence test, 5-category template). Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey) — the capability registry stays reserved for toggleable feature bundles per ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked; the ADR is corrected here alongside the code that realizes it. - gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand (prototype-pollution-safe); wired into default case; case 'state': removed; dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests. - tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY equivalence (recording-mock + runGsdTools end-to-end + pollution guard). - docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability- registry model (correction that did not land in the merged #2355). Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/ verify using this proven template. Closes #2360. * test(#2360): regenerate golden fixtures + allowlist for state cutover Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the install-parity fixtures (gsd-tools.cjs content hash changed), and the new tests/state-command-cutover.test.cjs is added to the lint-test-file-count allowlist under the 'state' prefix. * refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify) Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS (state landed in the pilot commit). init preserves its #1688 warnIfStaleBake pre-hook; validate binds the output emitter. Cutover test extended to assert all 6 are consumed + owned. Golden install-parity fixtures regenerated. |
||
|
|
15b3cc8690 |
docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a two-layer dispatch (registry families + leaf-verb table filling the prepared _dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions ADR-959 leaves open: full dissolution, family/leaf classification rule, shared parseFamilyArgs, and the capability-arm extraction shape. Phased under epic #2345 (P1-P4). Behavior-preserving; each cutover proven by the audit-command-cutover equivalence template. - docs/adr/2346-command-dispatch-completion.md (new) - docs/adr/959-*.md: Status Proposed -> Accepted + amendment section - docs/adr/README.md: index rows for 959 + 2346 - docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub - CONTEXT.md: seed glossary entry Closes #2346 (docs-only; no production code). |
||
|
|
fc913b37a5 |
refactor(#2268): gen:golden one-command fixture regenerator (#2275)
Phase 3 (convenience form) of golden-parity redesign (epic #2264). Adds npm run gen:golden (regenerates both fixture sets) and points the golden-parity/tree failure messages at it. Full CI-auto-comment deferred (documented in ADR-2264). Closes #2268. |
||
|
|
89b1bef881 |
refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267. |
||
|
|
ef5a5bc15d |
docs(#2265): ADR-2264 golden-install-parity redesign (#2270)
Phase 0 of golden-install-parity redesign epic #2264. Adds docs/adr/2264-golden-parity-redesign.md + index entry. Closes #2265. |
||
|
|
8cfbdf167f |
docs(#612): bracket phase-id convention ADR (PR-0)
PR-0 of the #612 tracer-bullet sequence: the ADR that locks the contract PR-1..PR-6 execute against. No production code. Rewrites the design for current next (v1.7.0-rc.5): phase-id surfaces now live in src/phase-id.cts under the ADR-2121 single-owner regime (core.cts retired, #1267), and M-NN is a shipped first-class convention rather than a no-adopter RC intermediate. States every blocking requirement from the approved-enhancement comment as an explicit design commitment: - terminal M-NN deprecation, end state two conventions (null + bracket) by forward consolidation via the migrator, not "no adopters" - EMIT/RENDER as one pure pair with fast-check round-trip properties, inside phase-id.cts under the #2128 token-source / drift-lint regime - single-sourced, generated PR-6 injection block + verify parity check - migrator dry-run-default / dirty-tree guard / atomic rollback (the current base's surgical reverse-rename #1542, a deliberate improvement over the requirement's literal "HEAD-sha reset" — flagged) + two-invariant fixture corpus + M-NN lift + HARD-REFUSE on absent project_code - everything gated on phase_id_convention === 'bracket'; null / M-NN paths byte-untouched (never gate on project_code) - concrete collision anchor normalizePhaseName('2-01.02-01') === '02', grounded in the live regexes at src/phase-id.cts:71/79 Guardrail: bracket is core config-gated behavior, not an add-only capability. Per-PR implementation map (PR-1..6) targets current module homes and the CARRY-FORWARD ledger. Adds the docs/adr/README.md index row. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fdf01ba588 | docs(#2207): ADR-2207 — STATE.md Status lifecycle & ownership | ||
|
|
474ca08e06 |
docs(#2171): record the statusline data-source scope boundary (#2178)
Records the statusline scope boundary decided during triage of #2160-2164: the statusline sources only local, read-only data (refine-existing + new-local), never credentials or external/network APIs. #2164 (account-usage segment) is out of scope on this boundary; #2163 (git) is in-scope but on the feature track; #2160/2161/2162 are approved enhancements. - docs/adr/2164-statusline-scope-boundary.md (new ADR, Accepted) - docs/adr/README.md (index row) - CONTEXT.md (### Statusline glossary/seam entry) - .out-of-scope/statusline-account-usage.md (#2164 rejection record) Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5695522d5f |
feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads: getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix (→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite), getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity destructures (antigravity is already on the descriptor-agents path). subagentToolkit flipped undocumented→full (Context7: antigravity.google/docs/cli/features); namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical golden parity for all 16 runtimes. UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's settings.json — non-destructive, idempotent, symmetric uninstall. Added to VALID_PERMISSION_WRITERS + the FinishPermissionWriter union. UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json registering the gsd-core companion MCP server (Gemini-successor mcpServers schema, best-effort — raw schema unpublished). Both writers dispatch from finishInstall. settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable, no absolute paths) is golden-tracked → only antigravity.json changes. Tests: declarative-reference-antigravity extended (source-grep guard across 4 modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) + antigravity-upgrades (permission-writer + mcp_config live-install, idempotency, user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md + connect-gsd-mcp-server docs updated; changeset (Changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7bbc243cf5 |
docs(#2144): index ADR-2143 in docs/adr/README.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a146812c5e |
docs(#2144): add ADR-2143 markdown table + mutation + fail-loud consolidation (Phase 0)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
12d50093e1 |
docs(#2121): add ADR-2121 phase-identifier parsing consolidation (Phase 0)
Phase 0 of the #2121 epic — an ADR-only PR that LOCKS the contract Phases 1-4 execute against. No production code lands here. Locks: - phase-id.cts as the single canonical owner of phase-identifier parsing. - New pure exports Phase 1 adds: parsePhaseFromProse (anchored; fixes the #2111 "Milestone v0.5 complete -> 5" class), stripConfiguredProjectCodePrefix / isForeignPrefixedPhaseQuery (config-aware; the #2104 fix's home), and roadmapPhaseLookupSources moved in as sole owner of the 3-source ordering (fixes the #2114 2-vs-3-source divergence). - Extend-never-mutate on the 12 existing exports (normalizePhaseName has a CRITICAL 84-symbol / 20-caller blast radius) — Hyrum's Law. - The exact exact->numeric->prefix-tolerant lookup ordering. - A behavioral anti-divergence contract: reference-identity guard + scripts/lint-phase-id-drift.cjs scanner, modeled on the repo's proven capability-precedence-parity / package-identity-drift patterns. #2104 remains blocked on PR #2105 and off this epic's critical path. Adds the docs/adr/README.md index row. Docs-only; no changeset required (no-changelog). Closes #2121 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
bc751a64ec | docs(#1990): point adr index at renamed file | ||
|
|
1171499f38 | docs(#1990): remove pre-rename ADR filename | ||
|
|
d0b8eacd3c | docs(#1990): rename ADR to Existing Code Onboarding | ||
|
|
e8fb05e965 | docs(#1990): index ADR-1990 in adr README | ||
|
|
3c7d722ed9 | docs(#1990): add ADR-1990 onboard projection module | ||
|
|
e3262d94d3 |
feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool (/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one runtime gate. - Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend (gate ladder: enabled -> Claude -> backend != inline -> nested+background host -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor + worktree, files_modified overlap -> separate stages, resumeFromRunId, budget). All interpolated identifiers validated script-safe; briefs JSON-quoted. - claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow) for orchestrator invocation. - Two gated loop contributions at wired points (execute:wave:post, plan:post); federated config keys (enabled/execution_backend/min_agent_sdk_version). - ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc. On any runtime lacking the Workflow tool, behaviour is byte-identical to today. closes #1143 |
||
|
|
0a7d41c3f0 |
docs(#1820): add ADR-1820 for the spec-section module seam, fallback toggle, and SPEC↔probe precedence contract
Documents the new architectural surface #1820 introduces, per the contributor-standards ADR requirement (a new Module seam that other code will depend on): - The spec-section detection Module seam (src/spec-section.cts) and its locked exported surface, supply rule, suffix-tolerant header invariant, and ownership boundary (detection only). - The workflow.specless_probe_fallback toggle as a policy decision (default-on, disableable cost-gate over the fallback INVOCATION path, not the verifier<->predicate contract) — records the maintainer 857:66 ruling rather than amending it. - The SPEC-supplied <-> probe-derived precedence & authoring contract: section-level precedence (a SPEC-supplied section is never re-run), one projectProhibitions serializer (no second producer), descriptor-less fallback predicates flag/abstain (never green, never auto-dismissed), no-silent-drop equality. Does not restate ADR-857/550/1606; cross-references them. Resolves the sole remaining review blocker on #1835. Refs #1820 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ |
||
|
|
d671171698 |
feat(#1575): complete agent-converter descriptor cutover for copilot/antigravity + surface path parity
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan (ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution + converter + normalize, matching install output byte-for-byte. - Pass skills:'*' sentinel for agents staging when no surface state modifications exist, so ALL agents are staged (not just those referenced by _calls_agents_). - Declare converted agents kind in copilot and antigravity capability.json; add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js. - Handle copilot .agent.md filename rename in both _copyStaged (install path) and _syncGsdDir (surface path). - Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor- parity.test.cjs asserts applySurface output is byte-identical to installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale- cleanup convergence and prune data-loss coverage. - Update ADR-1235 with cutover progress. Cline remains deferred (rules-only local branch + local/global complication). |
||
|
|
8de2ff9121 |
feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator Third-party capability gates declared via check.predicate were rendered for display but never evaluated (only built-in check.query gates fired; the security capability's gate worked solely via a hard-coded ship.md branch). Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts) that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a bounded sh -c command at the project root (via shell-command-projection.execTool), inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed. Wire a 'check predicate' subcommand into check-command-router.cts and extend the three generic workflow gate-dispatch sites (execute:wave:post, execute:post, plan:post) to route check.predicate gates to the new evaluator. The two-step gate contract (command-failure => onError; block => halt) is unchanged. - src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible - src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags - docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md - tests: 38 unit + integration tests (exit mapping, timeout, interpolation, property-based bijection, malformed-predicate fail-closed, real subprocess e2e) Closes #2008 * docs(#2008): backfill changeset pr number 2011 |
||
|
|
55604e9124 |
fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control The node-test fail-first proof accepted a deceptive content-independent negative test — one that reds merely because GSD_PROHIB_SUBJECT is set, ignoring the subject's content — whenever no cleanFixture was supplied, because #1346's causation control was opt-in. The proof's observed signal (RED) thus diverged from its target (RED caused by content) by default. Make the causation control mandatory for the node-test kind: a descriptor that omits cleanFixture is un-provable (fail-closed), never accepted under the weaker violation-only proof. When a clean fixture is present, fail-first is proven exactly as before (RED on violation AND non-vacuous GREEN on clean). The lint-rule kind is unchanged (its subject IS the linted file; no GSD_PROHIB_SUBJECT indirection). Breaking (Hyrum): a previously-green node-test prohibition with no clean fixture now hard-gates — blast radius is zero in-tree (no node-test prohibition ships today; only the lint-rule local/no-source-grep dogfood). Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in. Closes #1906 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ * docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test) Record the node-test mandatory-causation-control supersede across the governing surfaces: - ADR-1606 (the enforcement decision-of-record): addendum + Decision 4 annotated + the "Mandatory causation control — REJECTED" alternative flipped to accepted (premise no longer holds: zero in-tree node-test consumers). - ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph marked SUPERSEDED, pointing at ADR-1606. - spec-phase.md: check_clean_fixture is now REQUIRED for node-test (was "optional"). - CONTEXT.md: PROHIB.enforce.causation predicate updated. Regenerated the shipped-artifact cascade from the spec-phase.md edit (+149 B, well under the 40960 cap): 16 golden-install-parity fixtures and the workflow size baseline. Refs #1906 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
85ed50cc4f |
test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file that owns each subject-under-test, across 52 existing suites (state, config, frontmatter, roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard, health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe wrappers; 881 subtests conserved 1:1. No new test files. Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations (intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe. Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across 8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md + ADR-0002/443/1235/3524 test-file references. lint:ci green. Part of epic #1969. Closes #1972. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de3ba45d00 |
test(#1971): consolidate 48 gsd-tools CLI regression tests into subcommand suites
Fold 48 issue-named gsd-tools CLI regression files into the canonical test file that owns each subcommand subject (state, roadmap, phase, milestone, audit, config, router/dispatch, stats, verify, health, etc.), preserving every assertion and its origin issue number as provenance (block-scoped describe wrappers, 299 subtests conserved 1:1). No monolithic gsd-tools.test.cjs created — routes into 18 existing per-subject suites. Removes 48 tests/ files. Regenerates regression-name allowlist (271->231), ratchets the file-count allowlist across 6 buckets (audit/milestone/phase/roadmap/state/verify), and makes 10 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 10 stale ids). Repoints one CONTEXT.md symptom ref and ADR-3524's parity-test ref. lint:ci green. Part of epic #1969. Closes #1971. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4f779eda43 |
test(#1970): consolidate 61 install-suite regression tests into function suites
Fold 61 issue-named install/codex/runtime regression files into the canonical test file that owns each subject-under-test, preserving every assertion and its origin issue number as provenance (block-scoped describe wrappers, zero assertion loss — 652 subtests conserved 1:1). Routes: - codex-config.test.cjs +19 (codex config/toml/hooks/adapter/skill surface) - install.test.cjs +18 (node-runner norm, manifest, arg parse, finishInstall) - install-runtime-artifacts +13 (per-runtime conversion + emission) - install-minimal-hooks +6 (hook-event dialects + guards) - path-replacement +2 (opencode absolute pathPrefix) - install-write-confinement +2 (pristine dir writes) - install-regressions +1 (user-artifact preservation) Removes 61 tests/ files → 61 fewer CI processes. Regenerates the regression-name allowlist (271→222), ratchets the file-count allowlist (config 10→9, install 12→9), and makes 13 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 13 stale allowlist ids). Repoints ADR-0009's moved-test list. lint:ci green. Part of epic #1969. Closes #1970. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b51cbf96cf |
feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) (#1946)
* feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) * fix(#1943): re-export VALID_EXTENSION_EVENTS from gen-capability-registry (test import path) * fix(#1943): regenerate capability-registry.cjs + add changeset fragment * fix(#1943): import VALID_EXTENSION_EVENTS from validator, not gen-capability-registry (golden parity) |
||
|
|
69fef7c00e |
docs(#1683): Diátaxis host-integration docs + versioning policy + ADR-1239 Accepted — Slice 3 (#1940)
- how-to: author a host-plugin (external-author guide against the SDK surface) - tutorial: embed GSD in a new host (end-to-end programmatic-cli example) - reference: the Host-Integration Interface (axes, adapters, handshake, profiles) - explanation: interface versioning + deprecation policy (additive vs breaking, PROTOCOL_VERSION bumps, deprecation window) - ADR-1239 Status: Proposed → Accepted (Phases B–E shipped) |
||
|
|
3c13903dcd |
feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills in its mandatory init step, so .planning/config.json agent_skills.<type> reaches the agent on every runtime — including Cursor and /gsd-autonomous, where Skill()-delegated workflow bash init did not reliably execute. - gsd-core/references/agent-skills-bootstrap.md: shared contract (query + Read + dedup guard that skips when <agent_skills> is already in the prompt, so Claude's orchestrator-side injection never doubles) - 22 agents/gsd-*.md: one self-load line naming the agent's own type - gsd-core/workflows/autonomous.md: note that delegated agents self-load - tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS bijection + fast-check property) — Generative-Fix-Divergence guard - docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY row, Changed changeset Closes #1866 |
||
|
|
337eee59d5 |
docs(#1610): ADR for the workflow/agent size-budget ratchet (#1713)
* docs(#1610): ADR for the workflow/agent size-budget ratchet Gives the already-shipped size-governance decision (epic #1074; PRs #1089/#1096/#1097) its first ADR: per-file LF-normalized byte baseline (anti-creep across every workflow/agent .md) + loose tier hard caps (XL/LARGE/DEFAULT), measured in bytes not lines, with an explicit do-not-game-the-proxy clause (lazy extraction only). Distinct from the install-time skill-surface budget of ADR-0010/0011. Verified against tests/{workflow,agent}-size-budget.test.cjs + scripts/workflow-size.cjs: cap constants XL_CAP=98304/LARGE_CAP=61440/DEFAULT_CAP=40960/NEW_FILE_CAP=32768, and the largest XL orchestrator is plan-phase.md (~93,973B) ahead of execute-phase.md (~93,426B) — corrected from the draft. Closes #1610. * docs(#1610): de-rot tier-cap byte figures — cite baseline JSON not a drifting snapshot --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
dfff25f1bd |
docs(#1817): add adr-1817 state.md rebuild derivability contract (#1828)
* chore(context): scrub nul byte from consent-store predicate
CONTEXT.md line 206 (Capability Consent Store predicate) contained a
literal NUL byte between ${realpath(projectRoot)} and <id> documenting
the disk-key join format. The byte was intentional but made the file
binary-detected, breaking grep/rg searches (hit while preparing ADR-1817).
Replace with the 4-char \x00 escape. Preserves the byte-level disk-key
documentation; surrounding prose 'prototype-pollution-safe NUL-joined
keys' carries the semantic context. file(1) now reports 'Unicode text'.
* docs(#1817): add adr-1817 state.md rebuild derivability contract
Phase 0 of approved feature #1817 (epic). Lands the design contract for
the new `rebuild` transition in the STATE.md Transition Module (ADR-1769):
- ADR-1817 (new): `rebuild` is the capstone 11th transition. Six design
decisions: (1) core substrate, non-toggleable, same tier as the other
10; (2) section taxonomy — re-derivable (`## Current Position` prose
from frontmatter, `## By-Phase Progress` table from disk) vs preserved
(`## Session`, `## Decisions`, unknown sections) vs de-duplicated
(`## Session Continuity Archive`); (3) orphaned data is logged + dropped
with a structured audit entry in `## Rebuild Log` (ADR-1411 provenance
principle); (4) idempotency is a hard guarantee — a no-mutation rebuild
appends no log entry; (5) non-overlapping scope with `sync` (3
frontmatter fields, auto-triggered); (6) orthogonal to
`auto_prune_state` (rebuild reconciles with current canonical sources,
prune removes by retention policy).
- CONTEXT.md (STATE.md Transition Module section): list `rebuild` as the
11th intent; add contract predicates mirroring the ADR.
- docs/adr/README.md: add ADR-1817 to the index (Accepted).
Targets the #1776/#1761/#1591 body-drift cluster that survived ADR-1769's
per-field transitions. Phased per ADR-1817: this PR (Phase 0) closes
#1817; Phase 1 (#1827) lands `rebuildCore` + intent dispatch + drift-class
unit tests; Phase 2 (#1826) lands `cmdStateRebuild` CLI + dry-run +
integration tests + docs + changeset.
* docs(#1817): reword state-doctor alternative to satisfy docs-parity lint
The docs-parity-live-registry test scans every docs/*.md (including
docs/adr/) for slash-command tokens and asserts each one resolves to a
live command in the registry. The rejected-alternative #4 in ADR-1817
mentioned a hypothetical `/gsd:state-doctor` workflow, which tripped
the lint (`unknown command token(s): [/gsd:state-doctor]`).
Reword to 'standalone state-doctor workflow' (no slash prefix). The
extractor is aggressive — backticks and space-preceding tokens are both
extracted per the test's own polarity-invariant cases — so the only
sound fix is to not form a slash token at all for hypothetical names.
Verified locally: `node --test tests/docs-parity-live-registry.test.cjs`
now passes 31/31 (was 30/1).
|
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
21f0b4316f |
refactor(#1796): ADR-1769 Path A — finish STATE.md preservation consolidation (#1799)
Extract readModifyWriteStateMd's post-sync preservation block into a pure,
field-classification-table-driven applyStatePreservation in the STATE.md
Transition Module. progress / status / stopped_at now join current_phase_name
as table-governed (getFieldClassification), so a preservation-policy change is
a one-row table edit instead of a per-call-site patch.
This realizes the consolidation ADR-1769 / CONTEXT.md already claimed shipped
('Absorbs readModifyWriteStateMd post-sync preservation block') and routes the
#1264 preservation policy through the single field-classification table — the
bug class is now structurally guarded by the table, not just the call-site
shouldResync flag.
Behavior is byte-identical to the pre-amendment inline block (Hyrum-safe — the
15 readModifyWriteStateMd callers' observable preservation is unchanged):
- state/frontmatter/transition + bug regression suite: 847 pass
- phase/milestone/verify (other RMW consumers): 582 pass
- codex (gpt-5.5/high) adversarial review: CLEAN (58,564-case equiv sweep)
ADR-1769 amendment appended documenting #1796.
Closes #1796
|