f0bb0787c9f550c19f7ccfbaf2cfb4db98218db3
214 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
854c93533c | chore: sync next package version to 1.9.1 | ||
|
|
42f4f184c0 |
fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context verify-summary's Pattern 1 matched any backticked path-like token with no context check, so a prose mention of a future deliverable (`shared/types.ts` in a 'next phase will add…' sentence) was checked for existence and its absence failed the verdict on a healthy phase. #2685 added shape filtering but no context check. - src/verify.cts: both extraction patterns now require a claim label on the line (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no longer matches; genuine labeled claims still do. - gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION so relative claim paths resolve against the project root, not the raw cwd (subdirectory invocation no longer manufactures missing files). Regression tests: prose mention not treated as a claim; prose-only SUMMARY passes; absent claimed file still fails. * chore(#2844): backfill changeset PR 2910 --------- Co-authored-by: Test <test@example.com> |
||
|
|
79ed181ec0 |
fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows, tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow structural pre-pass then no-op'd silently — a hard execution failure read the same as 'optional dependency absent'. (A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is preserved; cmdArgs stays an array. POSIX untouched. (B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an absent binary. Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX negative-space test guards the unchanged bash -c callers. * chore(#2667): changeset fragment * chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true) * fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth CI caught two failures on the first push: 1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout cap's process-group reap (exit 124 never fired; hit the 30s harness backstop) AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY excluded now: real PE executables spawn fine directly; only .cmd/.bat are the CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test (node.exe spawned directly, exit 0). 2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the #2667 fallow pre-pass failure-KIND case statement; acknowledge it. --------- Co-authored-by: Test <test@example.com> |
||
|
|
4232a79396 | chore: sync next package version to 1.9.0 | ||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
82ca13f5a5 |
fix(#2736): write current_phase_name from the transition intent, not the lossy prose round-trip (#2821)
* fix(#2736): intent-first current_phase_name on transitions; dash-first prose precedence Primary: completePhase (adapter) and beginPhase (via readModifyWriteStateMd options) pass the intent-held display name to syncStateFrontmatter as an authoritative override, applied after every derive/preserve/carry-forward step — so the lossy prose round-trip can never destroy a name the transition just resolved. Names containing a parenthetical (`Closer-ruling measurement (D1a)`) now land in frontmatter verbatim instead of collapsing to the parenthetical (`D1a`). Secondary (#1695 AC #3 residual): parsePhaseFromProse prefers the em-dash name when it is a genuine name (not a status keyword, not a `Milestone:` tail), else falls back to the parenthetical — satisfying both first-party writer shapes (`N — Name (aside)` and `N (Name) — EXECUTING`). Still lossy for paren-containing names, which is why the intent-first override is the primary fix. plannedPhase carries no name in its intent, so it is naturally out of scope. Fixes #2736 * docs(changeset): backfill PR number for #2736 fragment * fix(#2736): drop an unnecessary type assertion on result.data StateTransitionResult.data is already `Record<string, unknown> | undefined`, so the cast was a no-op and tripped @typescript-eslint/no-unnecessary-type- assertion (CI lint-tests red on the first push; every test lane was green). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
0408276791 |
chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the ADR's swap amendment: a federated config slice lives inside a capabilities/<id>/capability.json, and three of the five key families had no capability directory until 5a created them. Four key families move to the lanes that use them; the central-schema removal and the federated addition land in this one commit because the exclusivity invariant fails the build on a key present in both. review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy ACROSS lanes and stay central. Two things the issue did not name, both found while building it: 1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys against manifest.validKeys only, and two of the four families (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>) were pattern-backed. That is not cosmetic: isCentralConfigKey consults those patterns and mergeFederatedConfig skips every key for which it returns true, so declaring a slice while the pattern survived would have shipped an INERT slice behind a green gate — the exact half-migrated shape the invariant exists to prevent. The gate now loads the patterns from the same manifest the runtime reads. 2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a federated key always resolves to its declared default. The three fallback guards in review.md checked only empty-or-"null", so a user who set the GLOBAL review.max_prompt_tokens would have silently lost trimming on the HTTP lanes. The guards now treat 0 as unset. D9 says review.models.<slug> is owned by "the lane whose slug it names". That is false for one lane: the shipped key is review.models.agy while the slug is antigravity. Ownership follows the lane; the key name is preserved, because renaming would break every config that sets it. Existing tests updated rather than left asserting the old world: config-get on a cleared federated key yields empty instead of not-found (what #2046 actually protects — never persisting the literal "null" — is unchanged and still asserted); the config-schema dynamic pattern representative moves to reviewer_instances; the prototype-pollution guard case moves to a surviving dynamic prefix so alert #26 keeps its coverage, with a new case asserting the old key is now rejected earlier; and Phase 2's harvest-widening inertness assertion becomes an ownership assertion, since Phase 4 is what consumes it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives A federated config key always resolves to its declared default, so an unset per-lane prompt budget needed a value the workflow could treat as 'not configured'. The first cut used 0 — which is wrong: 0 is already a LEGITIMATE per-lane budget meaning 'do not trim this lane' (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. The sentinel is now -1, which is not a valid token budget, so all three states stay distinguishable: unset falls back to global, an explicit 0 disables trimming for that lane, and an explicit N is used. Locked by three CLI round-trip tests. Surfaced by the isolated security reviewer before it crashed mid-run; verified independently against the shipped trim guard rather than taken on trust. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): update central-registration assertions and stay under the review.md cap The remote runner caught both; my local sweep missed the files. 1. tests/plan-review-convergence.test.cjs asserted the three local-server host keys are in VALID_CONFIG_KEYS. They are federated to their lane capabilities now, and the exclusivity invariant forbids a key living in both places. What #2306-local actually protects is that config-set ACCEPTS them, so that is what is asserted — via isValidConfigKey, the predicate config-set itself uses, which spans central and federated. A second assertion pins federated ownership, so a silent reversion back to the central schema fails too. 2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap is a red line, not a budget to raise. The three per-lane budget guard comments were near-identical; condensed to one terse line each. 61371 bytes, 69 to spare. Real extraction to workflows/review/modes/ is Phase 5b/6 work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs Isolated security review findings. MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse error and returned [], while its sibling loadCentralConfigKeys, reading the SAME file, writes to stderr and throws ExitError(1) on that identical failure class. Fail-open here defeats the gate this function exists to feed: with zero patterns, validateCrossCapability's pattern-collision check silently passes and an inert federated slice ships green. It was masked in the one production call site only because loadCentralConfigKeys runs first against the same path — a coincidence of ordering, not a guarantee, and this function is exported and called standalone. The two now share a contract: ENOENT is the legitimate absent case, anything else throws loudly. A single unparseable PATTERN is still skipped, which degrades to "checked less" rather than blocking every build. The branch had zero coverage; it now has two tests (malformed JSON, EISDIR). MINOR — docs/CONFIGURATION.md still listed review.models.qwen and review.models.cursor as settable, ~770 lines below this PR's own new Ownership section. Those lanes take no model flag, so they declare no model key and config-set now rejects them. Rows removed; the missing review.models.agy row added; the per-reviewer budget row corrected to name only the lanes that own a budget key, and to document that a per-lane 0 disables trimming for that lane. Also fixes a shadowed "raw" binding introduced by the fail-closed change, which made the generator unrequirable — caught immediately by its own --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2797): backfill changeset pr number to 2841 * fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next stays green, and it is currently blocking at least three unrelated PRs (#2841, #2832, #2827) with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees The existing loop comment attributes this to `git add .` rehashing O(n^2) blobs "before the object write had landed" and works around it by staging one path per iteration. That is not the cause: sequential execFileSync calls cannot race each other's object writes, and the failure persisted after that change — it simply moved to a lower commit index. The cause is that the git() helper inherited the runner's environment. A leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole operation. In every case `git commit` then cannot resolve a blob it just staged, which is precisely the error above. Verified by negative control: with GIT_DIR exported, this test fails on the unfixed helper (the git commands operate on the wrong repository entirely); with the helper stripping GIT_* it passes. The single-path staging is kept — it is genuinely less work — but it is no longer load bearing. Found while shipping #2797. Fixed in place rather than deferred: it is a defect surfaced during the work, and it is blocking other contributors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2452): build the base-advance commits empty, removing the lost-object class The base-ref guard has been failing in CI with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees It is currently red on at least three unrelated PRs (#2841, #2832, #2827) while next stays green. Two theories have now been tried and neither held. #1881 blamed `git add .` rehashing O(n^2) blobs and switched to staging one path per iteration; the failure moved from commit 32 to commit 25 and carried on. The preceding commit here made the git helper hermetic against leaked GIT_* environment — that IS a real vulnerability (with GIT_DIR exported the helper operates on the wrong repository entirely, proven by negative control) but it produces a different error than CI reports, so it is not demonstrably the cause either. Neither trigger reproduces off-CI, so this stops guessing at the trigger and removes the failure CLASS instead. The loop needs base-branch DEPTH and nothing else: no assertion reads these commits' contents, and base-side files cannot appear in `origin/base...HEAD` regardless. `--allow-empty` writes no blob and no tree, so there is no object for the index to reference and lose. It is also far less work than 60 write+hash+index cycles. The guard still proves its mechanism: the test asserts that a --depth=1 base fetch FAILS and a full fetch resolves, so a broken topology would surface immediately rather than passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Test <test@example.com> |
||
|
|
6a9babda69 |
chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3. - Five reviewers GSD never installs into become lane-only role:reviewer capabilities with no runtime body, no runtimeCompat and no install surface: gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail, which is now deleted outright. - The six hosts that are ALSO reviewers gain a reviewer body alongside their runtime body. Their runtime bodies are byte-identical to next -- verified per capability against the git blob, not asserted -- so no install behaviour moves. - KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived legacy alias for one release; where a capability carries both, the body wins and the slug appears once. Alias removal is Phase 7 (#2801). THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity, claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama, opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it, and the test asserts that literal list rather than a count. kimi-code is deliberately NOT declared here. It is net-new with no invoke_reviewers leg, so declaring it now would make it selectable but not invocable -- present in --all, selected, emitting an empty section for the whole 5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a reviewer at all and gains nothing. The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field, including probe and invoke sub-fields. All eleven are byte-identical, key order included. The epic's premise is that the manifest and the core descriptor describe the same lane with NO translation layer, and Phase 2's review already caught one divergence that every other test missed. Two ADR corrections folded in, as Phases 1-3 each did: 1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798 claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable: D9 assigns review.<host>_host to lane capabilities that do not exist until THIS phase creates them, and a federated config slice must live inside capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4. 2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs bin/lib/*.cjs modules, not capability directories -- antigravity, opencode and qwen appear zero times in it -- and gen-inventory-manifest --check passes with the five new dirs and no edit. Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported LANE_SLUG_RE. The prose had not followed the code. Closes #2798 * fix(#2798): catalogue reviewer capabilities in the generated matrix The capability matrix rendered exactly two tables, feature and runtime, via renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD role, so every role:"reviewer" capability was silently dropped from the first-party catalogue. The drift guard did not catch it, and could not: --check compares generated output against the committed file, and both omitted the five lanes identically, so it reported "up to date" while five shipped capabilities were invisible in the one document that is supposed to list what ships. A guard blind to an entire role is not guarding. This phase is what exposed it -- it ships the first role:"reviewer" capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect found while working is fixed in the current change, which overrides one-concern-per-PR). Verified red-before-green: with a lane row deleted from the matrix, --check now exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so there was nothing for the guard to compare. Phase 6 (#2800) still owns enriching the matrix with lane-specific detail (slug/flag/transport columns) and the locale parity gate. This is the narrower fix: the capabilities APPEAR at all. * fix(#2798): close two hardening gaps and record three limits durably Isolated security review (5 targets, no blockers) reproduced two gaps in the new deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for reuse and carries no other validation, so it must not depend on its caller. - A whitespace-only slug passed the length>0 test verbatim and occupied a roster entry it could never match. Slugs are now trimmed before the emptiness test. A blank body correctly falls through to the legacy alias rather than DROPPING the lane, which would have been worse than the blank slug. - KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there breaks import for EVERY consumer rather than degrading selection. It is now guarded, yielding an empty roster on a malformed registry. That is a visible degradation, not a silent one: under D4 an explicitly requested reviewer that is unavailable is an ERROR, so /gsd:review --claude against an empty roster fails loudly. This also removes an asymmetry -- the sibling capability-trust module documents its collectors as TOTAL and wraps them for exactly this reason. Also records three findings that previously existed ONLY in squash-merged PR bodies, which is not a durable record: - ADR-2782 D5 gains an implementation note explaining why the resolved host is deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds the resolved host; the loader has no config resolver, so folding it in would make the loader and lifecycle compute different signatures for one manifest and re-prompt forever. The binding is split: signature covers the SHA-pinned manifest fields, the consent record stores the resolved host, and Phase 5b re-resolves at invocation -- which is where rule 4 already puts the check. A reader comparing rule 1 to the code would otherwise conclude it is unimplemented. - CONTEXT.md's capability-trust entry still described THREE executable surfaces. Phase 3 added the fourth and made that false; corrected here, since it is drift this epic introduced rather than Phase 6's new-glossary-term work. - stableJson documents the NaN/Infinity/undefined -> null signature collision and why it is unreachable (JSON grammar has no such literal, so JSON.parse throws first). Reachability rests entirely on the ingest path staying JSON.parse-only, so the note lives where someone would break it. * chore(#2798): backfill changeset pr number to 2837 |
||
|
|
46e84d5e39 |
chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat (#2823)
* chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat Phase 2 of epic #2782 under ADR-2782. Delivers D1, D2, D3, D7, D8 and the four Phase-1 vocabulary amendments (A1-A4). - VALID_ROLES gains "reviewer"; the reviewer body is admissible on role:runtime (a host that is also a reviewer keeps one manifest) and on the new role:reviewer (a lane that is not an install target). A reviewer body on role:feature is an error: declaring one is an assertion of lane-ness. - validateReviewerBody + validateLaneProbe + validateLaneInvoke: nine closed enums, a transport discriminator selecting mutually-exclusive invoke sub-shapes, bounded probes (D7), and outputArg required-iff outputChannel is file-arg and forbidden otherwise. - Absent-safe (D4.1): only `undefined` is absent. null/{}/[]/false/0 are malformed assertions and error. 39 of 39 shipped capabilities depend on this. - collectReviewerWarnings: an unknown field inside the body warns, never errors, so a forward-built manifest degrades visibly instead of failing the build. - D8 uniqueness (slug / flags / reviewsSection) lives in validateCrossCapability, so it is enforced at build time over first-party AND at load time over the merged first-party union overlay set, with first-party-wins falling out of the loader's existing ordering rather than a new provenance check. - Config harvest widened past the role==="feature" branch in both the generator and the ownership loop. The often-cited cause of the stranded reviewer config keys -- the runtime body forbidding feature-only fields -- is not the mechanism: `config` is not in FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME. The cause is two harvest sites that never read it. Verified inert: no shipped capability declares config on a non-feature role, and the generated registry is unchanged. Three ADR corrections are folded in (Phase 1 set the precedent of amending in-phase): the misattributed config-stranding cause, D3's inverted profile-membership claim, and the specified capability folder names for lm_studio / llama_cpp, which would have failed the id kebab-case invariant. Closes #2795 * chore(#2795): collapse nine enum checks into one validateEnumField helper Standards-axis review findings, both applied: - Duplicated Code: the enum-membership + enumerate-the-members error shape repeated near-verbatim at nine call sites. Routing them through one helper makes "the error names its valid members" structural rather than a convention repeated nine times, where it would drift. That property is load-bearing until Phase 6 ships the prose reference, because these errors are currently the only documentation of the vocabulary. - Speculative Generality: the isReservedName() pre-check on every enum field was inert. A VALID_* set never contains __proto__/constructor/prototype, so membership alone already rejects them, and "must be one of: ..." is more actionable than "is a reserved name". The literal guards remain where they do real work -- the key-derived write sites in the registry generator and the claim() accumulator. The reserved-name test now asserts all three reserved names are rejected via enum membership, rather than one name via a branch that no longer exists. * fix(#2795): align lane slug grammar with Phase 1 and wire the load-time diagnostic channel Spec-axis review findings, both applied. (1) The slug grammar had diverged from Phase 1's core descriptor. Phase 1 exports LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/ (leading digit permitted); the manifest validator required a leading LETTER. A slug the core descriptor accepts -- a model-named lane such as 4o-mini -- would have been rejected by the manifest validator, which is exactly the translation layer ADR-2782 exists to delete. It was inert only because all eleven shipped slugs begin with a letter, so nothing else would have caught it until a third party shipped such a lane. The grammar cannot be reduced to one definition: Phase 1's module compiles to gitignored build output, and capability-validator.cjs is a committed plain .cjs that must load on a fresh worktree before build:lib has ever run. That makes this the repo's DEFECT.GENERATIVE-FIX class, so the duplication now carries a parity assertion -- laneSlugGrammarMatchesPhase1Descriptor -- which compares both the source grammar and the accept/reject verdict for a shared input set, and fails if the two ever drift again. (2) collectReviewerWarnings had exactly one caller: the build-time generator, which only ever sees first-party in-repo manifests. The real third-party overlay loader never called it and ValidatorModule did not declare it, so ADR-2782 D4.3 -- an unknown field inside a reviewer body is ignored WITH A WARNING -- surfaced nowhere at runtime, which is precisely the case D4.3 exists for. loadRegistry now collects those diagnostics on the accept path, behind a typeof-guard (an older built validator without the function still loads) and a try/catch (ADR-1244 D2's never-crash contract outranks a diagnostic). They land in a NEW OverlayMeta.diagnostics field rather than OverlayMeta.warnings, because warnings records capabilities that were SKIPPED and a consumer treating every entry as inactive would mislabel a working lane. Covered end-to-end by overlayLaneWithUnknownFieldIsAcceptedAndDiagnosed, which drives a real global-scope overlay through loadRegistry and asserts the lane is accepted, produces no skip warning, and yields a diagnostic naming the field. * fix(#2795): make the reviewer validators honour their documented totality contract Isolated adversarial review finding (MAJOR), reproduced by execution. validateReviewerBody documents itself as "TOTAL: returns an array of error strings for ANY input and never throws", and the overlay loader contracts every validator to RETURN errors -- #1461 OVL-1 records a validator that THREW and would have crashed every consumer of loadRegistry. The contract was false at ten sites: JSON.stringify throws on a BigInt and on a circular structure, and every enum/scalar rejection path interpolated the rejected value into its own rejection message. Reading the value could throw too, before any message was built, via a throwing getter or a Proxy get/ownKeys trap. Not reachable through a capability.json today -- every ingestion path is a plain JSON.parse of file text, which cannot express any of those shapes. Fixed anyway: the contract is stated on an EXPORTED function, and a caller must not have to re-derive today's reachability analysis to know whether it holds. Two layers, because serialization safety alone is insufficient: - describeValue() renders any value without throwing, so messages stay useful (a BigInt now reads "got: 10n" rather than degrading to a generic fallback). - A structural try/catch around validateReviewerBody and collectReviewerWarnings makes the guarantee absolute rather than argued, covering read-time throws that fire before any message exists. The same review found the property test guarding this contract was FALSE CONFIDENCE, which is the more important half. fc.anything() at default constraints emits no BigInt, no circular reference, no getter and no Proxy -- 20,000 sampled draws produced zero of each -- so the test was named for a contract its generator could not reach. Even withBigInt is insufficient under whole-value fuzzing, because the defect needs an exotic value in a specifically NAMED field and random key names never land on one. The property is now field-targeted across all twelve reviewer fields, and a companion test enumerates the shapes fast-check cannot generate at all (BigInt, circular, throwing getter, symbol, function, null-prototype) across scalar positions, array-element positions, and read-time traps. Verified red-before-green: with the fix reverted both property tests fail; with it restored all 119 pass. * chore(#2795): backfill changeset pr number to 2823 |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
09477f925e |
fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading The Workflow backend emitted every agent() call with no model, so model_overrides / model_policy / model_profile were silently inert on that path while the inline path honored them (ADR-1411). Neither existing suite contained the string 'model' at all. The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor') rather than hardcoding either, so it asserts backend parity rather than a fixed string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids reaching the code generator, the #2285 composed seam, CLI config-defaulting, and a fast-check round-trip property. RED expected: no model key is emitted anywhere, and --executor-model does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): thread the resolved executor model into the Workflow backend The Workflow backend emitted every agent() call with no model at all, so model_overrides / model_policy / model_profile_overrides / model_profile were silently inert on that path while the inline path honored all of them. The model was not dropped at the last step — it was absent from the whole seam: agentOptions() took no model, EmitInput had no field to carry one, and ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could not forward one. The generated script asserted the parity it broke. VERIFY-FIRST, which #2686 flags as the question that decides the fix: the Workflow tool's agent() DOES accept a per-call model. Its documented signature is agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? }) so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled out. ADR-1143:24's option enumeration omitting `model` is an incomplete enumeration, not a decision to exclude it. - agentOptions(p, executorModel) emits `model` only when it is a non-empty string that is not "inherit" (#2517: an empty model 404s on runtimes without native tier aliases). A non-string is a malformed config: omit, never throw. - executorModel threaded through EmitInput and ResolveWaveDispatchInput. - The CLI resolves gsd-executor from project config by DEFAULT rather than requiring a flag, reading the same source the inline path reads. An orchestrator that never learns about a new flag would otherwise silently keep the old bug. --executor-model exists only to pin/override. - ADR-1411 provenance: the generated header now states which model was applied, or that none resolved and why. A fallback must be a visible value. Compatibility: when nothing resolves, the emitted options object is byte-identical to before, so every existing caller and assertion is unaffected. Behavior change (Hyrum's Law): opted-in users move from session inheritance to the catalog-resolved executor model. Adding a `model` key also changes agent() opts, which invalidates the cached prefix of any in-flight resumeFromRunId run — a one-time re-execution. Both disclosed in the changeset. Fixes #2686 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): reject script-breaking model ids and share the emit predicate The isolated adversarial review found a BLOCKER in my own provenance comment, proven by execution (the emitted script exited 42 from an injected statement). U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment in EVERY engine — the ES2019 change legalized them inside string LITERALS only. So quoteString (JSON.stringify) is sufficient for the `model: "..."` object literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in a model id closed the comment and made the rest of the line live top-level code. The value is reachable from `.planning/config.json` (model_overrides / model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any non-claude runtime — attacker-influenceable in a cloned repo. `emitWorkflowScript` now rejects a string executorModel carrying any character in UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to phaseDir/runId, which is proof the codebase knew this hazard. Rejection is ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an emit failure to the inline backend WITH that reason, so the degradation is visible. A non-string stays on the existing defensive path (omit, never throw) — that is malformed config, not an injection attempt. Also from the reviews: - The predicate deciding "is this model emittable" was duplicated between the emission and the comment asserting it. Extracted to emittableModel() so a generated comment can never claim something the generator did not do — the exact failure class #2686 was filed for. - That predicate now trims and lower-cases before comparing, closing a real #2517-class gap: " " and "INHERIT" were previously emitted verbatim. - The adversarial test was pass-always against this very vulnerability — it asserted only that JSON.stringify appeared. Replaced with the real contract (rejection) plus an execution-level check that no LineTerminator survives into the comment. A raw U+2028 had also been committed into that test's fixture array where a tab was intended; both are now explicit \u escapes. - optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated model contained — silently not testing what it claimed. Now brace- and string-aware. Stale-test corrections in tests/fix-2285-*: three assertions froze the exact options literal `{ agentType: "gsd-executor" }`. The object legitimately gained an optional additive `model` key, so they now assert the invariant they exist to protect (agentType present, isolation absent) rather than a frozen literal. The CLI-vs-pure equality test pins --executor-model on both sides; otherwise it compared a config-resolved CLI run against a pure call given no model. CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): fix the options extractor and model the rejection path Two defects in my own test helper, caught by the full matrix: - optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The emitted shape is agent("brief", { ... }), so that never matched and the helper returned an empty array, making every assertion over it vacuously true. It now anchors on agent( and takes the first balanced, string-aware {...} after it. - The fast-check property predated the security fix and asserted ok:true for any generated string. Strings carrying an unscriptable character are now rejected, so the property models the real three-way contract: unscriptable -> ok:false; trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the trimmed value. Verified locally against the built module: omit values clean, both plans carry the model on the parity path, property passes 500 runs at seed 42. Test file re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are explicit \u escapes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): scope no-control-regex on the mirrored unscriptable-char class The class is the point of the assertion — those bytes are exactly what must be rejected — so the rule is disabled at that line rather than the class weakened. UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence the mirror. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2686): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c07734216f |
fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19 installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped with no citation and no evidence quote. Sourcing every axis independently against Kimi Code CLI's own docs (the issue's explicit requirement — `kimi` and `kimi-code` are distinct products) showed three values had been inherited from the Python `kimi` descriptor rather than sourced: - `embeddingMode` imperative -> declarative. Kimi Code plugins are a `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic API (docs/en/customization/plugins.md) — the same shape as `codex`. - `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md). The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not. - `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth bound is published, so the fail-closed sentinel applies over a guessed integer. `dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is registered with the host and `resolveDispatchType` maps every role onto coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance under Documentation gaps instead. Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/ backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499), and the install adapter is not selected by `embeddingMode` (install.js:543 always uses the imperative adapter). The one visible effect is the curated profile pin, which moves programmatic-cli -> declarative-cli. Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit member that has been in the closed vocabulary since kimi-code shipped. Same defect class and countermeasure as #2598: pin the corrected values and require the matrix to agree with the descriptor, because a descriptor/matrix disagreement is how the gap survived. Closes #2603 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed Surfaced by the orthogonal review of this change. `negotiateHostCapabilities` emits a sentinel-specific warning for every dispatch sub-axis carrying the documented `undocumented` value — namedDispatch, nested, background, subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell through to the numeric guard and reported `host dispatch.maxDepth is missing or not a number — treating as 0`. That message is indistinguishable from a genuinely malformed descriptor, so a correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this PR's kimi-code correction adds a seventh, which is why it is fixed here rather than left in place. The numeric guard keeps firing for genuinely malformed values; both paths still degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests, including the boundary case that the sentinel carve-out must not swallow a real malformed value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2603): backfill changeset PR number (#2687) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a3853472de |
fix(#2598): declare OpenCode subagent dispatch synchronous, not background (#2682)
* fix(#2598): declare OpenCode subagent dispatch synchronous, not background capabilities/opencode/capability.json advertised dispatch.background: true and dispatch.backgroundDispatch: true. negotiateHostCapabilities and every degradationFor / shouldFlattenDispatch consumer trusts these per-field values, so declaring a capability the host lacks OVERSTATES it — the opposite of the fail-closed posture the negotiation is built for. The issue's own citations needed checking before acting: the host-integration matrix (ADR-1239's designated deployment source-of-truth) documented `true` with NEWER evidence than the issue cited, and explicitly marked the issue's sst/opencode#5887 reference as a stale snapshot superseded by #2087. git log confirms #2087 deliberately flipped these from false to true, citing OpenCode v1.15.0/v1.17 as "background subagents enabled by default in all modes". Applying the issue as filed would, on that evidence, have REGRESSED a deliberate update. So the claim was verified against current upstream rather than either document. `packages/opencode/src/effect/runtime-flags.ts` on `dev` today reads: experimentalBackgroundSubagents: enabledByExperimental("OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS") `enabledByExperimental` falls back to the `experimental` flag and `bool()` defaults to false — the parameter is hidden from the model unless an operator opts in by env var. Upstream #29638 is still OPEN and confirms the session loop `tasks.pop()`s one subtask at a time. #2087's "default-on in all modes" reading does not hold against current dev. The issue's CONCLUSION is therefore right even though part of its evidence was superseded: concurrent dispatch cannot be relied on, so both fields are false. The matrix rows are corrected with the verified citation rather than reverted to the old #5887 quote, so the record shows why the value is false TODAY rather than re-asserting evidence that was legitimately superseded. Neighbouring sub-fields are untouched and pinned by test: namedDispatch, subagentToolkit, and isolation:'orchestrator-worktree' (which works via `opencode run --dir` at the OS process level and is unaffected — #2584 does not depend on this value either way). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2598): re-pin the dispatch contract tests to synchronous OpenCode dispatch gsd-test on the descriptor change came back FAILED (5 unique, both node versions). The failures were not incidental — they were deliberate contract-pin tests encoding #2087's decision, one named literally "background UPGRADE": tests/host-integration-descriptors.test.cjs - EXPECTED_FLATTEN[opencode] === false (background-eligible) - the derived background-eligible set pin tests/opencode-imperative-reference.test.cjs - "descriptor declares background dispatch true/true (v1.15/v1.17 upgrade)" - "background UPGRADE changes shouldFlattenDispatch: false now" So this is a recorded decision being reversed, not drift being corrected, and it is reversed on evidence: current upstream `dev` gates the capability behind OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS (default false) and upstream #29638 (OPEN) confirms the session loop still handles one subtask at a time. The issue is filed by the maintainer and explicitly directs "update golden-parity / validator fixtures as needed", which sanctions re-pinning. Behavioral consequence, verified: shouldFlattenDispatch(opencode) now returns TRUE, so GSD serializes opencode dispatch instead of trusting concurrency it cannot get. That is the correct fail-closed direction and is safe today — no shipped GSD flow drives OpenCode background waves (per the issue), and isolation:'orchestrator-worktree' is unaffected because it works at the OS process level via `opencode run --dir`, not via the native subagent. Each re-pinned test now asserts the retracted contract in the opposite direction — feeding the #2087 axes back in must still yield "would not flatten" — so a silent re-flip of either field is caught rather than merely un-asserted. lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2598): backfill changeset pr number (#2682) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d08c32048 |
fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable Every emitted script was rejected. Four invalid constructs, the first fatal on its own, so the Workflow backend could never dispatch a wave: 1. no `export const meta = {…}` first statement -> whole script rejected 2. resumeFromRunId("<id>") -> "resumeFromRunId is not defined". It is a Workflow TOOL INPUT parameter, not a script function. The run id still reaches the caller via summary.resumeRunId, to pass as that input. 3. budget(<n>) -> "budget is not a function". `budget` is a read-only object { total, spent(), remaining() } fed by the caller's token directive; a script cannot set it. Recorded as intent in a comment. 4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions". Now parallel([() => agent(…), …]) — passing agent() results directly also started every agent eagerly, before parallel() could bound concurrency. The single-plan stage had its own branch with the same parallel() defect; both branches are now one array-emitting path. Waves also emit phase() calls whose titles match meta.phases exactly, so progress groups correctly. Two secondary defects kept the script from ever being REACHED — which is why this shipped undetected: 5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no scriptable way" to introspect it and told callers to omit the flag, so gate 5 returned agent_sdk_version_unknown on every automated run while `capability state` still reported active:true. True for bash, false for Node: the router now reads the installed @anthropic-ai/claude-agent-sdk version, walking node_modules up the tree and reading package.json directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the SDK's exports map does not expose ./package.json. Precedence: explicit flag > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an unresolvable version still declines to inline. A too-old SDK now reports the truthful agent_sdk_version_below_floor instead of unknown. 6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any invocation without --runtime reported runtime_not_claude on an ordinary Claude project. Now delegates to runtime-slash.resolveRuntime. The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` snippet is removed rather than repaired: it was also shell-dependent — zsh does not word-split unquoted parameter expansions, so it collapsed to a single argv element, argValue() never matched, and the run failed into the same agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto- resolution removes the need for the construct entirely. Verified with the issue's own repro: no flags now reaches the version gate; an SDK above the floor yields backend:"workflow" with a script that parses as a real ES module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids Findings from the isolated review, all fixed. HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was already RED because of it. The registry embeds the fragment text INLINE, so the shipped/installed copy still taught the exact broken contract this PR fixes: the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown" guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of checking its exit code, so I recorded a red chain as green — checking $? now.) HIGH — three existing tests asserted the OLD broken shape and would have failed CI; none was touched by the first commit: tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…") tests/claude-orchestration.test.cjs — .includes('budget(') tests/claude-orchestration-command-router.test.cjs — .includes('budget(') Each now asserts the corrected contract: the id/pool reaches the caller via summary, and neither construct is ever CALLED. Two sibling assertions had also gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the new explanatory COMMENT contains that substring, not because anything is wired. Rewritten to assert the real property. MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked within a wave, but nothing checked wave ids across waves. That was harmless before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a matching meta.phases entry and the tool matches titles by exact string — two waves sharing an id would collapse into one progress group and misattribute the second wave's agents to the first. Rejected at validation, with tests either side of the boundary. MEDIUM — the fragment contradicted itself (its "Manifest construction" header still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never told the orchestrator to pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to a tool-invocation input, an implementer following only the fragment would have silently regressed phase-resume to a no-op. Both fixed. MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and docs/explanation/claude-orchestration-capability.md documented `resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output — teaching the bug as the feature. Updated to the real contract, including the required meta block and the thunk-array parallel() form. (The changeset is `Fixed`, so the docs gate exempts this; it is corrected because it is wrong, not because a gate demanded it.) LOW — the router's top-of-file comment still described the divergent `--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the fix; and inserting resolveInstalledAgentSdkVersion had orphaned resolveDetectionArgs' JSDoc above the wrong function. Both repaired. lint:ci now exits 0 (verified by exit code, not by reading output). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2590): backfill changeset pr number (#2681) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
b897070de3 |
fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard (#2656)
* fix(#2653): regenerate stale api-coverage.cjs + add artifact-sync guard The tracked build artifact gsd-core/bin/lib/api-coverage.cjs had drifted four days behind src/api-coverage.cts: PR 2551 landed the #2366 fix in the source without regenerating the compiled output, so the module that actually ships still carried none of it. Regenerate the artifact, and add scripts/lint-compiled-artifact-sync.cjs to lint:generated-sync so a tracked compiled artifact can never again silently diverge from its source. The check derives its file set from git ls-files rather than a hand-maintained list, and is regime-agnostic: if these artifacts are later untracked and gitignored per ADR-457, the tracked set becomes empty and the check passes trivially. Verified fail-first: the guard exits 1 against the previously-committed artifact (37731 bytes vs 38634 expected) and 0 after regeneration. Also fixes a defect this change surfaced in tests/no-phantom-issue-refs: its PHANTOM list still banned 2551 and 2361, but GitHub numbers issues and PRs from one shared counter and this repo has since reached 2654, so both now resolve to merged PRs. The guard was rejecting accurate citations of them — it failed this very commit for naming PR 2551 as the drift's provenance. Verified by replaying the guard's scan: 1 offender under the old list, 0 under the pruned one. Only 3182 is still a 404. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2653): backfill changeset pr number to 2656 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4a66d62d10 |
feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver
Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.
worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.
resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.
Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: rebuild tracked state-transition.cjs to match #2400 source
The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit
|
||
|
|
ec7978c0b4 | feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) | ||
|
|
d579daa3ed |
docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum * fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value * docs(changeset): backfill PR #2538 for Phase 6 (#2512) |
||
|
|
f654c24a3e |
feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505) * fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping * fix(#2508): remove leftover old-preamble lines (keep prose variant only) * fix #2508: prose-only reference file * test #2508: regen golden install parity after workflow preamble additions * fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines * docs(changeset): backfill PR #2525 for Phase 4 (#2508) |
||
|
|
c2a305c44d |
feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout PR 1 registered the kimi-code EoS descriptor with empty artifactLayout (SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test). PR 2 fills in the install surface: - src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill function. Today it delegates to convertClaudeCommandToKimiSkill (Python kimi-cli) because Kimi Code uses the same Agent Skills format + /skill: invocation per official docs. The distinct function name lets a future divergence land cleanly if Kimi Code's skill format evolves independently. - gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS. - capabilities/kimi-code/capability.json: artifactLayout.global now declares the skills kind with converter='convertClaudeCommandToKimiCodeSkill' + home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code docs: merge_all_available_skills = true default). - tests/installer-migration-install.integration.test.cjs: REMOVE the SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface. - Regenerated capability-registry + capability-matrix + golden install parity + install tree fixtures for kimi-code. * fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump - src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path can dispatch off the descriptor's converter string. - tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count 26 → 27 (added convertClaudeCommandToKimiCodeSkill). * fix(#2454): remove home override from kimi-code skills (inherit configDir) The home:'.kimi-code' override made the install plan resolve skills dest to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's temp configDir to miss the install. Removing it lets skills inherit configDir like most runtimes. * fix(#2454): kimi-code install contract surface is flat-skills (no agents) Kimi Code has NO custom named subagents (per official docs: 3 built-in coder/explore/plan only). The kimi-skills-agents surface expects agents/ gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed to flat-skills which only checks for skills/gsd-* dirs. * docs(changeset): Phase 2 kimi-code install layout Added (#2509) * docs(changeset): backfill PR #2520 for Phase 2 (#2509) |
||
|
|
bf8f320083 |
feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI) PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting GSD's kimi support into two distinct products per the user's directive: - kimi (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python) - kimi-code (new): Moonshot's Node Kimi Code CLI (~/.kimi-code, runtime: node, KIMI_CODE_HOME env) Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json descriptors, not hardcoded branches in install.js. The new descriptor uses the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs). Critical Kimi Code constraint reflected in the descriptor: hostIntegration.dispatch.namedDispatch: false hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan'] hostBehaviors.namedSubagentsSupported: false Kimi Code's official docs confirm only 3 built-in subagents with NO custom- subagent registration (the [subagent] table only has timeout_ms). The kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in kimi-code's artifactLayout. Schema adjustments: - subagentToolkit set to 'undocumented' (the existing escape hatch); the schema enum (full/read-only) lacks a 'limited'/'built-in-only' value. A follow-up PR can extend the schema enum to add 'built-in-only' as a first-class axis value reflecting Kimi Code's documented model. Registration: - capabilities/kimi-code/capability.json (new descriptor, modeled on codex) - bin/install.js: allRuntimes array + --all list + --kimi-code flag - gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases (kimi-code, kimicode, kimi_code) - src/runtime-name-policy.cts: FALLBACK_ALIASES map - gsd-core/bin/lib/capability-registry.cjs: regenerated via scripts/gen-capability-registry.cjs --write Tests: - tests/multi-runtime-select.test.cjs updated for the new runtime count (18) + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19. Out of scope for PR 1 (follow-up PRs in the sequence): - Install-time decision logic (kimi vs kimi-code detection / prompt) - agent-install-check semantics for kimi-code (verify Agent Skills presence) - cmdAgentSkills fallback returning subagent prompt content - Workflow template mapping (named agents → built-in coder/explore/plan) - Migration guidance for users currently on 'kimi' who are actually on Kimi Code - Schema enum extension for subagentToolkit: 'built-in-only' Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS). * fix(#2454): complete drift-guard registrations for kimi-code runtime The drift guards caught every surface that pins runtime enumeration. Each update is mechanical, driven by the guard's named failure mode: - src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code - src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the isKimiCode predicate generator - bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19 - gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry (null/null/null — same as kimi, no model tier defaults until configured) - docs/reference/capability-matrix.md: regenerated via scripts/gen-capability-matrix.cjs --write (kimi-code row added) - tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP: kimi-code → '.kimi-code' - tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated) The capability-registry is already regenerated from the prior commit. * test(#2454): update drift-guard tests for kimi-code runtime registration Multiple drift guards pin runtime enumeration counts and option numbering. Each update is mechanical, driven by the guard's named failure mode: - tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17); 'all 16 flags' → 'all 17 flags' in test names + messages. - tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice test for kimi-code (option 11). Prompt test updated for new numbering. - tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs); EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per docs, same as Python kimi/opencode). - tests/global-config-home-fragment.test.cjs: table-count test renamed 13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier). * fix(#2454): empty artifactLayout for kimi-code (PR 1 scope) The skills kind requires a converter (existing converters are per-runtime like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only registers the descriptor; the actual Agent Skills converter (and a new 'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside the install-time decision logic. Empty artifactLayout.global is valid and means 'nothing to install yet via the layout seam'. Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs (localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as option 11 in install.js's buildRuntimePromptText (renumbered downstream options 11..17 → 12..18, All 18 → 19). * fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode) The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved) for the new kimi-code runtime id. Property names with hyphens are awkward for consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing the conventional PascalCase flag name. The 16 prior single-word runtime ids are unaffected (the regex finds no hyphens). * fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures - tests/runtime-flags.test.cjs drift guard: use proper kebab-case conversion (isKimiCode → kimi-code, not 'kimicode') so the registry comparison doesn't false-positive on hyphenated runtime ids. - tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae single-choice tests for the renumbered options (kilo 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16). - tests/install.test.cjs: Kilo integration option 11→12, prompt test regex updated. - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1. The kimi-code install produces the standard GSD install layout (skills, contexts, references, etc.) — 436 paths, same shape as other runtimes that have no custom converter yet. * fix(#2454): add kimi-code install contract + global config home fragment - src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path instead of falling through to the default '.claude'. - tests/installer-migration-install.integration.test.cjs RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for PR 1; PR 2 will specialize once the Agent Skills converter lands). - tests/multi-runtime-select.test.cjs: fix space-separated-choices test for the renumbered kilo option (11 → 12). - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: regenerated after rebasing onto current next (new planner-reversibility.md from #2471 etc. now included). * test(#2454): skip kimi-code install contract until PR 2 ships install layout The end-to-end install test (tests/installer-migration-install.integration .test.cjs) asserts every allRuntimes entry installs a runtime-specific artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the capability descriptor + flags + labels, but the install LAYOUT (Agent Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and self-removing — PR 2 removes the entry alongside adding the install surface, restoring the contract loop to full coverage. * fix(#2454): restore compact model-catalog.json format (M1 review) Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations' commit used python json.dump(indent=2) which inflated the file from 165→607 lines (every nested entry got expanded) and lost the trailing newline. The semantic change was just a 3-line kimi-code entry. Restored the original hybrid format (top-level indent=2 + inner entries' one-line style) and added kimi-code in matching form. Regenerated golden install parity + install tree fixtures since the model-catalog.json hash changed. * fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code) CI lint-tests job failed on the glossary drift guard (scripts/check-glossary-refs.cjs --check): ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but bin/install.js's allRuntimes array has 18. ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js (missing from CONTEXT.md's list: kimi-code). Missed in the prior commits because gsd-test does not run the glossary check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims to 18 values + kimi-code in the member list. * chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511) * docs(changeset): Phase 1 kimi-code runtime Added (#2511) * test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands * docs(changeset): backfill PR #2519 for Phase 1 (#2511) |
||
|
|
d39ca0c004 |
Merge pull request #2490 from open-gsd/docs/2481-adr-1239-effort-axis
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort |
||
|
|
8cf747724f | chore: sync next package version to 1.8.0 | ||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
909a3b180b |
fix(#2470): install pi's extension as gsd.js so pi actually discovers it (#2478)
* test(#2470): failing-first — pi extension must satisfy pi's auto-discovery filter pi auto-discovers extensions/ entries through isExtensionFile(), which accepts only .ts and .js. GSD installs its extension as gsd.cjs, so pi silently skips it: no /gsd command, no error, no log line. Encodes pi's discovery PREDICATE rather than a literal filename, so the contract keeps holding across future renames, and adds the migration-006 test matrix for retiring the stale gsd.cjs left in pre-fix installs. Red until the fix lands. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): install pi's extension as gsd.js so pi actually discovers it pi auto-discovers extensions/ entries via isExtensionFile(), which accepts only .ts and .js and skips everything else silently. capabilities/pi declared the dest as gsd.cjs, so the extension installed correctly and was then ignored forever: no /gsd command, no error, no log line. Install it as gsd.js. The in-repo source stays pi/gsd.cjs — tests require() it directly and .cjs is unambiguous CommonJS; only the installed name has to satisfy pi, and pi loads accepted files through jiti, which handles CJS and ESM alike. (The reporter's premise that ~/.pi/agent/package.json declares "type":"commonjs" does not hold — pi never writes that file.) Renaming an installed artifact requires a migration record, so add 006 to retire the stale gsd.cjs from pre-fix installs; without it the old path drops out of the manifest and uninstall can never remove it. The migration plans nothing for an unmanifested gsd.cjs: emitting remove-managed there would have the executor downgrade it to preserve-user and mark it blocked, failing the install for anyone who hand-placed their own file. Also pins body-parser >=2.3.0 (GHSA-v422-hmwv-36x6). The advisory reaches the production tree transitively via the Claude Agent SDK and fails the npm-integrity gate, blocking any PR; pinned via the existing overrides idiom. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): address orthogonal review findings + register migration checksum Code review: - pi/gsd.cjs's install docstring still told readers to copy the file to extensions/gsd.cjs — the exact silently-broken state this PR fixes. Anyone following it recreated the bug. - Two stale extensions/gsd.cjs comments in install-minimal-hooks.test.cjs. Security review: - _installNativePluginIfDeclared confined nativePlugin.dir but joined nativePlugin.file onto the validated directory unchecked, so a descriptor whose file carried .., an absolute path, or a NUL byte would have written outside configHome. Not reachable in a shipped build (descriptors are first-party and compiled into the capability registry), but file is exactly the field this PR changes. Confine the full dest path instead; for a well-formed descriptor this resolves identically to the previous mkdir(dir) + join(dir, file). Covered by four new write-confinement tests. Also register migration 006 in the #670 EXPECTED_CHECKSUMS baseline — shipped migration bodies are locked to a committed checksum and a new migration fails CI until it is listed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): never dereference a symlinked managed path when snapshotting fs.copyFileSync follows symlinks, so a managed path replaced by a link had the REFERENT's bytes copied into the migration journal's rollback and backup trees — a gsd.cjs symlinked at a private key would land that key's contents under gsd-migration-journal/. Deletion was already safe (fs.rmSync unlinks the link, never the target); the copy was not. Nothing GSD installs is ever a symlink, so the faithful snapshot of a symlinked managed path is the link itself. copyPreservingSymlink recreates it, which keeps rollback fidelity (restore re-creates the same link) while never reading the referent. Scoped the pre-delete to the symlink branch only, so the regular-file path keeps copyFileSync's overwrite-in-place and a mid-restore failure cannot destroy the destination. The restore-side existence check moves to lstat, since existsSync follows a link whose target is gone and would silently skip the restore. This lives in the engine all six migrations share, so 000-005 are hardened too. Also regenerates the pi golden-parity hash: correcting pi/gsd.cjs's own install docstring changes the extension's content, which the golden suite caught. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): symlink-preserve the in-apply failure-recovery restore too The previous commit routed three copy sites through copyPreservingSymlink but missed a fourth: the catch block inside applyInstallerMigrationPlan, which replays rollback snapshots taken earlier in the SAME apply attempt. Those snapshots are symlinks precisely because of that commit, so the raw copyFileSync there dereferenced them and wrote the referent's bytes to the LIVE install path — worse than the journal-tree leak it was meant to fix, since it is user-visible and at a predictable location. Verified by experiment rather than assertion: with the pre-fix line restored, the managed path comes back as a REGULAR FILE containing the referent's bytes; with the fix it comes back as a symlink and the bytes appear nowhere. The accompanying test injects the failure by letting the delete succeed and then throwing once, modelling a later step failing after the delete. That ordering is load-bearing — an earlier draft injected before the delete, which leaves the live path in place, so the pre-fix copyFileSync hit a same-file collision and threw instead of leaking. That draft passed against the bug it was written to catch; this one fails against it. Adds the missing rollback() coverage as well: a restored symlinked managed path must come back as a link pointing at its original target. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2470): read the backup location from the journal, not the plan The new backup-content assertion read backupRelPath off result.plan.actions, where it is always null: the planner reserves the field and apply chooses the concrete location, recording it in the journal. The assertion therefore failed on "backup path must be recorded for the user" rather than on anything about the behavior it was written to check. Read it from the journal, which is the authoritative record. Verified by executing all four new test bodies in-process against the built engine — the backup file exists and holds the locally patched content. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2470): backfill changeset pr number to 2478 * chore(#2470): backfill changeset pr number to 2478 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a54feb4216 |
fix(#2440): per-counter progress ratchet — total_plans always takes derived value (#2468)
* fix(#2440): per-counter progress ratchet — total_plans always takes derived value Two sites fixed (targeted — existing body-only write tests preserved): Site A — read path: shouldPreserveExistingProgress (state-document.cts:167) removed total_plans from the all-or-nothing ratchet check. It now joins total_phases as an always-derived counter. Only completed_phases and completed_plans keep ratchet behaviour (they are monotonic). This fixes gsd-tools query state.json reporting stale total_plans when a curated completed_plans triggers the ratchet. Site B — write path: applyStatePreservation (state-transition.cts:162) gained a deriveProgressKeys opt-in flag. When true (passed by cmdStatePlannedPhase only), total_plans and total_phases take the derived (post-sync) value instead of the wholesale curated restore. When false (the default — state.update, state.patch), the existing #3242 wholesale protection stays fully in force. This fixes the state planned-phase verb writing a stale total_plans. The opt-in approach preserves all 8 existing #3242/#1264/#500 body-only write tests that assert wholesale progress preservation during non- progress updates. Tests: - tests/state.test.cjs: 4 unit tests for shouldPreserveExistingProgress (total_plans upward/downward/equality + completed_plans ratchet active). - tests/state-transition.test.cjs: 2 #2440 regression tests for deriveProgressKeys=true (total_plans takes derived; boundary at equality). The existing !resync wholesale-restore test stays unchanged (default behavior preserved). References: #2440; #1446 (total_phases read-path fix — same principle); #3242 Bug A (body-only preservation — protection preserved via the opt-in gate); ADR-1769 (applyStatePreservation table-driven preservation). * chore(#2440): backfill pr:2468 in .changeset/mellow-eagles-chatter.md |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
d04e287fa9 |
fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases detectApiIntegration fired on any integration verb co-occurring anywhere on a line with any API noun, treated / as a word boundary (so a first-party Next.js src/app/api/... route path matched the noun "api"), and read any capitalized word before API/SDK/REST/GraphQL as a service name behind a fixed stopword denylist (so threat-model prose like "Resolver-only API" fired). Because the verify:pre seal gate is BLOCKING, a phase touching no external API could not reach UAT without fabricating a coverage matrix. The compound rule now requires the verb and noun to share one clause (sentence punctuation and table-cell walls end a clause) within a bounded word gap. Non-prose spans are excluded before matching: fenced code (already), inline code spans (new stripInlineCode in the markdown-sectionizer seam), and path-shaped tokens. The <Service> API surface rule requires proper-noun position — a clause-initial capitalized word is ordinary English and needs dependency evidence (URL / package reference) on the same line — and rejects compound modifiers ("Resolver-only", lowercase after the hyphen). A phase that integrates no external API now has a first-class, reasoned way to say so: a COVERAGE.md containing "No external API integration: <reason>" satisfies the gate (declaration + rows is contradictory and blocks). The true-positive path is pinned by regression tests: every default-vocabulary positive still fires, including the widest word-gap pairing and the surface-rule-only shape. Fixes #2365 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2365): tighten api-coverage detector per Codex review (round 2) Applies the Codex review findings on the initial #2365 fix: - S-1: a COVERAGE.md "no external API integration" declaration is the human override for a fallible detector, so it must PASS even when detection still fires — but the contradiction is now SURFACED in the gate output (overridden signal count + terms) instead of passing silently. - S-2: verb/noun pairing is now a term-group nearest-pair merge walk over precomputed word ordinals (computeWordStarts / minWordGap), not a match×match cross product — a hostile line repeating one pair thousands of times stays linear instead of going quadratic. - FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`) are kept as noun/dependency evidence rather than being fully masked, so a genuine dependency reference inside code ticks still corroborates. - C-1: the <Service> API surface rule now scans every candidate in every clause; a rejected first candidate no longer shadows a later genuine service. - Cross-clause binding: a verb may bind a noun in the immediately following clause only when its own clause names a service object, within a tight gap — admits "Integrate Stripe, exposing its endpoints …" without re-admitting the unrelated-clauses false-positive class. - Internal-descriptor negative evidence ("internal Payments API", "the internal endpoint") never pairs; URL/scheme matching generalized beyond http(s). All 5 acceptance criteria still hold: the three reported false positives are clean and "integrate the Stripe API" still fires. Built .cjs committed alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) + lint:regression-names clean; affected suites 256/256 green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3) Codex's second-round review found the round-2 tightening had over-corrected into FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal gate silently let through (the catastrophic class, since a missed API surface is worse than a dismissable false positive). Retuned the detector to be explicitly fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external API integration" declaration dismiss the residual false positives. Fail-open false negatives fixed (all now detect): - F1 clause-initial `<Service> API` with a plain follower ("Stripe API for payment processing") — dropped the follower-allowlist / corroboration gate on clause-initial surfaces; a service that is not a stopword, descriptor, or compound modifier is a real name from any clause position. - F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an alphabetic final label now contributes its API nouns; a first-party route path (no dotted host) still does not. - F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the compound path no longer filters nouns on "internal"/"first-party" (Codex: the qualifier can describe the vendor's own API, not the consuming project's). - F4 long single integration clause — removed the word-gap cap entirely: it could not separate a 21-word genuine clause from an 18-word internal one, so the clause boundary is now the whole relationship test. - F5 lowercase cross-clause service — cross-clause binding no longer requires a capitalized "service object". New false positives fixed (all now clean): - F6 a URL token that swallowed a trailing clause comma, merging two clauses — trailing clause punctuation is kept literal so the split survives. - F7 a capitalized internal component authorizing cross-clause binding — the new gate requires a dependent elaboration, not a new coordinate clause opened by a conjunction ("…, then document…"). - F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol and locality descriptors are rejected in the `<Service>` position. - Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms), semantics preserved (148 sectionizer tests unchanged). Net simplification: the fail-closed model removed the round-2 minWordGap / groupByTerm / follower / corroboration machinery (350 insertions vs 445 deletions across the touched files). Under fail-closed, three round-2 negative tests now correctly detect (integration verb + "internal"-qualified noun, and the distant-same-clause case); none were trek-e acceptance FPs. Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown- parsing) + lint:regression-names clean; all 8 review findings reproduced as regression tests, both directions. Built .cjs committed alongside the .cts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4) Codex's round-3 adversarial review found the fail-closed retune had introduced new holes in both directions. Resolved: Fail-open false negatives (now detect): - External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an integration surface and contributes an endpoint noun even when the host names no vocabulary word. A bare domain link with no path ("https://example.com") stays a non-signal, so "Integrate … from example.com, document …" is still clean. - Locality qualification ("internal", "private") no longer leaks across a sentence or clause boundary: only plain spaces may separate the descriptor from the service, so "The cache is private. Stripe API …" now detects. - Cross-clause binding: the fragile head-word cap (which could not tell a genuine "Connect … to Stripe payments, exposing its endpoints" from an unrelated "Integrate … from URL, document …" — both 4 words after the verb) is replaced by a participial-continuation rule: a verb binds a noun in the next clause only when that clause begins with an "-ing" elaboration. This fixes the 4-word-head false negative AND the false positive below at once. False positives (now clean): - Cross-clause no longer binds a finite continuation regardless of separator: "Wire the settings form. Document endpoint props." / "…; document …" / "…, document …" are separate actions, not elaborations. Perf (quadratic → linear): - The trailing-punctuation peel is a backward char scan instead of an unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms). - SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms). Consumer fail-open (blocking gate): - readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is unreadable. On a read error the gate BLOCKS ("could not read the phase scope …") instead of silently certifying no-integration from partial scope — an unreadable plan could be the one describing the integration. Documented fail-closed tradeoffs, now pinned with tests so they are not "fixed" back into a fail-open: a clause-initial capitalized common word before "API" ("Payment API", "Search API") reads as a service name; a long clause pairs a verb with a distant noun; and a CommonMark inline code span that wraps a newline is matched within-line only. Codex judged these acceptable because the COVERAGE.md declaration is a cheap override. One documented limitation remains out of scope: "Integrate Stripe, and authenticate requests with its API" (a coordinate finite clause whose noun refers back by pronoun) needs coreference resolution, beyond a lexical detector. Verified: 379/379 affected + command-router tests pass (+14 new regression tests covering every round-3 finding, both directions); tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5) Round-4 review confirmed the detector's two most complex features generate findings in both directions no matter how they are tuned, because they need a vendor dictionary + coreference the issue rules out in principle. Per the operator's "ship the robust core" decision, both are removed and their gaps are documented rather than chased further: - Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused a fail-open on finite continuations ("Integrate Stripe; use its OAuth endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing endpoint terminology…" — wrongly fired). Detection is now same-clause only. - URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo link) and recreated routine UI-phase false positives. An external URL is evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1"). Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future maintainer does not re-add the heuristics that caused the false positives above: a service named only in a clause separate from its API noun, and a bare external host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md declaration and rare in real phase prose ("integrate the X API"). Also fixed from the round-4 review: - Qualification now survives markdown emphasis ("The **internal** Payments API" stays clean) while still not crossing a sentence/clause boundary. - readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the phase directory or reading the roadmap fallback — not only per-plan-file failures; a missing directory/section remains a legitimate no-op. The declaration-override path surfaces scope_read_error so an incomplete-scope override stays visible. - SERVICE_SURFACE_API_RE length-bound comment no longer overclaims. Net: the detector is same-clause verb+noun + `<Service> API` surface, with path/code/inline masking and a fail-closed posture. All five acceptance criteria hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review) The round-5 sanity review confirmed the detector simplification is sound (all acceptance positives fire, all required negatives clean) and flagged one real blocker plus a nit: - Blocker: readPhaseScope's roadmap fallback could still silently pass an UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent", no exception reached isRealReadFailure, and the blocking gate certified empty scope. Fixed at the source: read the roadmap directly and honor the function's OWN documented contract — null only on ENOENT (genuinely absent), otherwise throw. Both existing callers already wrap it in try/catch expecting that throw, and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by a new e2e test (unreadable roadmap fallback → block). - Nit: the detectApiIntegration JSDoc still described the removed cross-clause participial binding and "every external hostname counts" — corrected to the actual same-clause-only behavior and the names-a-vocab-word URL rule. Verified: full unit suite green; tsc + eslint + lint:regression-names clean. Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): backfill changeset PR number (#2397) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): sync generated capability-registry + recapture install goldens CI surfaced two generated-artifact staleness issues (all failing test shards + lint-tests traced to these, not to a logic defect): - gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API integration" declaration section) but did not regenerate the registry, which embeds an inline copy of that fragment. Regenerated via `gen-capability-registry.cjs --write` — the diff is exactly the fragment text sync. Fixes `lint:generated-sync` and the "committed registry is in sync" + "registry integration" tests. - The 18 golden-install-parity fixtures were stale by exactly one hash line each — `gsd-core/references/api-coverage.md`, which this PR edits and which is a hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that single hash per runtime and nothing else. Fixes the `golden parity — *` tests. No source or behavior change — generated artifacts only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): flip representative-corpus manifest to assert the fixed behavior The #2371 representative corpus (merged into next after this branch was cut) is a known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test fails loudly the moment #2365 is fixed, at which point — per its own contract in representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the assertion checks expectedDetected instead. This is that moment. Removed currentBuggyOutput from the three detector fixtures (nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now asserts detected:false, which the fail-closed same-clause detector satisfies. Notes updated to describe the fix rather than the bug. The #2366 matrix corpus is left untouched — that tripwire belongs to its own PR (#2374). Corpus test: 7/7 pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): skip chmod-000 fail-closed e2e tests on Windows The three fail-closed gate tests induce an unreadable plan / directory / roadmap with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync still succeeds, so the gate never reaches the read-error path and the assertion fails on the windows-latest CI leg. The fail-closed LOGIC is platform- independent (readError → block) and is fully exercised on the macOS/Linux legs; only the method of inducing EACCES is POSIX-specific. Guard the three tests to skip on win32 as well as root, mirroring golden-install-parity's win32 skip. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged; this closes the test/process-hygiene findings. Major: - CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate). - Replaced the banned wall-clock assertion in the "hostile repeated-term line" test (Clock Seams rule — no elapsed-time asserts) with a deterministic signal-count assertion, which also directly verifies the term-dedup that keeps pairing linear (one signal for a 10k-pair line, not thousands). - Rewrote the three fail-closed read-failure tests: instead of chmod 0o000 (a no-op under root / on Windows, the pattern the repo's IO-failure convention avoids) they now exercise the newly-exported `readPhaseScope` in-process and inject the failure by monkeypatching fs.readFileSync/readdirSync to throw, restoring in finally. Deterministic and platform-independent (no skip needed), and they add the ENOENT-is-absence case that the chmod tests couldn't express. Minor: - Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40} service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN (200) on the declaration reason. - Added a fast-check property that fuzzes the tokenizer / clause splitter / masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector is total (never throws), shape-stable, holds detected <=> signals, and is deterministic. readPhaseScope is exported for the in-process tests. Verified: 125 detector + 19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) + lint-regression-test-names + lint-test-file-count clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d0bacc2517 |
fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10 hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates exited 127 ("command not found") on such hosts, and the gates — which only distinguish 0/124/other — misreported a passing build or test as a FAILURE. Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU `timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES, 128+signum on signal), inherits stdio so pipes/redirects work, and reaps the whole process group so a watch-mode runner cannot outlive its budget. Runs before gsd-tools' flag parsing so the wrapped argv stays opaque. Hardened per adversarial review: - On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant that traps SIGTERM was otherwise orphaned holding stdout, hanging captured gates (the exact watch-mode hang the feature prevents). - Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it. - Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout). - Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command position so prose "timeout 30 seconds" no longer false-positives. Resolution lives once in the CLI; all 10 sites call the shared verb. A parity guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if a bare `timeout`/`gtimeout` execution reappears (the portable `command -v timeout` probe form is intentionally allowed). Also fixes the identical bug in the zh-CN checkpoints translation, updates the tests that asserted the old strings, trims a redundant phrase in gsd-verifier.md to keep it under its size hard cap, and refreshes the size baselines + golden install-parity fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2351): add changeset (#2426) * chore: regenerate golden/size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6b196ef638 |
fix(#2423): finalize job syncs next package.json after final release (#2437)
* fix(release): finalize job calls sync-next-version.cjs to bump next after final release (#2423) The release pipeline's 'finalize' job shipped X.Y.0 to npm 'latest' and merged to main, but never bumped 'next' to match. 'scripts/sync-next-version.cjs' exists exactly for this — its docstring promises to run 'for every release type (rc / hotfix / final)' — but it was wired only into the 'rc' job (release.yml:479), not 'finalize'. As a result, after 1.7.0 shipped on 2026-07-15, 'next' stayed at 1.7.0-rc.6 and every npm script banner on 'next' (and feature branches cut from it) reported the stale rc version. Regression of #1104 — closed incomplete (covered rc only, not final). This patch: - adds a 'Sync next branch to the published release' step to the 'finalize' job, mirroring the rc job's pattern at line 479 (uses VERSION from inputs.version rather than PRE_VERSION from steps.prerelease, since finalize does not run the prerelease step); - gates it on !inputs.dry_run + continue-on-error:true (matches rc); - adds tests/release-finalize-syncs-next-version.test.cjs — a structural YAML assertion that fails against the pre-fix workflow and passes after. Failing-first demonstrated during development; the test parses job blocks by indentation rather than grep so it stays valid as the file grows. Root-cause diagnosis: scripts/sync-next-version.cjs:14 docstring admits 'used by release.yml's rc job, which has no next-targeting PR of its own'. release.yml:506-685 (finalize job) had no sync-next-version step before this patch. Every prior rc.N release has a matching 'chore: sync next package version to 1.7.0-rc.N' commit; there is no such commit for 1.7.0. * chore(release): sync next package version to 1.7.0 (#2423) Replays the canonical 'chore: sync next package version to <v>' commit that the release pipeline's rc job auto-produces via scripts/sync-next-version.cjs, for the 1.7.0 final release that shipped on 2026-07-15 (commit |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
50efae13ce |
fix(#2305): stage the shared guard hooks Kilo's native plugin spawns (#2327)
* fix(#2305): stage shared guard hooks for Kilo — drop skipSharedHooksInstall Kilo's capability descriptor declared BOTH hostBehaviors.nativePlugin (a plugin that spawns the shared PreToolUse guard scripts as subprocesses) AND hostBehaviors.skipSharedHooksInstall:true, which suppresses staging of hooks/*.js into the Kilo config dir. The plugin's runHook treats an absent hook script as a silent allow, so every guard it spawned (gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard) no-opped on every Kilo install. OpenCode uses the byte-identical plugin with hook staging on and is unaffected — it is the reference shape. The skip flag predates Kilo's plugin surface: it dates to #1821 (hooks were dead weight for a runtime with no hook consumer), and #2093 added the hooks-dependent nativePlugin without revisiting it. - capabilities/kilo/capability.json: remove skipSharedHooksInstall (regenerated gsd-core/bin/lib/capability-registry.cjs accordingly) - bin/install.js: correct the stale #1821 comments claiming Kilo has no plugin surface - tests/kilo-upgrades.test.cjs: install-fixture tests (global + local) asserting the guard scripts land where the plugin's walk-up resolves them; an end-to-end test driving a disallowed out-of-worktree write through the REAL installed Kilo tree and asserting the guard rejects it; a cross-runtime descriptor invariant (nativePlugin and skipSharedHooksInstall:true must never coexist) - tests/kilo-imperative-reference.test.cjs: flip the pinned assertion - golden fixtures regenerated (kilo now stages the 24 hook files, same set as OpenCode) Fixes #2305 * fix(#2305): warn loudly when a guard hook script is missing (runHook) runHook's absent-file branch returned a silent exit-0 allow — the mechanism that let #2305 ship undetected: with the hooks bundle never staged on Kilo, every PreToolUse guard the plugin spawned resolved to "file not found → allow" with zero signal anywhere. Keep the adapter's design contract (a missing hook must never break the tool call — pinned by the existing adapter test) but make the absence loud: console.error once per hook file, naming the unresolved path and the remediation. Applied identically to .kilo/ and .opencode/ plugin copies (byte-parity guard). Golden parity fixtures regenerated (the installed plugin file's hash changed). Fixes #2305 * chore(#2305): add changeset fragment * test(#2305): include gsd-workflow-guard.js in the staged-guards regression list The native plugin spawns four guards on write-like tool calls — the regression test's PLUGIN_GUARD_HOOKS list covered three. Staging itself was already asserted via the golden fixtures (the full bundle), but the named per-guard assertion should cover every guard the plugin actually dispatches. Surfaced by cross-AI review of PR #2327. * test(#2305): update the #1821 tests that encoded Kilo's false no-plugin premise The #1821 hook-copy test asserted Kilo must receive no staged hooks — the exact behavior this PR reverses (and the cause of all 8 CI failures). Kilo moves from the ZCode "no dead hooks" loop to the OpenCode group, with positive assertions on the new contract: the three guard hooks the plugin spawns, hooks/lib/git-cmd.js, and plugins/gsd-core.js all staged. The integration runtime contract flips kilo packageJson to true (the CommonJS marker ships with the bundle), and the pi contract comment no longer cites Kilo as a no-plugin runtime. * chore(#2305): scope the queued #1821 changeset fragment to ZCode only The fragment still claimed the installer skips hooks for Kilo — rendering both it and this PR's fragment into the same release would ship two contradictory statements about Kilo's install behavior. It now claims ZCode only and notes that #2327 reverses the Kilo half. * chore(#2305): rename changeset fragment to the generator naming convention 2305-kilo-stage-guard-hooks.md -> loud-guard-hooks.md, matching the <adjective>-<noun>-<noun> shape npm run changeset generates (review nit). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
81f7ab4df1 |
refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) Cutover all 55 remaining case arms from runCommand's switch to HOST_COMMAND_ROUTERS. runCommand now contains only its default case (~40 lines): the three-layer dispatch (capability → overlay → host table) plus the unknown-command diagnostic. The 73-case switch is dissolved. Each case body was relocated verbatim to a module-scope route*Command function via a brace-matching extractor; inner break; statements (from _dispatchNonFamily early-exit patterns) were converted to return; (5 arms affected); loop break; statements preserved. Closes #2384 * chore: retrigger CI |
||
|
|
062f3fda90 |
chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test Adds tests/fixtures/representative/ — a permanent corpus of verbatim, incident-sourced fixtures (never author-invented) from #2286, #2347, #2365, #2366, each labeled with its expected gate verdict in a MANIFEST.json and driven through the real CLI gate entrypoint via tests/representative-corpus.test.cjs. Adds a document-shaped fast-check property test alongside the existing writer-seeded bijection test in tests/api-coverage.test.cjs: the existing generator produces rows and renders them through the writer, so the document shape is a constant and it cannot fail against a decoy table; the new one generates the document space instead. Two gates (#2365, #2347) are still open, so their corpus/property assertions are marked with node:test's official `todo` option — the test executes and reports its failure without affecting the process exit code (https://nodejs.org/api/test.html#test-options). The audit-uat corpus (#2286, fixed by #2317) is a normal passing assertion, proving the methodology works end to end and not just cataloguing gaps. Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples; a negative fixture must come from a source that doesn't know the gate exists. No production src/*.cts changes — validation only, per #2371's scope. Closes #2371 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields Standards-axis review findings, all fixed: - Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen) that were copy-pasted between the parse/render bijection test and the new document-shaped property test in tests/api-coverage.test.cjs — a future edit to one could have silently desynced the two properties. Hoisted to a single module-scope declaration both tests reference. - Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome -> expectedReason. It asserted against the gate's `reason` field, but this codebase already has a real, different `outcome` field at parser altitude (extractDecisions' DecisionOutcome) — naming the manifest field after the wrong altitude's term was exactly the ambiguity the "Fixture provenance" rule this PR adds exists to eliminate. - Removed the unused `role` field from three MANIFEST.json files (never read by any test) and wired the previously-dead per-fixture `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's `results` array — catches a regression that moves items between the two fixture files while preserving the aggregate total, which the existing total_items check alone would miss. No changes to test intent or coverage — same assertions, correctly named and fully wired. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while preparing this PR — unrelated to #2371's own changes, but a defect found while working is fixed in place rather than deferred. Leftover from #2368/#2370 (merged just before this branch rebased onto it): the case 'capability' arm that needed these two requires was relocated to bin/lib/capability-command-router.cjs, which already requires both modules directly (lines 24-25) and is their only real consumer (cmdCapabilityState, resolveCapabilityRuntimeState, cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had zero other references in the file and were never re-exported — confirmed via grep across the file and its module.exports. Behavior-preserving: Node's require cache means the underlying modules still load exactly once via capability-command-router.cjs's own requires; gsd-tools.cjs never used its now-removed local bindings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): replace todo-marked assertions with characterization tests gsd-test's own JSONL result parser (gsd-test-runner's internal/pipeline/parse.go, verified directly against that repo's source) has no concept of node:test's `todo` option — it only recognizes kind:"pass"|"fail" and hard-errors on anything else. A { todo: true } test whose body throws is counted as a real failure in gsd-test's own verdict, exactly as if it weren't marked todo — proven by an actual gsd-test run against this branch, which reported outcome:"failed" with all six todo-marked assertions (the property test plus five representative-corpus fixtures) in the failure list, each carrying the correct raw node:test `todo` field the tool's parser simply doesn't read. Replaces todo with characterization: MANIFEST.json now carries both the correct target verdict (expected*) and the exact current observed verdict (currentBuggyOutput, directly verified against live CLI output for all five fixtures). Tests assert currentBuggyOutput — an honest, non-vacuous pin of today's known-broken reality that passes today and will fail loudly the moment the referenced fix changes the observed output, at which point the assertion should be flipped to expected* and currentBuggyOutput deleted. The document-shaped property test switches from throwing fc.assert to non-throwing fc.check (returns RunDetails per fast-check's own docs) and asserts report.failed === true directly, for the same reason. Updates all prose (CONTRIBUTING.md, the fixture READMEs) that previously claimed todo would be respected — that claim was factually wrong for this repo's actual tooling and must not ship. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: regenerate golden-install-parity fixtures after rebase onto next Rebasing onto the current next (which now includes #2381's todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real conflicts in all 18 golden-install-parity fixtures — expected, since both branches changed the same gsd-tools.cjs hash entry. Resolved by taking one side to unblock the rebase, then regenerating fresh from source via npm run gen:golden and verifying the result; every file's diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs, correcting a stale intermediate hash from the arbitrary conflict pick. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e6c16efa6d |
refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) (#2382)
* refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) Relocate 13 case arms from runCommand's switch to HOST_COMMAND_ROUTERS: - resolve: resolve-model, resolve-granularity, resolve-execution (3) - git: git (1) - config: config-ensure-section, config-set, config-set-model-profile, config-get, config-new-project, config-path, migrate-config (7) - research: research-store, research-plan (2) Each case body moved verbatim to a module-scope route*Command function (closures over config/commands/output/_dispatchNonFamily preserved). dispatchHostCommand extended to pass defaultValue + workstreamContext (needed by config-get and config-path). No new files, no logic change. Closes #2373 * chore: retrigger CI |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
cf004df678 |
refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in runCommand's default case after capability/overlay dispatch, before the unknown-command error. Migrates 'state' as the pilot: removes the hardcoded case 'state': arm; state now dispatches default -> dispatchHostCommand -> routeStateCommand, byte-identical to the old path (proven by the new state-command-cutover equivalence test, 5-category template). Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey) — the capability registry stays reserved for toggleable feature bundles per ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked; the ADR is corrected here alongside the code that realizes it. - gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand (prototype-pollution-safe); wired into default case; case 'state': removed; dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests. - tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY equivalence (recording-mock + runGsdTools end-to-end + pollution guard). - docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability- registry model (correction that did not land in the merged #2355). Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/ verify using this proven template. Closes #2360. * test(#2360): regenerate golden fixtures + allowlist for state cutover Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the install-parity fixtures (gsd-tools.cjs content hash changed), and the new tests/state-command-cutover.test.cjs is added to the lint-test-file-count allowlist under the 'state' prefix. * refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify) Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS (state landed in the pilot commit). init preserves its #1688 warnIfStaleBake pre-hook; validate binds the output emitter. Cutover test extended to assert all 6 are consumed + owned. Golden install-parity fixtures regenerated. |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ff9cb6069f |
fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active' but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller outside their own CLI router, and execute-phase.md declared an execute:wave:pre hook point that the workflow body never rendered — so claude_orchestration.enabled:true had zero effect on real runs. Approach B (maintainer-chosen): - execute-phase.md now renders the execute:wave:pre hook (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75, immediately before each wave's Agent() dispatch — fixing the latent dead-hook gap for any pre-wave capability. - Move the claude-orchestration contribution execute:wave:post -> execute:wave:pre (a pre-wave backend selector belongs before dispatch, not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md with prose instructing the orchestrator to call resolve-wave-dispatch before step 3. Unrelated wave:post contributions (ui.safety-gate, drift, external-job, mempalace) untouched. - New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result; exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is a real non-CLI-router, non-test caller of both functions. Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool absent, SDK below floor, execution_backend:inline, malformed input) or an emit failure resolves to inline with a byte-identical result shape — no regression to the default-off execute-phase path. Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a fast-check composition property, capability.json contribution assertions, and a source-contract guard that execute:wave:pre is now actually rendered. Dependent registry-shape assertions updated in-scope. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
20ff405cb3 |
feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline New statusline.state_format config, enum full|compact (default full — existing rendering untouched). "compact" renders the state segment as "<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 · executing" — dropping the milestone name and progress bar (the two biggest width costs) and collapsing narrative statuses to a single keyword. Per the #2162 approval conditions, the keyword set is the canonical vocabulary from normalizeStateStatus() in state-document.cjs (discussing/planning/executing/verifying/completed/paused) — no parallel hand-rolled list, so the vocabularies can't drift — and the canonical stuck state "paused" renders uppercase as PAUSED (no new "blocked" lifecycle state). Statuses the normalizer passes through unrecognized fall back to their first word capped at 16 chars. Lifecycle scenes preserved: active_phase wins over the body phase number, milestone completion renders "complete", idle-with-next-action renders "next <action> <phases>". Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2162): changeset fragment for PR #2175 * fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format - register statusline.state_format in the fix-1628 coercion-bypass matrix - 15/16/17-char boundary tests for the shortGsdStatus fallback cap - changeset body ends with the (#2162) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage - compact renderer gates the milestone-complete scene behind the absence of an in-flight phase id, mirroring formatGsdState's if/else precedence (Scene 1 beats Scene 3); regression test covers the non-atomic active_phase + percent=100 STATE.md shape - direct config-set accept/reject test for statusline.state_format plain strings (ENUM_KEYS matrix covers only the JSON coercion shapes) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests) Re-review Major: gating done on !phaseId held completion back for the legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100 regardless of phaseNum, so compact must too. Gate is now !s.activePhase. The phaseNum-only test now expects 'complete' and cross-checks the full renderer; a parity test feeds identical inputs to both renderers. Re-review Minor: shortGsdStatus gets fast-check property coverage (totality, canonical fixed points, separator safety, fallback shape). Golden fixtures regenerated for the hook byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
27415c780e |
fix(#2252): exclude PLAN-REVIEW artifacts from plan count (#2263)
* fix(#2252): exclude PLAN-REVIEW artifacts from plan count The loose /PLAN/i fallback in isRootPlanFile matched *-PLAN-REVIEW.md, inflating plan counts. Added PLAN_REVIEW_RE exclusion before the fallback. * docs: backfill changeset PR number (#2263) * fix: regenerate stale capability-registry after next merge |
||
|
|
8b70db343b |
fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 completePhaseCore was writing the overloaded bare 'Milestone complete' on the last phase — the same string space the milestone-close verb owns for terminal state. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts). Milestone termination ('<version> milestone complete' / 'Awaiting next milestone') remains solely with milestoneCompleteCore. Status lifecycle: Ready to plan → All phases complete → <version> milestone complete → Awaiting next milestone. Changes: - src/state-transition.cts: completePhaseCore status value - src/phase.cts: #2028 guard comment - tests/state-transition.test.cjs: assertion + test name - tests/phase.test.cjs: 8 assertion updates (positive + negative) - tests/state.test.cjs: normalizeStateStatus test case + reset regex - tests/workstream.test.cjs: fixture status to terminal value - gsd-core/workflows/progress.md: Route D label - gsd-core/workflows/transition.md: Route B label - CONTEXT.md: Status lifecycle glossary entry (ADR-2207) - .changeset/brave-geese-jump.md * test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline Workflow file edits (progress.md, transition.md) changed install payload hashes and pushed past the committed workflow-size baseline. Regenerated all 17 golden-install-parity fixtures + claude-local via the standalone gen script (which now also covers the local-scope claude layout). Updated workflow-size-baseline.json and agent-size-baseline.json via size:baseline. * fix(#2204): correct claude-local golden hashes + document gen-script limitation The gen-script's claude-local generation produces macOS-specific hashes incompatible with Linux CI (local-scope install embeds platform-varying node-runner paths). Reverted to manual update using Linux FAILURES.md +actual hashes for the 2 changed workflow files. Added explanatory comment in the gen script. * test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary Addresses orthogonal code-review findings (Medium #1 + #2): - Add isCompletedInventory test cases for ADR-2207 status lifecycle (terminal 'milestone complete' → true; intermediate 'All phases complete' → false; archived → true; active statuses → false) - Clarify CONTEXT.md glossary: note that isCompletedInventory intentionally excludes the intermediate value * docs: backfill changeset PR number (#2259) * docs(#2204): add Status lifecycle table to state-md reference (ADR-2207) |
||
|
|
c1885df9e5 |
chore(#2143): prohibition-with-teeth + migrate remaining ad-hoc table sites — Phase 4 (final) (#2253)
* chore(#2143): prohibition-with-teeth + migrate remaining table sites — Phase 4 Phase 4 of epic #2143 (ADR-2143 §7). Completes the markdown table/mutation consolidation by (a) giving the ad-hoc-parsing prohibition teeth and (b) migrating the last ad-hoc table sites onto the shared seam. - src/markdown-table.cts: new formatting-preserving `updateTableCell` primitive (self-contained, ragged-row-tolerant header/delimiter/cell-range scan; splices only the target cell's raw span, preserving all other bytes incl. padding/CRLF; no-op-preserves-padding when a transformer returns the current value). Exports splitTableRow/isDelimiterRow/findTableStartOffset for tolerant reuse. - eslint-rules/no-adhoc-markdown-parsing.cjs: TABLE-REGEX detector extended to `new RegExp(<literal|static-template>)`; new `.replace()`-mutation detector for roadmap/state/content receivers with a table/section-shaped pattern. - scripts/lint-table-schema-drift.cjs (wired into lint:ci): fails if a TABLE_SCHEMA header drifts from its authored table; tests import its logic (single source). - Migrated onto the seam (behaviour-preserving vs pre-Phase-4 HEAD, verified byte-diff old-vs-new): roadmap.cts cmdRoadmapUpdatePlanProgress, phase.cts cmdPhaseComplete + traceability, milestone.cts cmdRequirementsMarkComplete, uat.cts read path, state.cts metrics/decisions/By-Phase. - Incidental correctness gains from the migration: a decoy table can no longer swallow a phase-progress update (## Progress scoping); a ragged neighbouring row no longer silently aborts an edit; completing integer phase N no longer touches a decimal sub-phase N.x row; record-metric no longer drops trailing section content or duplicates the ## Performance Metrics section. - Kept justified allow-adhoc-markdown markers only where genuinely not a table (security.cts <|role|> token) or a loose non-GFM section (uat human-verify). Two orthogonal isolated reviews (correctness/adversarial + security) passed; correctness found 4 behaviour regressions in the first migration pass, all fixed and re-verified byte-identical-or-better vs OLD. Surfaced for maintainer (pre-existing, ambiguous domain logic, NOT changed here): templates/state.md places a By-Phase table under ## Performance Metrics while cmdStateRecordMetric assumes a Plan|Duration|Tasks|Files table. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): match traceability row by first-cell value, not Requirement header Phase 4's migration matched the REQUIREMENTS.md traceability row by a column literally named `Requirement` (`row['Requirement']`), but real tables head that column `REQ-ID`. The by-name lookup found nothing, so `phase complete` and `requirements mark-complete` left the Status cell `Pending` (regressed #2769 / #2203, caught by gsd-test — 8 failures, both node 22/24). - src/phase.cts, src/milestone.cts: match the row by its FIRST cell's value (the requirement-ID column) regardless of that column's HEADER name, via `Object.values(row)[0]` (updateTableCell builds the record in header order). This mirrors OLD's first-cell `\|\s*<id>\s*\|` anchor, restoring header-name independence while keeping the seam. - src/milestone.cts hasTable: broadened from `Requirement`-only to also recognize `Requirement ID` / `REQ-ID` / `REQ ID` headers, kept in sync with the now-positional rowMatch/hasRow so a REQ-ID-headed table participates in the ADR-2143 §6 write-set and the #2140 table_unmatched drift check (it was silently omitted before — a checkbox-only partial reconcile against a REQ-ID table could report as fully reconciled). The `Requirement`-headed path is byte-identical to OLD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): replace stale structural milestone guards with behavioural suite The `milestone.cjs regex global state fix` block was a source-structure guard (allow-test-rule: structural-regression-guard) — it readFileSync'd the compiled milestone.cjs and asserted removed regex idioms (`tablePattern.test`, `afterTable !== reqContent`, `doneTable = new RegExp(...)`). Phase 4's migration deleted those regexes (table update is now updateTableCell), making the assertions obsolete. Per the Test Cleanup rule, replace them in-PR with a behavioural suite driving the compiled CLI: - multi-ID mark-complete flips all IDs (guards the lastIndex/global-state class), - Pending->Complete flip under both `REQ-ID` and `Requirement` headers (#2769), - idempotent already_complete detection with no corruption, - REQ-ID-headed table participates in write_set (traceability entry, applied), - REQ-ID-headed table trips #2140 table_unmatched drift on a missing row. Pruned the now-nonexistent structural-regression-guard entry from the lint-allow-test-rule-refs allowlist (the source-text-is-the-product entry for the same file remains valid). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): backfill PR number 2253 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): record-metric targets its own metrics table, not By-Phase velocity `state record-metric` appended its per-plan row (`| Phase 1 P1 | 5min | 3 tasks | 4 files |`) into the FIRST table under `## Performance Metrics` — which on a real template-derived STATE.md is the By-Phase velocity table `| Phase | Plans | Total | Avg/Plan |`, polluting it on EVERY plan completion (execute-plan.md:414 is a per-plan call). The command's own metrics table is `| Plan | Duration | Tasks | Files |`, which the template does not ship, so the row never reached it; the scaffold branch also emitted a wrong `| Phase | Plan | Duration | Notes |` header matching neither the row nor the canonical table. Pre-existing (predates Phase 4); surfaced while migrating this site and fixed here per no-defer, on the user's explicit go-ahead. - src/state.cts cmdStateRecordMetric: locate the metrics table by its own header shape (`Plan|Duration|Tasks|Files`, via splitTableRow/isDelimiterRow) rather than "first table in the section". When the section exists but has no metrics table (only the By-Phase table), self-heal by appending a fresh **Per-Plan Metrics:** table to the END of the section body — By-Phase table, Recent Trend and footer preserved verbatim, no duplicate `## Performance Metrics` heading, created stays false. Absent-section scaffold header corrected to the canonical `| Plan | Duration | Tasks | Files |`. Ragged-tolerance + None-yet preserved. - Not touching templates/state.md (golden-install-parity hashed) — record-metric self-creates the table on first use instead. Failing-first regression test (tests/state.test.cjs) demonstrates the By-Phase pollution on the pre-fix build, then green after. Verified: no pollution, self- heal idempotency, both-tables isolation, content/heading preservation, flags, None-yet, corrected scaffold header (23-check adversarial harness + all existing record-metric scenarios). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteSection seam primitive (level-bounded whole-section removal) ADR-2143 §4 shipped withSection/collectSection (replace a section BODY) but no way to DELETE a section (heading + body). Phase 4 suppressed the phase-remove section delete instead of building it. deleteSection(content, predicate, opts) locates the section via the collectSection machinery and splices out from the heading's start offset to the next same-or-higher-level heading — so a level-3 `### Phase N` delete stops at a following level-2 `## Progress`, never past it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove no longer deletes ## Progress on last-phase removal updateRoadmapAfterPhaseRemoval deleted a `### Phase N` detail section with a greedy raw regex whose lazy scan, on the LAST phase, ran to EOF and destroyed the following `## Progress` heading and its entire tracking table — silent data loss, uncovered by tests (removal tests only exercised a middle phase). Migrated onto the new deleteSection seam (level-bounded, stops at `## Progress`); dropped the allow-adhoc-markdown SECTION-DELETION suppression. Failing-first regression (tests/phase.test.cjs) removes the LAST phase and asserts the ## Progress heading + table survive; middle-phase removal is byte-identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteTableRow seam primitive (row removal, ragged-tolerant) Sibling of updateTableCell: locates the first GFM table, matches a DATA row by predicate (ragged-tolerant record build, header order), and splices out that row's whole line preserving every other byte. Returns {ok:false,reason} on no table / no match. Enables migrating the phase-remove Progress-table row delete off its ad-hoc regex (ADR-2143 §7 — the "future row-delete seam" Phase 4 punted). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove deletes the Progress row via deleteTableRow The Progress-table row delete used a whole-document regex with two defects: (a) `\.?\s` required whitespace after the phase number, so a COMPACT row `|2|Beta|` was never deleted (stale row left behind); (b) unscoped — it could strike a row in a different table (e.g. an earlier `| Phase | Requirements |` table). Migrated onto deleteTableRow, scoped to the `## Progress` section (mirrors deriveProgressFromRoadmap), matching the row by first-cell phase number (integer zero-pad-insensitive; decimal exact; removing `2` never touches `2.5`). Both allow-adhoc-markdown suppressions removed. New behavioural tests: compact unpadded row deleted; padded byte-parity on the surviving rows (their ordinal correctly renumbers via the pre-existing renumber block). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): deleteTableRow leaves no dangling newline on last EOL-less row Deleting the final row of a table with no trailing EOL sliced from the row's start to end-of-string, stranding the newline that terminated the previous line. Back rowStart over the preceding \r?\n in that branch so the table ends cleanly. (Caught by the primitive's own unit test on gsd-test; local scenario checks missed the no-trailing-EOL edge.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): migrate read-only section-collects onto collectSection Six hand-rolled `## Section` read-extract regexes replaced by the collectSection seam (behaviour-preserving; extracted bodies feed the same downstream parsers): state.cts matchSessionSection (## Session / ## Session Continuity) + ## Blockers, smart-entry.cts ## Blockers, audit.cts ## Current Focus + ## Open Questions. Removes 6 allow-adhoc-markdown "pending #1372" suppressions. Incidental fix: the old Session regex `## Session[ \t]*\n` silently failed on a CRLF `## Session\r\n` heading (Windows STATE.md), nulling all session fields; collectSection is CRLF-safe, so session state now resolves on Windows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): fence-safe state-transition section writes + dedup stripFrontmatter - milestoneCompleteCore's `## Current Position` and `## Operator Next Steps` section resets used fence-blind raw regexes that a fenced `##` inside the body could truncate/mis-target (#2130/#2067/#2080 class). Migrated onto a fence-aware tokenizeHeadings-based helper (resetSectionVerbatim) that is byte-identical to the old output on the canonical path (9/9 fixtures) and correctly ignores a fenced fake heading (proven robustness gain). - mutateCurrentPositionFirstTime: hand-rolled locate+splice → collectSection + replaceSection (byte-parity). - stripFrontmatter was inlined byte-identically in state.cts AND state-transition.cts; hoisted the single canonical copy into frontmatter.cts (both call sites now import it) + unit tests — eliminates the divergence risk per CLAUDE.md "Generative Fix Divergence". Removes 3 allow-adhoc-markdown / #1372 markers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): name-address By-Phase sum + uat parse, eslint recall hole, catches - state.cts By-Phase "Total plans completed" sum: positional 2nd-cell regex → name-addressed splitTableRow read (correct on a reordered header, where the old code silently summed the wrong column). Marker removed. - uat.cts parseVerificationItems: loose pipe regex → splitTableRow within the existing table/numbered/bullet union scan (item list byte-identical; does NOT reintroduce the reverted strict-parseMarkdownTable item-drop). Marker removed. - eslint no-adhoc-markdown-parsing: close the `new RegExp(identifier)` recall hole — resolve a const-declared table-shaped regex identifier (mirrors the .replace() detector) + RuleTester cases; param/call args stay out (boundary). - commands.cts: delete a lying comment that claimed the scaffold date "stays on raw UTC / deferred" — #2136 already moved it to realClock.localToday(). - Empty catches (classified, not blind-swept): removed 4 dead try/catch; fixed 3 error-hiding (phase-insert decimal-dir I/O collision now fails loud; phase-remove rename partial-failure surfaced; milestone-archive true count via finally); left best-effort swallows with justification comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): extractFencedBlock seam + migrate api-coverage named fence parseCoverageMatrix extracted its ```coverage fenced block with an ad-hoc regex (the last real allow-adhoc-markdown suppression). Added extractFencedBlock to the markdown-sectionizer seam (reuses stripFencedCode's CommonMark fence engine — info-string match, ~~~/backtick, nesting, indent) and migrated onto it; byte- parity on the parsed CoverageMatrix across 8 fixtures. Only security.cts:367 (a genuine `<|role|>` protocol-token false-positive, not a GFM table) remains marked in src/ — the "prohibition with teeth" goal (nothing grandfathered but a true FP) is met. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): By-Phase row insert is name-addressed (insertTableRow seam) updatePerformanceMetricsSection's INSERT-new-row branch located the By-Phase table with a canonical-column-order-only regex + a hardcoded positional row literal, so on a reordered header it silently inserted nothing — inconsistent with the now name-addressed UPDATE and SUM halves of the same function. Added insertTableRow (markdown-table seam sibling of updateTableCell/deleteTableRow: name-addressed, header-order-agnostic, EOL-preserving) and migrated the branch onto it, mapping By-Phase values by column NAME. Canonical-order output is byte-identical; a reordered header now inserts a correctly-mapped row; a pre-existing CRLF mixed-EOL splice glitch is incidentally fixed. Retired the now-dead byPhaseTablePattern const. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase-list checkbox flip via updateBullet seam Added updateBullet (markdown-sectionizer): a fence-aware, offset-tracked single-bullet write primitive (GFM 1–4-space marker tolerance) — the write counterpart to read-only iterateBullets. Migrated mutateMilestonePhase's phase-list checkbox flip (`- [ ] Phase N …` → `- [x] … (completed <date>)`) off its whole-slice regex onto it, same milestone-slice scope + clock seam. Byte-identical across simple / idempotent / metachar-title / double-space / CRLF scenarios. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): scope the Progress-ordinal renumber to ## Progress via seam phase remove's integer-renumber decremented Progress-table phase ordinals with a whole-document `content.replace(/(\|\s*)(\d+)(\.\s)/g, …)` — unscoped, so it also rewrote any `| N. …` cell in an unrelated/decoy table (same class as the batch-2 row-delete scoping bug). Migrated onto updateTableCell, scoped to the ## Progress section, decrementing each affected row's leading phase ordinal by column name. Byte-identical on canonical Progress tables + multi-row + decimal-sibling cases; a decoy `| 3. … |` row before ## Progress is now correctly left untouched. The sibling heading / checkbox-bullet / PLAN.md-filename / Depends-on-prose renumbers are not GFM-table mutations (outside ADR-2143's table/section mandate) — left as-is. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): review fixes — scope traceability write, restore Current Position H3-stop Adversarial review of the remediation (BLOCK verdict) — all 9 findings fixed: - F1 (BLOCKER): requirements mark-complete / phase complete flipped the checkbox but NOT the traceability row on the shipped template, because updateTableCell bound to the FIRST table (## Out of Scope, no Status column) instead of the ## Traceability table — the #2140 silent-divergence class, re-introduced by the seam migration and missed by tests (fixtures had Traceability first). Scoped the write + hasRow probe to the ## Traceability section slice (updateTraceability Cell helper) in milestone.cts + phase.cts. Failing-first tests on the Out-of-Scope-before-Traceability layout; the #2769 first-cell match preserved. - F2 (MAJOR): mutateCurrentPositionFirstTime restored to locateCurrentPosition (STOP_H2_PLUS) — collectSection's default H2-stop swallowed a level-3 subsection and the field regexes clobbered it (#2130 class). - F3/F8: Progress-ordinal renumber re-escapes via escapeCell + keys padding recovery by row index (was de-escaping `\|` and losing padding on dup values). - F4: insertTableRow escapes cell values internally. - F5: updateBullet accepts a tab after the marker (`[ \t]{1,4}`). - F7: resetSectionVerbatim consumes CRLF blank lines (byte-parity on CRLF). - F6/F9: corrected two misleading comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): data-loss + CRLF-session user-facing fixes (#2253) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): de-flake the G10 windsurf ReDoS-guard wall-clock assertion The G10 test asserted `elapsedMs < 1000` for a 200k-char payload — a wall-clock assertion (CLAUDE.md: never assert on wall-clock time) that flaked on a loaded node24 bench at ~1.1s. It was redundant: runHook's spawnSync `timeout: 10000` already SIGKILLs a catastrophic-backtracking hook, so the exit-0 assertion is the real ReDoS guard. Removed the timing assertion; kept exit-0 + documented the subprocess-timeout mechanism. Surfaced (not caused) by this branch's gsd-test runs loading the bench; unrelated to the markdown-parsing changes but fixed in place per the no-flaky-tests rule. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |