f0bb0787c9f550c19f7ccfbaf2cfb4db98218db3
99 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9ac0dfad58 |
chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared context-composer seam. Its success condition is that review-prompt output does not change, and the only authority on "did not change" is the behavior that shipped before the refactor. Capture that behavior now, while it is still the live implementation. 47 characterization cases, every `expected` value computed by executing the current implementation rather than hand-authored — the independence CONTRIBUTING.md "Fixture provenance (#2371)" asks for. A corpus is only worth what it can detect, so this one was validated by mutation rather than assumed. Five deliberate defects were injected and each must be caught by at least one case: - the note reserve deducted unconditionally instead of only under pressure - the pressure test relaxed from `>` to `>=` - a no-op head-shrink still setting the shrunk flag - the per-plan floor dropped from the proportional share - drop order reversed Two of those exposed real holes in the first cut of this corpus, and the cases that close them exist because of it: - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a floored plan group, and the 1024-char floor absorbed the entire trim, so the mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap, which makes the strict inequality observable as context kept vs omitted. - No case reached proportional-truncate at all — B6 and B7 both hard-failed the min-set pre-check first, leaving planTruncationPct at 0 across every case and the floor semantics entirely unexercised. Rebudgeted to 700 and 1100 so the min-set fits and the truncate step is actually reached; they now record 40.20% and 48.80%. The A4/A10 families sweep the pressure boundary from both sides, which is where this function has regressed before: CONTEXT.md's LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band, because the suite paired a trivially-fitting budget with a trivially-overflowing one and never sampled between them. A4 pins that nothing is trimmed from the cap down to 81 tokens under it; A10 pins that pressure fires at +1. Together with A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures. Two facts the corpus establishes that the design notes had wrong: - "" and null sections are NOT distinguished. applyBudget uses truthy checks throughout, so an empty-string section is treated as absent: not rendered, not dropped, never recorded in `omitted`. B13b pins this while the ladder is actively trimming, where only the non-empty `research` is dropped. - Sizing matters. B12/B13 were first written at a budget where both hard-failed the min-set check and returned "", so comparing them compared two empty strings and proved nothing. Committed as its own commit, ahead of the refactor, and regenerated against the pre-refactor implementation, so the oracle is demonstrably independent of the change it will adjudicate. Refs #2929 * refactor(#2929): extract the context-composer seam from prompt-budget Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer — per-runtime artifact emission — but it is walled inside the cross-AI review pipeline. Lift it into a shared seam so later phases can call it, without changing what the review pipeline emits. ADR-1671 specifies the composer as "priority + binary-search cutoff to a per-runtime budget". Read against the code it generalizes, that contract cannot express the thing being generalized. applyBudget is not a cutoff: it is a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor; instructions and roadmap are never touched at all. A cutoff composer sorts by priority and discards the tail — it has no way to say "shrink this one", "truncate that one but never below 1 KB each", or "these three are the only droppables, in this order". Building to the literal contract and routing prompt-budget through it would have silently changed review-prompt output, which is the one outcome this phase forbids. So shrink strategies are the core abstraction here, and cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. That is an elaboration of the ADR's intent, not a departure from it, and ADR-1671 is updated to say so. Three decisions worth stating: - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan of surviving fragments and never a string. assemblePrompt's rendering is prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and owning it in the composer would force emission to adopt prompt-shaped rendering. The split is what lets one seam serve both consumers. - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires for emission caps. The existing code converts a token budget to a character budget with a hardcoded `* 4`; that assumption is now an explicit `charsPerUnit` inverse, which is precisely what a byte unit needs in order to reuse this. - The entry point is `composeWithinBudget`, not `applyBudget`. That name already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter being an unrelated graph-edge budget. A third would make every symbol search in this repo ambiguous, and it already misresolves: preflight and impact queries for "applyBudget" return graphify's. Behavior is unchanged and proven so: all 47 characterization cases reproduce byte-identically, and the corpus is mutation-validated rather than merely green (see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and from eighteen mutable accumulators to two, both inside a helper copied verbatim. estimateTokens deliberately stays in prompt-budget and keeps its exact math: src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins plan estimates and recorded actuals to that same scale, so moving or changing it would silently break the calibration loop. Refs #2929 * docs(#2929): document the context-composer seam and amend ADR-1671 Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new domain modules), and a mutation-matrix entry for the new module. The ADR amendment is the substantive part. ADR-1671 specified the composer as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function the platform generalizes, so the ADR now records shrink strategies as the core abstraction with cutoff as one strategy among them, reserved for per-runtime emission in Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned against that contract and would otherwise be planned against a mechanism that does not work. The mutation-matrix entry is not bookkeeping. Stryker scores per module against a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the extracted code unmeasured while prompt-budget's own score floated free of the logic it used to cover. context-composer gets its own entry at the same floor. Refs #2929 * test(#2929): pin the effectiveBudget rounding mode in the parity corpus An isolated correctness review found a real blind spot: mutating `Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator happened to produce a whole number, so floor, round and ceil all agreed and the rounding mode was entirely unpinned by a corpus whose whole job is to pin observable behavior. Three cases fix that by straddling the .5 boundary: A11 95 * 0.90 = 85.5 floor 85, round 86 -> the two disagree A12 97 * 0.90 = 87.3 floor and round agree; ceil (88) does not A13 93 * 0.85 = 79.05 same guard at a non-multiple-of-10 margin, so the margin arithmetic is exercised and not just the budget A11 alone catches the round mutation; all three catch ceil. Regenerated against the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so the expanded corpus keeps the independence property the original capture had. The corpus is now mutation-validated against seven injected defects, every one caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink setting its flag, the truncate floor ignored, drop order reversed, and both rounding-mode changes. Refs #2929 * feat(#2929): flexReserve floors and the byte-stable isolate prefix Two of issue #2929's "Done when" items were unimplemented rather than deferred, and an isolated review flagged them alongside my own audit. Both are part of ADR-1671's composer contract, so shipping the seam without them would have left Phases 3-4 building against a contract that does not exist yet. flexReserve is a per-fragment floor in measure units that every strategy must respect, which is what makes it different from the pre-existing floorChars: that one is a chars-denominated detail of proportional-truncate alone and is retained unchanged. A floored fragment is never dropped, is never head-shrunk below its floor, and raises its own proportional cap. A fragment already smaller than its floor is untouchable outright. Metadata gains `floored`, listing the ids whose floor actually prevented a trim — a guarantee no caller can observe is a guarantee no test can hold you to. isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed, never dropped, but still counted, because a prefix excluded from accounting would silently under-count real context. Metadata gains `isolatePrefix` so a caller can hash or assert on the exact bytes. Declaring an isolate fragment after a non-isolate one throws: a prefix that is not at the front is not a prefix, and accepting it would make the cross-runtime stability claim meaningless. Adds tests/context-composer.test.cjs for the exact new semantics and tests/context-composer.property.test.cjs for the five invariants, including the budget-monotonicity property the issue names explicitly. Both are registered in the mutation matrix, since coverage does not migrate with relocated code. prompt-budget uses neither feature, and its output is unchanged: all 50 corpus cases still reproduce byte-identically. Refs #2929 * chore(#2929): allowlist the prompt-budget parity suite The parity corpus needs its own test file and that makes prompt-budget a three-file module against a limit of two. The lint offers consolidation or an allowlist entry with justification; the entry is the right call here. Consolidation would mean folding the characterization suite into prompt-budget.test.cjs, which is the one thing that should not happen to it. The parity suite is a distinct concern with a distinct lifecycle: it is generated rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own scoring target, and its failure means something categorically different from a unit-test failure — not "this behavior is wrong" but "observable output moved". Burying it inside a general unit file would obscure exactly that signal. The allowlist is an identity ratchet, so this entry pins today's three exact filenames: adding a fourth still fails, and dropping back to two requires removing the entry. Refs #2929 * fix(#2929): register the new module with two gates it was missing The remote matrix caught three defects that no local check could, because the local runner is blocked in this repo and these suites had therefore never executed. Eight failures, identical on node22 and node24, so nothing environment-shaped. Two are the new-module ripple. A net-new src/*.cts lands in six places and this change had reached four of them — .gitignore, INVENTORY, the manifest, and the CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet baseline (a deliberate review-visible mirror of the matrix floors, which every COVERED module must carry). Both are now registered, the ratchet at the same floor of 66 the matrix declares. The third was a test asserting an outcome it had made impossible. It set budget:1 alongside a 400-char required fragment, so the group budget came out at -99 and the proportional-truncate step was skipped entirely — the deliberate "non-positive group budget is skipped, never clamped" rule inherited from the original ladder. Nothing was trimmed, and the test then asserted a truncation. Rebudgeted so the step actually runs, with the arithmetic written out in a comment so the next reader does not have to re-derive why 120 rather than 80. Fixing that surfaced a genuine bug in the composer. `floored` is documented as recording fragments whose flexReserve prevented a trim that would otherwise have happened, but the push sat in the else-branch of "content did not change", so it only fired when nothing was trimmed at all. A fragment truncated to a reserve-raised cap has also had a trim prevented — 40 characters' worth in the test above — and was silently absent from the field that exists to make the guarantee observable. The condition was already right; it was in the wrong branch. Now recorded on both paths: a drop prevented outright, and a truncation capped higher than the share alone would have allowed. Parity is unaffected — prompt-budget never sets flexReserve, so the branch is unreachable from every corpus path, and all 50 cases still match. Refs #2929 * chore(#2929): backfill changeset PR number (#2958) * chore(#2929): correct the corpus case count in the changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
c19d3d7bda |
chore(#2453): resolve the uniform-router-param conflict and clear the warning floor (#2489)
npx eslint . reported 0 errors, 53 warnings. The floor eroded the signal: a
genuinely new warning had to be spotted against noise, so 'lint is clean' was not
usable as a check. This takes it to zero.
Fifty of the 53 were one category in gsd-core/bin/gsd-tools.cjs — all unused
ARGUMENTS (error x30, cwd x10, raw x7, args x3), never unused variables. They are
the Command Routing Hub's uniform handler signature, function routeX({ args, cwd,
raw, error }), declared identically whether or not a handler uses all four
members. argsIgnorePattern: '^_' is structurally in conflict with that convention:
satisfying it would mean _-prefixing ~50 parameters and making the signature
non-uniform across the table. Option 1 of #2453: disable args checking for that
file only, keeping varsIgnorePattern intact so genuinely dead variables (the #2379
class) still surface — verified by injecting an unused variable, which still warns.
This is the config decision #732 explicitly deferred ('Severities stay warn (no
config change in this pass)').
The remaining three predate the issue's count of 51 and are real defects, not
suppressions:
- tests/workflow-compat.test.cjs — the step-9 lookahead used (?=\*\*10\.|\z).
\z is a Perl/Ruby end-of-input anchor with NO meaning in JavaScript; it matched
a literal 'z', so the lazy span silently stopped at the first z whenever **10.
was absent, truncating the captured step and letting the assertion pass against
a partial block. Corrected to $.
- tests/debugger-prevention.test.cjs — an unused RegExp built one line above the
one actually used. Removed.
- tests/installer-migrations.test.cjs — try/finally inside a test body, which
CONTRIBUTING prohibits ('verbose, masks test failures, not an approved
pattern'); the unused t was the symptom. Converted to t.after().
Closes #2453
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2cbf186420 |
chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112, #2118) were already fixed tactically on next; this introduces the reusable structural contracts and rewires the primary #2140 site onto them. - src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the per-surface write-set (`WriteOutcome {surface, applied, requirement?}`, `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports `Result` from here (single source; distinct from command-routing-hub's Result). - requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT, per-surface write-set; `write_set_complete` is true only if every surface of every requirement applied — structurally forbidding the #2140 OR-into-one-flag masking, including across a multi-ID batch (adversarial-review regression). Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed). - deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow; RoadmapProgress return contract unchanged. - commit --files (#2112) and milestone complete --dry-run (#2118) left as-is (single-surface commit / pre-mutation preview — not genuine multi-surface writes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2244): backfill changeset PR number (#2251) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d49ac81306 |
chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a canonical seam and migrate the pilot reader. - Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed {columns, rows} addressed by column NAME; ragged rows are typed parse errors, not silent), a single-source TABLE_SCHEMAS registry (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with variants under one id), matchTableSchema, and findTableBySchema. Result<T> is scoped to this seam (distinct from the dispatch Result). - Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the position-anchored regex to name-based resolution via the seam — fixes #2137 (the 5-column milestone-grouped Progress table previously returned all-null). - Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route fast.md's log_to_state through it, retiring the inline `awk NF-2` column arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped (| and newlines) and the STATE.md read-modify-write is atomic under readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230). - Writer/reader/template parity test guards TABLE_SCHEMAS against drift (ADR-2143 §3 Generative-Fix-Divergence). Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md + INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md. Behaviour-preserving for the canonical 4-column Progress table; the named bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2242): backfill changeset PR number (#2248) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): escape backslash before pipe in markdown-table cell escaping CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl. literal backslashes) round-trip exactly. Added backslash round-trip tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself "pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via a new seam helper findTableWithColumns (first table whose header is a superset of Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while staying seam-based and preserving its `## Progress` scoping (#2012/#1445). Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale state.test.cjs assertion that predated the Phase-1 migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a4d8eb4a78 |
chore(#1867): eslint-ignore generated ui-consideration-probe.cjs (SHIP-01)
ci-preflight (full test:unit) caught a Phase-1 registration gap: the
tsc-generated gsd-core/bin/lib/ui-consideration-probe.cjs was not in the
eslint.config.mjs global ignore list, failing 551-eslint-bin-lib-coverage
("each bin/lib/*.cjs is linted xor ignored according to migration state",
ADR-457/#537). Generated artifacts are ignored, not linted — added it next to
the sibling probe adapters (edge-probe/probe-core/prohibition-enforcement).
Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
|
||
|
|
6e7e3111fb |
test(#2126): fix #1259 real-eslint CPU starvation — lint a non-type-aware .cjs clean target
The prohibition-enforcement real-runner tests linted src/clock.cts (a .cts) as their clean target. Under eslint.config.mjs's type-aware block for src/**/*.cts (recommendedTypeChecked + parserOptions.project: tsconfig.build.json), each eslint spawn loaded the WHOLE tsconfig.build.json program (~2s, CPU-heavy). The real-runner tests spawn eslint repeatedly; under --test-concurrency those full-program type-checks oversubscribed the bench CPU and blew the 60s subprocess bound -> fail-closed (intermittent, load-dependent — passed 24241/24241 in an earlier run, failed here). Root fix (not a retry/timeout bandaid; measured projectService = no faster since a single-file .cts lint still loads type info): add tests/_ff_lint_clean.cjs, a KNOWN-CLEAN lint-scoped .cjs companion to _ff_lint_violation.cjs, with a flat-config block enabling local/no-source-grep so the clean pass stays non-vacuous. Repoint the 6 src/clock.cts real-runner usages (5 targets + the FF-02 toothless violationFixture) at it. Each spawn is now ~0.8s non-type-aware (no whole-program load) — starvation removed. All 6 tests' semantics verified in-process (SF-01 greens; toothless/fail-closed stay unverified); full-repo `eslint .` green. Refs #2126, #1259 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
015c3a7fda |
fix(#2071): extract install-time effort resolvers so effort sync stops requiring the un-shipped bin/install.js
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/) with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts) required the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never present there. So `effort` config changes silently never reached installed agents without a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring "runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../" anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated). Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig + resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers) out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four symbols back from the new module (it still calls them + re-exports them), so there is one source of truth and no duplication/drift. The lazy manifest read is repointed from the package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`). Scope note: this is one of four instances of the anti-pattern; the other three are already shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept out of this fix. Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp home (the golden-parity helper) and runs the issue's exact repro (`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for bin/install.js. Fail-first verified: against pristine next the same test throws `Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly. New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md + INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib (excluded from golden parity), so no golden fixtures change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |
||
|
|
603593d41d |
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest defaults to WATCH mode in an interactive TTY — exactly where a user runs `gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never exited and the orchestrator waited indefinitely. Recovery needed the user to manually prompt "something blocking?". Fix — one shared helper + a bounded, surfacing timeout on the test-command gates: - New pure module src/normalize-test-command.cts + `gsd-tools query normalize-test-command` verb: rewrites a resolved command to a best-effort one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script whose package.json runner is watch-vitest → `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged — never double-flagged). Named `normalize-test-command` (not `test-*`) so the file does not match node --test's default `test-*` discovery glob. - The three gates that HUNG or silently-continued route through that ONE helper and bound execution with `timeout $(config-get workflow.test_gate_timeout)` (new config key, default 600s): the regression gate (extracted to execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen — it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124. - verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint on 124, staying under its frozen 40960-byte tier cap. Security hardening (review): the normalizer only rewrites a runner named as a standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never mangled), is length-capped and uses only linear-time split-based scanning (no super-linear backtracking on an adversarial `workflow.test_command`), and reads package.json only when it is a regular file (never blocks on a FIFO via `--dir`). Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores, inventory manifest/index. All 16 golden-install-parity fixtures + workflow size baseline regenerated for the changed shipped files; bin/lib is excluded from the parity manifest. Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route through the shared helper + configured timeout + exit-124 watch-mode hint; verify-phase asserted as normalize-only/already-bounded). tests/execute-phase-active-flags.test.cjs repointed at the extracted step; tests/planner-language-regression.test.cjs allowlist comment updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
080bacdb4b | Merge branch 'next' into codex/gsd-onboard | ||
|
|
e3262d94d3 |
feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool (/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one runtime gate. - Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend (gate ladder: enabled -> Claude -> backend != inline -> nested+background host -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor + worktree, files_modified overlap -> separate stages, resumeFromRunId, budget). All interpolated identifiers validated script-safe; briefs JSON-quoted. - claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow) for orchestrator invocation. - Two gated loop contributions at wired points (execute:wave:post, plan:post); federated config keys (enabled/execution_backend/min_agent_sdk_version). - ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc. On any runtime lacking the Workflow tool, behaviour is byte-identical to today. closes #1143 |
||
|
|
7ef834cabc | feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them | ||
|
|
a5298c1fc0 | refactor: project onboard routing in init | ||
|
|
8de2ff9121 |
feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator Third-party capability gates declared via check.predicate were rendered for display but never evaluated (only built-in check.query gates fired; the security capability's gate worked solely via a hard-coded ship.md branch). Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts) that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a bounded sh -c command at the project root (via shell-command-projection.execTool), inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed. Wire a 'check predicate' subcommand into check-command-router.cts and extend the three generic workflow gate-dispatch sites (execute:wave:post, execute:post, plan:post) to route check.predicate gates to the new evaluator. The two-step gate contract (command-failure => onError; block => halt) is unchanged. - src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible - src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags - docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md - tests: 38 unit + integration tests (exit mapping, timeout, interpolation, property-based bijection, malformed-predicate fail-closed, real subprocess e2e) Closes #2008 * docs(#2008): backfill changeset pr number 2011 |
||
|
|
5ae4ea4c84 |
feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib. |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1388e7a362 |
feat(#1683): published Host-Integration SDK surface + smoke test — Slice 2 (#1939)
The SDK entry (src/host-integration-sdk.cts) is the single PUBLIC surface a host-plugin author imports: the negotiated schema + classification, the five adapters (declarative/imperative/model/hook/state), and the serialized handshake. Frozen so the public shape cannot be mutated. Everything else in gsd-core stays internal. tests/sdk-smoke.test.cjs imports ONLY from the SDK entry and builds a third-party host-plugin end-to-end (compose adapters + handshake + classify) — proving an external author can wire a host without gsd-core internals (#1683 AC). ESLint-ignore + inventory manifest kept in sync for the new tsc-emitted module. |
||
|
|
47ecf997d6 |
feat(#1683): serialized capability-exchange handshake — Phase 6 Slice 1 (#1937)
* feat(#1683): serialized capability-exchange handshake — Phase 6 Slice 1 The out-of-process wire form of Phase 1's negotiateHostCapabilities: SDK hosts (pi, VS Code) that cannot share object refs exchange a JSON capability set over a wire boundary (MCP-style initialize). src/handshake-serialized.cts provides buildHandshakeRequest (host side) + handleHandshakeRequest (engine side, delegates to negotiateHostCapabilities), both JSON-round-trip-safe. CONSISTENCY is the contract: a serialized request yields the same NegotiationResult as the in-process call for the same axes (asserted in the test). Pure + additive; the companion MCP server or an SDK host binds it to a real transport. * fix(#1683): ignore tsc-emitted handshake-serialized.cjs (ADR-457) + refresh inventory manifest * fix(#1683): correct inventory manifest — drop junk smart-entry.cjs, add handshake-serialized.cjs The local gen-inventory-manifest scan captured a stale untracked smart-entry.cjs (not on next, no src/*.cts) and missed the freshly-built handshake-serialized.cjs. Aligned to the canonical clean-build report: + handshake-serialized.cjs, - smart-entry.cjs. * fix(#1683): type the JSON wire round-trips (no-unsafe-return) |
||
|
|
41193a44bd |
feat(#1681): ADR-1239 Phase C-2 — gsd-mcp-server bin entry + lifecycle test [slice 3b] (#1810)
* feat(#1681): ADR-1239 Phase C-2 — gsd-mcp-server bin entry + lifecycle test [slice 3b] Phase 4 slice 3b (closes #1681). The companion MCP server bin entry so any MCP-consuming host connects via 'npx gsd-mcp-server' (or its bin on PATH) and gets GSD command (point 1) + state IO (point 5) with no bespoke plugin. - gsd-core/bin/gsd-mcp-server.cjs: #!/usr/bin/env node shim requiring ./lib/mcp-server.cjs + runServer({stdin, stdout}); non-zero exit on fatal error (justified n/no-process-exit disable). Mirrors gsd-tools.cjs. - package.json: add 'gsd-mcp-server' bin entry. - tests/gsd-mcp-server-bin.test.cjs: 3 process-lifecycle tests — initialize + tools/list round-trip + clean exit, malformed-line -> parse error + server keeps running, empty stdin -> clean exit. Synchronous spawnSync (bounded; server exits on stdin EOF, no orphan). Phase 4 trust-gate (#1806) + loader wiring (#1808) + server module (#1809) + this bin/lifecycle slice = all of #1681's deliverables. Concrete host binding -> Phase 5 (#1682). npm-integrity + eslint + security + inventory all clean. * docs(#1681)+chore(changeset): how-to for the companion MCP server + Added fragment docs/how-to/connect-gsd-mcp-server.md — Diataxis how-to guide for connecting any MCP-capable host to gsd-mcp-server: goal-oriented flow (add config → restart → verify), real-world per-host conditionals, troubleshooting, and a trimmed reference table. Explanation/reference linked out (ADR-1239, capability-trust- model) per Diataxis boundary rules rather than mixed in. .changeset/humble-seals-rest.md — type: Added (first user-reachable surface of the epic: a new bin command). The how-to doc satisfies the docs-required gate. * fix(#1681): move gsd-mcp-server shim to top-level bin/ (out of the runtime-copied tree) The shim at gsd-core/bin/gsd-mcp-server.cjs was inside the tree the installer copies into every runtime config dir, so it leaked into all 16 runtimes and broke golden-install-parity. The MCP server is a PACKAGE bin the host spawns (npx gsd-mcp-server), not a per-runtime artifact — so it belongs at top-level bin/ alongside install.js (which is also never copied into a runtime config). - gsd-core/bin/gsd-mcp-server.cjs -> bin/gsd-mcp-server.js (require path now ../gsd-core/bin/lib/mcp-server.cjs). - package.json: bin entry -> bin/gsd-mcp-server.js. - tests/gsd-mcp-server-bin.test.cjs: SHIM path updated. - eslint.config.mjs: add bin/gsd-mcp-server.js to the bin/install.js block (drops the n/no-process-exit disable — the n plugin isn't loaded for that block, so the disable referenced an undefined rule). golden-install-parity 16/16 restored; lifecycle + unit tests green; eslint 0; lint:ci all ok. |
||
|
|
2b38356275 |
feat(#1681): ADR-1239 Phase C-2 — companion MCP server module (points 1 + 5) [slice 3a] (#1809)
* feat(#1681): ADR-1239 Phase C-2 — companion MCP server module (points 1 + 5) [slice 3a] Phase 4 slice 3a. A minimal, dependency-free stdio JSON-RPC 2.0 server exposing two of the six interface points so any MCP-consuming host (Claude/Codex/OpenCode/ VS Code/Gemini/Cursor/Cline/Hermes) can drive GSD with no bespoke plugin: - point 1 (command): tool gsd_invoke_command -> createHub/dispatch. - point 5 (state IO): tools gsd_read_state / gsd_write_state -> the Phase 3 stateIO seam (filesystem default). - src/mcp-server.cts: handleMessage(request, ctx) pure JSON-RPC handler (initialize / tools/list / tools/call) + runServer({input, output}) thin line-delimited-JSON loop over injectable streams. 3 tools wired to the existing engine surfaces. NO new dependency (hand-rolled JSON-RPC; the repo ships only claude-agent-sdk + ws — an MCP SDK is a separate packaging call). - tests/gsd-mcp-server.test.cjs: 9 tests (initialize, tools/list, state read/write round-trip, command dispatch, unknown tool / missing name / unknown method / notification / parse error, injectable-stream round-trip). Bin entry / packaging / manifest-version-sync / process-lifecycle docs -> slice 3b. This slice ships the importable, tested server surface a host (or the bin shim) drives. Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST + injection-scan audit. All clean locally (9 tests + security 15/15 + inventory + eslint 0 problems). * chore(changeset): add Changed fragment for companion MCP server module (#1681) |
||
|
|
21b81ea068 |
feat(#1681): ADR-1239 Phase C-2 — external-descriptor trust gate (configHome confinement) [slice 1] (#1806)
* feat(#1681): ADR-1239 Phase C-2 — external-descriptor trust gate (configHome confinement) [slice 1] Phase 4 slice 1. Load-time, fail-closed configHome write-confinement for installed third-party host-plugin descriptors — defense-in-depth on top of the existing opt-in/schema/consent/first-party-wins loader gates + Phase 2's install-time assertDestWithinConfigHome (#1679 AC3). - src/external-descriptor-trust.cts: isPathConfined(target, root) pure cross-platform containment primitive + assertDescriptorConfined(descriptor, configHome) — walks runtime.artifactLayout global/local destSubpaths, throws fail-closed (naming descriptor + path) on the first escape. Rejects ../escape + absolute-outside-root. Missing layout / invalid entries skipped. - tests/external-descriptor-confinement.test.cjs: 7 tests (containment primitive, benign passes, global/local/absolute escapes rejected, missing layout, invalid entries). Load-time twin of Phase 2's install-time gate — rejects malformed/escaping descriptors BEFORE consent even matters. NOT wired into loadRegistry yet (slice 2, 4 callers, medium blast radius); this ships the reusable gate + tests. Not the ADR-1577 prompt-injection breaker (separate concern, shared word trust). Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST + injection-scan audit. All clean locally (7 tests + security + inventory + eslint 0 problems). * chore(changeset): add Changed fragment for external-descriptor trust gate (#1681) |
||
|
|
abd21b2968 |
feat(#1680): ADR-1239 Phase C-1 — hook-bus + stateIO seams [AC4] (#1805)
* feat(#1680): ADR-1239 Phase C-1 — hook-bus + stateIO seams [AC4] Phase 3 slice 4 (AC4, final #1680 slice). The last two adapter seams behind the negotiated hookBus/stateIO axes: - src/hook-bus.cts: createHookBus({bus}, {hostEmit?}) -> host/engine/none. engine = in-process pub/sub (handler errors isolated); host = host-owned, fail-closed emit until a host emitter is bound; none = silent no-op (degrade to rule-text). PORTABLE_EVENT_FLOOR = SessionStart/PreToolUse/PostToolUse/ Stop/SessionEnd (the claude dialect all hook hosts share). - src/state-io.cts: createStateIO({io}, {backend?}) -> filesystem (today's behavior — straight fs) / sandboxed-storage / session-log-append (fail-closed seams until a host backend is bound). Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST entries for both new .cjs; injection-scan 'act as' substring audit; unused-import check. All clean locally (11 tests + security 15/15 + inventory + eslint 0 problems). Phase 3 (#1680) seam layer now complete. Concrete host binding -> Phase 5 (#1682, D15/D18). * chore(changeset): add Changed fragment for hook-bus + stateIO seams (#1680) |
||
|
|
b152f7e64c |
feat(#1680): ADR-1239 Phase C-1 — model adapter seam (passive + active) [AC3] (#1804)
* feat(#1680): ADR-1239 Phase C-1 — model adapter seam (passive + active) [AC3] Phase 3 slice 3 (AC3). Two model-layer adapters selected by the negotiated modelMode axis (host-integration.cts): - passive: formalizes today's tier routing from src/model-resolver.cts — resolveModel delegates straight to resolveModelForTier (byte-for-behavior). This is the CLI runtimes (claude/gemini/codex/opencode/cursor/...): GSD injects prompts / a per-agent model field. - active: a host-supplied sendRequest seam (VS Code vscode.lm / pi providers) — GSD calls the model through the host. Ships as a fail-closed seam (throws until a provider is bound); Phase 5 wires a concrete provider. createModelAdapter({modelMode}, {sendRequest?}) — factory gating throws on invalid mode. Proactive CI gates applied: ADR-457 eslint ignores entry + INVENTORY-MANIFEST cli_modules entry + comment-wording audited for the injection-scan substring trap. * chore(changeset): add Changed fragment for model adapter seam (#1680) |
||
|
|
da368311ea |
feat(#1680): ADR-1239 Phase C-1 — imperative embedding adapter (composes loadRegistry) [AC2] (#1803)
* feat(#1680): ADR-1239 Phase C-1 — imperative embedding adapter (composes loadRegistry) [AC2] Phase 3 slice 2 (AC2). The engine-as-library path: createImperativeAdapter composes loadRegistry({includeInstalled:true}) — first-party-wins + consent + fail-closed gates, identical trust semantics to the CLI — and binds the engine surface behind the SAME HostIntegrationInterface the declarative adapter (AC1) satisfies, plus a registry accessor for the composed capability set. - src/adapter-imperative.cts: createImperativeAdapter({runtime}, {loadOptions}) → ImperativeAdapter (kind:'imperative' + .registry + install/uninstall delegating to install-engine). Thin: delegates the loop, does not reimplement. - tests/adapter-imperative.test.cjs: kind (16 runtimes), registry composition (loadRegistry called with includeInstalled:true), loadOptions pass-through, install/uninstall delegation, fail-closed construction. - eslint.config.mjs + docs/INVENTORY-MANIFEST.json: ADR-457 ignores entry + cli_modules entry for the new emitted .cjs (the two drift gates that bit AC1, applied proactively here). Concrete host binding (OpenCode/VS Code/pi) deferred to Phase 5 (#1682). * test: remove dead readStateMd helper from bug-1760 test readStateMd was defined but never called (writeStateMd is the only state-md helper this test uses). Clears the lone no-unused-vars warning so the repo lints fully clean (0 problems). No behavior change — test still passes 2/2. * chore(changeset): add Changed fragment for imperative embedding adapter (#1680) * fix(adapter-imperative): reword comment to avoid injection-scan substring match The prompt-injection scan regex 'act\s+as\s+(?:a|an|the)' was matching the 'act as the' substring inside 'contract as the declarative adapter' (contrACT AS THE). Reword 'contract as' -> 'shape as' — no 'act' substring, identical meaning. Clears the 'lib source files are clean' + 'codebase prompt injection scan' security-gate failures. |
||
|
|
c642ed0ec5 |
feat(#1680): ADR-1239 Phase C-1 — declarative embedding adapter + minimal HostIntegrationInterface [AC1] (#1802)
* feat(#1680): ADR-1239 Phase C-1 — declarative embedding adapter + minimal HostIntegrationInterface [AC1] Phase 3 slice 1 (AC1). Names + bounds today's projection path behind the common HostIntegrationInterface that both declarative + imperative adapters will satisfy. - src/embedding-adapter.cts: minimal HostIntegrationInterface (kind + runtime + install/uninstall) + ADAPTER_KINDS. The full 6-point binding surface (command/dispatch/model/hooks/state/artifact) is DEFERRED until the imperative adapter (AC2) fixes the shape — ADR-1239 lists the wire-shape as an open question; freezing it now risks rework across Phases 3-6. - src/adapter-declarative.cts: createDeclarativeAdapter({runtime}) factory. Delegates in-process to install-engine installRuntimeArtifacts / uninstallRuntimeArtifacts (the SAME engine functions bin/install.js uses), so output is byte-identical to today's install (gated by golden-install-parity). Module-ref call style = monkeypatch-friendly for tests. Lossy by design: projects files, does not drive the loop (that's the imperative adapter, AC2). - tests/adapter-declarative-equivalence.test.cjs: kind classification (all 16 runtimes), install/uninstall delegation with exact args (the byte-identity link), fail-closed construction (missing/invalid runtime throws). Purely additive — no install.js/install-engine changes. Unblocks AC2 (imperative adapter) + AC3/AC4 (model/hook/state seams) as follow-up slices. * chore(changeset): add Changed fragment for declarative embedding adapter (#1680) * fix(lint): ignore tsc-emitted embedding-adapter/adapter-declarative .cjs (ADR-457) The new src/embedding-adapter.cts + src/adapter-declarative.cts modules' emitted gsd-core/bin/lib/*.cjs artifacts must join the ADR-457 ignores list (lint the src/*.cts source, not the emitted .cjs). Without this, the .cjs is linted under js.recommended where @typescript-eslint/no-require-imports is undefined, so the verbatim-copied eslint-disable directive surfaces as 'Definition for rule not found' — failing the lint-tests CI job. Also restores the clean line-level disable in adapter-declarative.cts (valid in the .cts source context where the rule IS defined). * fix(docs): add adapter-declarative + embedding-adapter to INVENTORY-MANIFEST cli_modules The new ADR-1239 Phase C-1 modules' built .cjs artifacts must be registered in docs/INVENTORY-MANIFEST.json's cli_modules array or the 'docs/INVENTORY-MANIFEST.json matches the filesystem' drift test fails on CI. Mirrors the existing sorted entries. |
||
|
|
744bb7aaee |
refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase (#1775)
* refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase Lands the Phase 1 substrate per ADR-1769: - src/state-transition.cts (new Module): - Field-classification table (FieldClass enum + FIELD_CLASSIFICATION rows) - STATE_MD_SECTIONS constants block - Pure transitionCore(content, intent, deps) dispatch - beginPhase intent implementation (first-time + #3127 resume paths) - src/state.cts:cmdStateBeginPhase — collapses ~190 lines to a thin dispatch onto transitionCore via readModifyWriteStateMd. The lock, no-op write guard, and #1230 post-sync delta heuristic stay in the RMW seam; the body-mutation policy moves to transitionCore. - tests/state-transition.test.cjs (24 tests): - Substrate invariants (table enum, section constants) - Characterization: 6 first-time body field updates - Characterization: 5 #3127 idempotency-guard resume behaviors - Characterization: 3 Current Position section mutations - Characterization: Current focus body text line (#1104) - Property (RULESET.TESTS.property-based-testing): beginPhase status propagation + FIELD_CLASSIFICATION own-property contract - Resume Current Position mutation (preserves Plan/Phase/Status) No external behavior change. Full state.test.cjs regression (177 tests) plus bug-3127/#3242/#905/#948 pass. Property tests surfaced two pre-existing quirks (state-document.cjs greedy \s* on whitespace-only field values; Object.prototype method leakage on FIELD_CLASSIFICATION lookups for strings like 'toString') — documented in test comments; fix-out-of-scope for Phase 1. Closes #1771 * refactor(#1771): ADR-1769 Phase 1 codex review corrections Addresses 3 blocking findings from codex gpt-5.5/high review: 1. FIELD_CLASSIFICATION shape (state-transition.cts): - Was flat FieldClass enum (collapsed source + preservation) - Now two-column {source, preservation} rows per ADR-1769 §4 - Added missing fields verified via Memtrace against buildStateFrontmatter (state.cts:1633-1653): gsd_state_version, last_updated, last_activity_desc, progress.{total_phases, completed_phases, total_plans, completed_plans, percent} - Field 2-7 preservation dispatch can now consult the table 2. Prototype-pollution hardening (state-transition.cts): - Table is now Object.freeze(Object.assign(Object.create(null), {...})) - getFieldClassification() helper uses Object.hasOwn; returns null for inherited prototype methods (toString/valueOf/__proto__) - Old code: FIELD_CLASSIFICATION['toString'] returned the function 3. STATE_MD_SECTIONS aligned to canonical template: - Verified against gsd-core/templates/state.md via Memtrace - Was: 8 entries including non-template sections (## Session, ## Decisions, ## Operator Next Steps, ## Session Log, ## Roadmap Evolution) - Now: 6 canonical top-level sections (## Project Reference, ## Current Position, ## Performance Metrics, ## Accumulated Context, ## Deferred Items, ## Session Continuity) Also: beginPhase now consults getFieldClassification() per touched field (codex finding: 'table not consulted by transitionCore'). Unknown fields raise immediately — adding a field without a table row is caught at runtime. Repo-hygiene catches from gsd-test (not node --test, which missed these): - gsd-core/bin/lib/state-transition.cjs added to eslint.config.mjs ignore list (ADR-457 tsc-generated) - docs/INVENTORY-MANIFEST.json regenerated via node scripts/gen-inventory-manifest.cjs --write Property test for Object.prototype leakage tightened to verify getFieldClassification() returns null for toString/valueOf/__proto__. Ref #1771 * fix(#1771): cast Object.create(null) to satisfy @typescript-eslint/no-unsafe-assignment ESLint CI failed on src/state-transition.cts:72:14 — Object.create(null) returns `any`, which leaked through Object.assign to the typed `FIELD_CLASSIFICATION` declaration. Adding an explicit cast to `Record<string, FieldClassification>` eliminates the unsafe-assignment while preserving the null-prototype protection codex review recommended. gsd-test: 21888/21888 PASS. * fix(#1771): add 'see #1771' to allow-test-rule exemption per ADR-456 CI lint-allow-test-rule-refs failed: 'New allow-test-rule exemption without an issue ref — add `see #NNN` per ADR-456'. Updated comment on tests/state-transition.test.cjs to reference the Phase 1 issue. * fix(#1771): remove unnecessary allow-test-rule exemption The exemption was added speculatively. The test file does not use readFileSync + .includes()/.match()/.startsWith() on source content — it calls transitionCore() with string literals and verifies results via stateExtractField() and array .includes() on the updated[] array. No exemption needed. |
||
|
|
4ced0a64cc |
feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint * chore(#1561): backfill changeset PR number (#1767) --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
e075a41c86 |
feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD (#1755)
* feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD Addresses #1754 (approved-enhancement). Detects when the running gsd-tools.cjs is outside the project root while a project-local install exists — the shadowing scenario from #1748 where a stale global canary CLI (retired @gsd-build/sdk) silently overrides project-local GSD. Implementation (Node CLI entry-point, not shell snippet — avoids bloating 93 workflow files past their size caps): - src/cli-skew-check.cts: pure function checkCliSkew({resolvedPath, projectRoot, projectLocalExists}) → string|null. Compares paths via path.relative; returns a warning when the resolved CLI is outside the project root AND a project-local install exists. Includes @gsd-build/sdk removal hint when the path matches. No I/O (pure), no gsd-sdk literal (avoids bug-2801 lint). - gsd-core/bin/gsd-tools.cjs: wired at startup via the existing findProjectRoot resolver. Non-blocking (try/catch; advisory stderr warning, never gates). - eslint.config.mjs: registers the new ADR-457 generated artifact in the ignores. - tests: 6-case suite (skew/no-skew/legacy/normalization); all green. - Golden fixtures regenerated (UPDATE_GOLDEN=1) for the new compiled artifact. - docs/how-to/update-gsd.md: Diátaxis reference note for the skew warning. Full suite: 3354 pass, 0 regressions (1 pre-existing local AGENTS.md failure). lint:ci green. Closes #1754 * chore(#1754): backfill changeset pr placeholder * chore(#1754): regenerate INVENTORY-MANIFEST for the new cli-skew-check source module --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
871621c3c8 |
feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) (#1742)
* feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) ADR-1703 Phase 6 of the cross-platform portability epic (#1702). Adds the second production-code portability AST rule + the ADR-mandated glob expansion to bin/install.js and scripts/build-hooks.js. - eslint-rules/require-fs-op-fallback.cjs: flags an unguarded fs.rename / fs.renameSync (the atomic-publish primitive named first in DEFECT.WINDOWS-FS-OPS.symptom) that is NOT inside a try/catch whose handler references a transient errno ('EPERM'/'EBUSY'/'EACCES' or a *RETRY_ERRNOS set) AND NOT behind a Windows platform guard. A catch that silently swallows or cleans-up-and-rethrows without an errno check does NOT satisfy the defect's 'never silently swallow' clause. copyFile/unlink are deliberately not flagged (they are the fallback primitives per the defect's own fix-forward). Scope narrowed to rename per Phase 5's precision discipline; documented on #1740. - src/shell-command-projection.cts: export retryRenameSync(from, to) — the drop-in bounded-retry helper over the existing atomicRenameWithRetry. - 27 bare fs.renameSync sites across 11 modules routed through retryRenameSync (capability-lifecycle/lock/source, installer-migrations, milestone, phase, planning-workspace, roadmap-upgrade, runtime-hooks-surface, state, workstream). Idempotent on POSIX; resilient to AV/indexer transient locks on Windows. - eslint.config.mjs: register rule at error on src/**/*.cts; new focused portability-rules block covering bin/install.js + scripts/build-hooks.js (ADR-1703 L124-126 glob expansion — both files are compliant: zero rename violations). - tests: 15-case RuleTester suite; portability-rule-disable-ban extended (PROTECTED_RULES + scans bin/install.js/build-hooks.js with shebang handling); ci-test-scope portability-lint selection rule. - CONTEXT.md DEFECT.WINDOWS-FS-OPS predicate rewritten to point at the rule; docs/contributing/cross-platform-portability-rules.md reference + how-to. Closes #1740 * chore(#1740): backfill changeset pr:1742 * fix(#1740): tighten require-fs-op-fallback precision (codex review HIGH-1/HIGH-2) Addresses two false-negative findings from the codex (gpt-5.5/high) adversarial review of PR #1742: HIGH-1 — a catch that REFERENCES a transient errno but only rethrows (no retry/fallback) was marked compliant. The DEFECT.WINDOWS-FS-OPS fix-forward requires retry, not just recognition. Fix: catchHandlerHasRetrySignal now requires a loop `continue` backedge OR a `return <call>` delegation; a bare rethrow is flagged. The misleading `/* retry logic */` valid test is replaced with a real retry loop, and the rethrow-only shape is added as invalid. HIGH-2 — the nested-try ancestor walk treated an OUTER errno-catch as protecting the rename even when an INNER catch intercepted/swallowed the error (the outer catch is unreachable). Fix: isInsideTransientErrnoTryCatch now stops at the NEAREST enclosing TryStatement WITH A CATCH HANDLER whose block contains the rename (try-finally is skipped — it doesn't catch); outer catches are no longer consulted. The unsound nested-try valid test is converted to invalid, and a try-finally-skipped valid case is added. Verified: 17 RuleTester cases pass; zero new production violations (the 27 fixed sites use retryRenameSync; the real retry loops — atomicRenameWithRetry, capability-ledger/consent, build-hooks — remain compliant via continue/errno); lint:ci green; disable-ban + vocab-drift green. --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
9d52043f50 |
feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5) ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content (src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result (path.basename excluded — returns a separator-less filename) interpolated into an @-reference / config-dir markdown body without .replace(/\\/g,'/') normalization, per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT. Build-and-assess found the canonical defect site (computePathPrefix) already compliant and only 1 src/ hit — a false positive (path.basename in a status message) — eliminated by narrowing (exclude basename; require a real @-ref/ config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention. The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES; CONTEXT.md predicates + how-to doc updated. - RuleTester suite (26 cases) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1733): add changeset for Windows agent-skills path-leak fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd). On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet. Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors. Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci: re-run golden-install-parity on src/lib + installer changes (close drift guard) golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it must re-run whenever the built lib could change. ci-test-scope selected it for neither src/** nor installer changes, so a source-only edit (e.g. #1691's milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the 'TS runtime sources' and 'installer and package layout' selection rules, with behavioral regression tests for each. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
b307c4cfde |
refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move) (#1735)
* refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move) Relocate the runtime-artifact install cluster out of the 12,490-line bin/install.js into a dedicated src/install-engine.cts -> install-engine.cjs: installRuntimeArtifacts, uninstallRuntimeArtifacts, installOpencodeFamilySkills, and their cluster helpers (_copyStaged, snapshot/restore, legacy migration, GSD-entry pruning, preserve/restoreUserArtifacts, OpenCode-family converters, USER_OWNED_ARTIFACTS). - bin/install.js imports the engine and re-exports the moved symbols for back-compat; getCommitAttribution STAYS in install.js (impure config I/O + argv explicitConfigDir global) and is injected via a resolveAttribution param. - 17 test files migrated to import the moved symbols from the engine. - Bookkeeping: eslint built-artifact ignore, .gitignore, INVENTORY manifest+row, CONTEXT.md Install Engine Module glossary seam. Behaviour-preserving: install output is byte-identical for all 16 runtimes (golden-parity harness #1730) — the only delta is the new install-engine.cjs file shipping in the installed gsd-core/bin/lib/ tree. Closes #1734 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1734): backfill changeset PR number (#1735) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
f08b177215 |
feat(#1726): G1-G6 portability AST rules; fix all offenders; delete the ratchet (Phase 4) (#1731)
Phase 4 of epic #1702. Closes #1726. |
||
|
|
c57e0d56c2 |
feat(#1720): no-unguarded-nonportable-exec AST rule; retire the regex script (Phase 3) (#1723)
Phase 3 of epic #1702. Closes #1720. |
||
|
|
481d121dd4 |
feat(#1711): no-posix-mode-bit-assert AST rule (Phase 2) (#1718)
Phase 2 of epic #1702. Closes #1711. |
||
|
|
a72dbfa58c |
feat(#1707): no-path-literal-in-assert AST rule + portability foundation (#1710)
Phase 1 of epic #1702. Closes #1707. |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a46109b97 |
enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior. |
||
|
|
ba96c70b14 |
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a deterministic classifier that `verify-work` consumes to route deliverables to auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic. - New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage block (extractFrontmatter can't — its `-` items are scalars-only; this is a focused parser, sibling of parseMustHavesBlock), validates each entry, and classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE typed-IR surface. Exposed via `uat classify-coverage --summary <f>`. - Auto-pass is the narrow proven case only: strict-boolean human_judgment:false AND non-empty all-`pass` verification AND zero validation errors. Everything else — judgment, empty/failing verification, malformed entry — routes to the human (fail-safe). A malformed block falls back to legacy prose extraction and surfaces an error; an absent block is byte-identical to pre-#1602. - execute-plan create_summary populates the block (fail-safe default human_judgment:true); verify-work extract_tests consumes it; create_uat_file marks auto-passed entries `source: automated`. - Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY, eslint/gitignore registration, and Diataxis docs (COMMANDS reference + USER-GUIDE explanation) updated. - Behavioral tests via the CLI (no source-grep); parser-robustness regressions for the null-entry/comment-header/mis-indent cases found in adversarial review. Closes #1602 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
405ae9b3b7 | refactor(#1557): add runtime artifact install plan module (#1560) | ||
|
|
e7855bc217 | fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) | ||
|
|
9219af3360 |
feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch. Closes #1433. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1abb0d3427 |
feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability install command + consent gate are Phase 4). NEW modules only; not yet wired into bin/install.js. - src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/ git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install), tarball (https download + sha512 integrity-before-extraction + tar), registry (stub). SECURITY: install never executes capability code (copy/extract only, --ignore-scripts); integrity verified before staging; symlink members rejected at interior, source-root, AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction; npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell); git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet. - src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest: readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent, prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir). - 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex adversarial rounds; all 12 findings fixed + regression-tested. - .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen. Closes #1432 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1432): add changeset for capability source resolver + ledger Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |