12ee75a509cbfd8fe2ae2871056fa81777d7dcf5
60 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
12ee75a509 |
chore: point MSD at git.golem15.com/golem15/msd-core
Some checks failed
Tests / PR mergeability (push) Successful in 1m37s
Tests / Base branch health (push) Successful in 11s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 2m16s
Tests / plugin-validate (push) Successful in 1m6s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 25s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24) (push) Failing after 19s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 19s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 9s
Close Draft PRs (sweep) / Sweep open draft PRs (push) Successful in 8s
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
Tests / Required tests (push) Has been cancelled
The fork lives on the golem15 Gitea forge, not GitHub. Package identity now parses either host and derives in-place raw URLs for Gitea; the identity-drift lint accepts the new host; README drops GitHub-only badges and the npm quickstart in favour of the checkout installer. |
||
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
b2d50ffd83 |
fix(#4489): make capability-registry.test.cjs's extractShellBlocks CRLF-safe (#4584)
* fix(#4489): make capability-registry.test.cjs's extractShellBlocks CRLF-safe Second, independent copy of the #4409 CRLF-fragile line-splitting bug, explicitly flagged as out of scope there ("other duplicated helper in file sibling test files not part of the shadowing chain, tracked separately if divergent"). Same fix: content.split('\n') -> content.split(/\r?\n/), matching src/text-lines.cts's splitLines() and the already-fixed sibling copy in tests/runtime-launcher-parity.test.cjs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4488): regenerate INVENTORY-MANIFEST.json for tdd-red-evidence.cjs's ADR-457 untracking Discovered while validating #4489's push: #4488's merge (untracking gsd-core/bin/lib/tdd-red-evidence.cjs per ADR-457) left docs/INVENTORY- MANIFEST.json stale, since that file was previously listed as a tracked shipped artifact. Removed via node scripts/gen-inventory-manifest.cjs --write. docs/INVENTORY.md already described this file as gitignored (no update needed there -- it already documented the intended state). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * revert: undo incorrect INVENTORY-MANIFEST.json edit from 57f1457a27 The prior commit removed tdd-red-evidence.cjs's manifest entry based on a false premise: a stale tsconfig.build.tsbuildinfo (gitignored, untouched by git checkout/rebase) told tsc the file's compilation was already current even though git's own checkout had deleted the actual output file during this branch's rebase onto #4488's merge (a tracked-in-old-tree, untracked-in-new-tree transition deletes the working-tree file regardless of the new .gitignore entry). tsc's incremental cache doesn't verify its recorded output still exists on disk, so it silently skipped re-emitting it. Confirmed real root cause: deleting tsconfig.build.tsbuildinfo and rebuilding fresh correctly re-emits gsd-core/bin/lib/tdd-red-evidence.cjs (it is gitignored now, not deleted -- src/tdd-red-evidence.cts is unaffected by ADR-457's tracked-vs-gitignored distinction and always compiles). The manifest's own purpose (per its docstring) is 'every shipped surface derived entirely from the filesystem' -- this file still ships via the normal build, so it belongs in the manifest regardless of git-tracking status. Net result matches next's own INVENTORY-MANIFEST.json byte-for-byte; this correction should not have been needed at all had the build cache been fresh when the prior commit was made. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
18c899def5 |
enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a262ad6b61 |
fix(#4148): dispatch wave-pre step hooks (#4185)
* fix(#4148): dispatch wave-pre step hooks External capabilities can render step hooks before a wave, but the execute workflow consumed only contributions and silently skipped every step. Reuse the shared dispatch contract before executor spawning and pin the capability-validator boundary with a red-first regression. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * test(#4148): pin wave-pre dispatch ordering * test(#4148): pin wave-pre dispatch contract * chore(#4148): bind upstream changeset PR * chore(#4148): restore fork changeset identity * fix(#4148): align wave-pre dispatch contract Mirror the sibling wave-post all-shapes clarification while pruning redundant prose so the rebased workflow remains below its frozen byte ceiling. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * chore(#4148): restore upstream changeset identity * fix(#4148): align wave-pre capability guidance * docs(#4148): identify wave-pre manifest input Name the third-party manifest trust origin at the wave-pre dispatch boundary so the reviewer-requested validation guidance matches wave-post. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * docs(#4148): preserve execute-phase byte budget Remove a redundant advisory label while retaining the non-blocking contract, keeping the reviewer-required trust-boundary wording at the enforced 93,400-byte ceiling. * fix(#4148): mark wave-pre manifest-input validation as security-relevant Reviewer nit on PR #4185: wave-pre's step-dispatch sentence had the (third-party manifest input) parenthetical but dropped the ⚠ marker that wave-post's parallel sentence (execute-phase.md:1044) carries, losing the visual flag that this validation is security-motivated. Trims the redundant "of one" from "not one shape of one" to reclaim the 4 bytes the marker adds — the ADR-857 byte-margin gate (tests/claude-orchestration.test.cjs) leaves zero slack at the 93,400-byte ceiling. * fix(#4148): trim wave-pre step-dispatch prose to clear ADR-857 byte ceiling Merging next's unrelated growth (#3990's TDD_APPLICABLE conditional) pushed execute-phase.md 116 bytes past the 93,400-byte ceiling, failing CI on all three platforms. The security-relevant ⚠ marker and ref.command validation call-out (added per prior reviewer nit) are preserved verbatim per the pinned regression test in capability-registry.test.cjs; only the non-pinned connective prose is trimmed. * fix(#4148): recalibrate execute-phase.md self-imposed margin, restore security marker next grew execute-phase.md by ~230 bytes across two unrelated merges during this fix (#3990's TDD_APPLICABLE conditional, then a further step-extraction commit), consuming this test's own self-imposed 93,400 safety buffer under ADR-857's actual, unmodified 93,600 ceiling (docs/adr/857-capability-system.md:22). The wave-pre step-dispatch sentence cannot shrink further without dropping one of the pinned substrings this same test file asserts on (kind=="step", loop-hook-dispatch, never blocks or redirects executor spawning, Validate `ref.command`). Raises the self-imposed margin to 93,550 (still 50 bytes under the real, untouched ADR ceiling) and restores the ⚠ marker the prior reviewer round required for the ref.command validation call-out, which byte pressure had dropped. * fix(#4148): restore full ref.command validation wording, drop self-imposed margin Adversarial review (agy/gemini-3.8-flash-high) flagged two issues in the prior CI-recovery commit: 1. Trimming "in-context before any shell use" from the step-dispatch warning weakened the inline operational instruction (the reader is told WHAT to validate but not the specific in-context-not-shell mechanism the referenced loop-hook-dispatch.md:45-51 threat model requires). Restored it - the merge with next since the last commit freed enough real margin (77 bytes under the untouched 93,600 ADR-857 ceiling) to afford it without any margin change. 2. The prior commit self-imposed margin bump (93400 to 93550) was, on reflection, the wrong lever: it is a number this PR invented, not an ADR value, and re-bumping it every time next grows execute-phase.md is a losing pattern (already needed twice in one session). Removed the redundant assertion; the same line existing bytes-under-93600 check against the real, frozen ADR-857 ceiling (docs/adr/857-capability-system.md:22) is the actual invariant and is untouched. workflow-size-budget.test.cjs tier hard cap (98304 bytes, extract-not-bump by design) remains the correct backstop for runaway growth. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
f16ff7d1b3 |
enhance(#3545): widen no-source-grep with fold+hooks, migrate 76 sites (#4161)
* feat(#3545): widen no-source-grep with one-hop path-fold and hooks dir Resolve a readFileSync() path argument that is a bare Identifier one hop back to its VariableDeclarator initializer before classification, and recognize `hooks` as a source directory alongside bin/lib/gsd-core/src. Measured (epic #3464 phase 7): fold+hooks together newly flag 76 unsuppressed sites across 18 files that were previously invisible to identifier-indirected or hooks/-rooted source reads. Neither widening alone is sufficient — hooks-only surfaces 0 new sites, confirming #3520's prior finding that the identifier-indirection gap must close first. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3545): migrate 76 sites newly flagged by the fold+hooks widening Per-site classification: rewrite behaviorally (require() the real module, assert on its actual exported behavior) wherever the read was a proxy for code behavior; add a site-scoped `// allow-test-rule: <reason> (#3545)` marker only where the raw source text genuinely is the product under test (codex-config.test.cjs's adapter-header-contract checks, install.js structural-wiring guards with no exported symbol, AST-parse fixture inputs, etc.) — each marker cites an existing repo-sanctioned category from CONTRIBUTING.md's allow-test-rule exception table. Also converts two try/finally test bodies (introduced during this same migration) to the required t.after() cleanup pattern per CONTRIBUTING.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3545): re-baseline effective-exemption ceiling to 81 The fold+hooks widening's own newly-detected sites are now suppressed by site-scoped markers, moving them from invisible into the tightly-ratcheted effective-exemption count. Ceiling rises from 10 to 81 (the exact measured high-water mark, grace unchanged at 2) — a deliberate, measured re-baseline per the widening working as intended, not an ordinary ceiling bump. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3545): use canonical allow-test-rule category tokens 4 markers added during migration cited an issue ref correctly but didn't use one of CONTRIBUTING.md's seven recognized category tokens, unlike every other marker in this change. Cosmetic only — same suppression lines, same effective/live counts (81/81, 0 live). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3545): correct stale phase-artifact path in test comment Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4dfc46bbe7 |
enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate * feat(#3348): add context-drift pre-check gate for plan-phase Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md effective last-changed time (git commit time, falling back to mtime for uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently reuses an upstream artifact that predates a decision added to CONTEXT.md after that artifact was derived from it. Deterministic, no model call. New `gsd_run verify context-drift <phase>` command, sibling to the existing verify.codebase-drift/verify.schema-drift gates in the drift capability. Warn-only by default (workflow.context_drift_precheck), with an opt-in workflow.context_drift_action: block escape hatch. Wired at plan:pre in plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions. * fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution * fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe The #3859 empty-diff guard decides whether a submodule bump would land using `--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules` config. The real `git commit -- <paths>` that follows was never given the same override, so under a bare `diff.ignoreSubmodules=all` repo config the two calculations disagree: driven on git 2.39.5 (Debian bookworm, the linux-node24 test-matrix image), the guard correctly stands aside but the scoped commit itself then silently fails (exit 1, no error text) for a gitlink bump it had just confirmed would be recorded, surfacing as commit_failed instead of committed:true. Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the probe and the commit it protects can never diverge. Harmless when no submodule path is involved (driven: identical outcome on an ordinary scoped file, with and without the flag). * fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal * fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a current-milestone enumeration). This PR's refactor pass lifted that block into a shared helper, resolvePhaseDirByToken, also used by the new cmdVerifyContextDrift — the guard tracks exemptions by enclosing function name, so the readdirSync now lives in an unexempted function and started firing. Move the exemption to resolvePhaseDirByToken (same written reason, now covering both callers) instead of migrating to listAllPhaseDirs, which would introduce two real behavior deltas here: it catches readdirSync failures internally (old code let them throw) and sorts results by phase number before matchPhaseDirs picks matches[0] (old code used raw, OS-dependent readdirSync order). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): satisfy lint:ci — slash form, capability registry regen - docs/features/context-drift-gate.md used the deprecated /gsd: colon form; docs are never passed through the install-time slash-form converters, so lint-docs-command-form requires the hyphen form. Regenerated docs/FEATURES.md from the corrected fragment. - Regenerated gsd-core/bin/lib/capability-registry.cjs after editing capabilities/drift/capability.json (lint:generated-sync). * fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit` subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for every scoped commit call, breaking 17 position-based assertions in the commit-files pathspec regression suite that read `a[0] === 'commit'` to find the commit invocation among recorded git calls. `execGit` already accepts an `env` option merged onto `process.env` before spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0` as a per-invocation config override functionally identical to `-c key=val`, expressed via env instead of argv. Passing that env alongside the existing commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged) fixes the real commit's effective diff.ignoreSubmodules to match the empty-diff guard's probe without moving anything in argv position 0. Scoped to exactly the canScope branch, matching the probe's own preconditions and leaving no behavior change for commits the probe never evaluated. No test file changes needed — the 17 previously-failing assertions test argv[0] against the array passed into execGit, which never changes. * fix(#3348): register verify-context-drift in the check subcommand router The drift capability's new plan:pre gate declares check.query "verify.context-drift", which normalizes to `check verify-context-drift`, but no such subcommand was routed — phase6-capstone-conformance's uniform-block-field test failed with "Unknown check subcommand" for every declared gate query. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot of the drift capability's config keys. #3348 legitimately adds two new keys (workflow.context_drift_precheck, workflow.context_drift_action) for its own plan:pre context-drift gate — extend the expected set (and clarify the assertion message) without weakening the test's exactness. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction #3348 (an earlier commit on this branch, e4b80ad81) extracted cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block into the shared resolvePhaseDirByToken helper (also used by the new cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's function-scoped exemption from cmdVerifySchemaDrift to resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer contains a line the guard's detectors match, so it needs no exemption. tests/phase-locator.test.cjs's E2 test still pinned the exemption to the old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to match the same "migrated call site's exemption must move, not duplicate" pattern the test already applies to cmdRoadmapAnalyze and cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the still-exempt list and add symmetric assertions that it no longer carries the exemption while resolvePhaseDirByToken now does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs 'always exits 0 (query command contract)' included the no-phase-arg case, which contradicts the file's own earlier 'errors with usage message on missing phase arg' test (that case legitimately exits 1 via the Usage error) — drop it from the always-exits-0 cases. 'degrades to mtime comparison outside a git repo' and '...in a repo with no commits' relied on real wall-clock ordering between two back-to-back writeFileSync calls to prove CONTEXT.md is newer than RESEARCH.md; on a fast filesystem both can land in the same mtime tick, producing a tie that computeContextDrift's strict `<` correctly treats as not-stale, so stale_artifacts comes back empty. Make both tests deterministic via explicit fs.utimesSync instead of relying on timing (CONTRIBUTING.md: never assert elapsed wall-clock time). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture The "all plan:pre when-keys false" fixture explicitly disables every known workflow.* plan:pre toggle, but didn't yet know about the new workflow.context_drift_precheck key (defaults to true), so the new drift context-drift gate stayed active and broke the empty-activeHooks assertion. Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3348): backfill changeset PR number (pr:0 -> 4147) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b848b23861 |
feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md - Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's existing render + generic contribution dispatch pattern - Inject planner-targeted contribution fragments into the planner prompt, after AGENT_SKILLS_PLANNER, matching D-08 ordering - Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions - Record quick.md's byte-growth rationale in this commit trailer for every gsd-core-verbatim runtime Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does. * test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call < agent-skills placeholder < injection block < spawn ordering) assertions, all extracted via a brace-bounded slice anchored on the literal injection instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER} and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain brace pairs). Add a capability-registry.test.cjs describe block proving the registry-wide D-07 exclusion is meaningful: at least one plan:pre contribution exists, every plan:pre contribution has a non-empty into/fragment.inline, and the registry as a whole carries at least one non-planner-into contribution. Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's real content, and buildContract() still yields exactly 5 entries. Verified red-without-Task-1 by temporarily reverting quick.md to its pre-f30de9cc content and re-running these three suites (D-08 failed as expected), then restored via git checkout and re-confirmed green. * docs(#3778): note quick planning also renders plan:pre in the tutorial The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the plan:pre hook set. Since quick.md now dispatches the same hook set (f30de9cc), the sentence understated the capability's real reach. * feat(#3778): add changeset fragment * chore(#3778): reference the upstream issue in the changeset fragment The fragment was the only one of 81 in .changeset/ without a trailing (#NNNN) reference or a bold lead-in. serializeChangelog auto-appends only the pr: field, so the rendered CHANGELOG entry carried no link back to issue #3778. * test(#3778): scope the D-07 registry assertion to what it actually proves The registry-wide non-planner check was named "D-07 exclusion is meaningful", which overclaims: it proves only that `into` takes non-planner values somewhere in the registry, not that anything is excluded at plan:pre. Every plan:pre contribution is currently into: "planner", so the filter is a forward-looking safeguard there. Narrowing the assertion to plan:pre (as review suggested) would fail today. Asserting plan:pre is all-planner would be brittle — it would break the day a legitimate non-planner plan:pre contribution lands, which is exactly when the safeguard starts doing work. So the assertion is unchanged and only the name and comment are corrected. * chore(#3778): point the changeset fragment at the upstream PR The fragment carried pr: 3, the fork staging PR. changeset lint derives the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that would read as pr-field drift. Point it at open-gsd/gsd-core#3934. * test(#3778): require contributions in Quick revision prompts * test(loop-host): require Quick auxiliary registration * fix(#3778): preserve contributions in Quick plan revisions * fix(#3778): validate Quick as a planner contribution host * fix(#3778): tighten Quick contribution contract * test(#3778): drop unnecessary source-contract exemption * fix(#3778): require Quick planner target coverage * docs(#3778): describe targeted auxiliary coverage --------- Co-authored-by: davdittrich <davdittrich@gmail.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
15af0f5536 |
enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
86fa2917d7 |
enh(#3866): dispatch step and contribution hooks at verify:pre (#3869)
* test(#3866): pin that verify:pre must dispatch every hook kind verify-work.md's verify_pre_hooks step dispatches only `kind == "gate"`, so getWiredKinds reports verify:pre -> {gate} and gen-capability-registry rejects any capability declaring a step or contribution there. The verify lane is therefore closed to capabilities that want to contribute to what UAT covers rather than refuse to let it start. Failing-first: the step, contribution, and exact-kind-set rows are RED; the pre-existing gate row is a green regression pin so the new arms cannot orphan the arm verify:pre already had. Refs #3866 * feat(#3866): dispatch step and contribution hooks at verify:pre verify_pre_hooks dispatched `kind == "gate"` only, so getWiredKinds reported verify:pre -> {gate} and gen-capability-registry's validateHooksWired rejected any capability declaring a step or contribution there. A capability could refuse to let UAT start; it could not contribute to what UAT covers. Add contribution and step arms mirroring execute:wave:post, deferring to references/loop-hook-dispatch.md and carrying its ref.command in-context validation guard ahead of any shell-use prose. A verify:pre step is advisory: it never blocks the start of UAT and an erroring step is routed by its own onError. The gate arm and its check guard are untouched. Give extract_tests an additive consumption seam for the artefacts those steps declare via the existing steps[].produces field -- no new registry field, no new ordering, no invented filename. Manifest-supplied artefact names are validated in-context against an allowlist and resolved only inside PHASE_DIR. With no producing step the derivation is unchanged, pinned by test rather than asserted in prose. Review findings folded in: the artefact-name allowlist (isolated adversarial pass), the artefact-shape contract and the seam-inertness tests (spec axis), and the reference/how-to split so one constraint has one source of truth (standards axis). Closes #3866 * chore(#3866): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
ea594300d9 |
fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard * fix(#3606): validate hook-kind coverage at call sites and dispatch generically * fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction * fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex * chore(#3606): regenerate install-tree fixtures for new wave-post fragment * chore(#3606): sync canonical launcher preamble into new fragment * fix(#3606): keep fragment preamble ahead of first gsd_run mention * fix(#3606): revert sync script's preamble move in explore.md * chore(#3606): regenerate derived manifests post-rebase * chore(#3606): allowlist peer test files - base was red on the count lane * chore(#3606): regenerate inventory for peer's verify-command-grounding doc * chore(#3606): grounding test maps to its own module by longest prefix * chore(#3606): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9e4f0e99ad |
fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash
bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.
Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.
The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.
Refs #3631
* fix(3631): exclude derived bytecode caches from the consent digest
RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.
collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.
Three properties were preserved deliberately, each pinned by a test:
- The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
first would have turned the exclusion into a way to smuggle a symlink past the check;
a symlink named x.pyc still throws.
- Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
caps guard the WALK; the digest answers a different question, and exclusion must not
become an unbounded-bytes hole.
- An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
directory marker was the sharper half of this bug: an empty __pycache__ flipped the
hash before any .pyc existed, so a *.pyc-only filter would have left it live.
The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.
Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.
Fixes #3631
* fix(3631): narrow the digest exclusion after two isolated security reviews
The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.
HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.
FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.
Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.
Narrowed to what is actually defensible:
- a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
the walk still recurses and hashes every non-excluded child.
- .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
- a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
- declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
basename are now rejected in both validator copies — a file named .pyc can contain
perfectly valid JavaScript, so the exclusion must not be reachable from a declared
surface.
Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.
Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.
Refs #3631
* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind
Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.
HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.
The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.
The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.
Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.
Changeset rewritten: it still described the rejected wholesale-exclusion semantics.
Refs #3631
* chore(3631): backfill changeset PR number (#3650)
---------
Co-authored-by: sim <sim@local>
|
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
967bddba37 |
fix(#3384): strip mcp__* tool grants from zcode-installed subagents (#3483)
* fix(#3384): strip mcp__* tool grants from zcode-installed subagents ZCode's dispatcher treats every mcp__<server>__* entry in an agent's tools: frontmatter as a required MCP server and hard-fails the subagent spawn (CONFIGURATION_ERROR) when it is not connected, whereas Claude Code treats the same grants as an optional allowlist. ZCode shared Claude's verbatim agents copy (converter: null), so all 8 MCP-granted agents failed to spawn out of the box with zero MCP servers configured. Add convertClaudeAgentToZcodeAgent — a line-surgical converter that filters mcp__* entries out of the frontmatter tools: grant list (both inline comma and YAML block-list shapes) and preserves every other byte. Declare it on both of zcode's capability.json agents entries and cut zcode over to the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES) so the legacy inline loop stops deleting+re-copying the converted agents raw. Claude Code, Kimi, and Gemini install behavior is unchanged. * chore(#3384): link changeset fragment to pr 3483 --------- Co-authored-by: sim <sim@local> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ce38d44811 |
fix(#2777): remove stale codex local home metadata (#2831)
* fix(#2777): remove stale codex local home metadata * chore(#2777): add changeset for codex local layout metadata --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c2a305c44d |
feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout PR 1 registered the kimi-code EoS descriptor with empty artifactLayout (SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test). PR 2 fills in the install surface: - src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill function. Today it delegates to convertClaudeCommandToKimiSkill (Python kimi-cli) because Kimi Code uses the same Agent Skills format + /skill: invocation per official docs. The distinct function name lets a future divergence land cleanly if Kimi Code's skill format evolves independently. - gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS. - capabilities/kimi-code/capability.json: artifactLayout.global now declares the skills kind with converter='convertClaudeCommandToKimiCodeSkill' + home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code docs: merge_all_available_skills = true default). - tests/installer-migration-install.integration.test.cjs: REMOVE the SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface. - Regenerated capability-registry + capability-matrix + golden install parity + install tree fixtures for kimi-code. * fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump - src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path can dispatch off the descriptor's converter string. - tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count 26 → 27 (added convertClaudeCommandToKimiCodeSkill). * fix(#2454): remove home override from kimi-code skills (inherit configDir) The home:'.kimi-code' override made the install plan resolve skills dest to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's temp configDir to miss the install. Removing it lets skills inherit configDir like most runtimes. * fix(#2454): kimi-code install contract surface is flat-skills (no agents) Kimi Code has NO custom named subagents (per official docs: 3 built-in coder/explore/plan only). The kimi-skills-agents surface expects agents/ gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed to flat-skills which only checks for skills/gsd-* dirs. * docs(changeset): Phase 2 kimi-code install layout Added (#2509) * docs(changeset): backfill PR #2520 for Phase 2 (#2509) |
||
|
|
20ff405cb3 |
feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline New statusline.state_format config, enum full|compact (default full — existing rendering untouched). "compact" renders the state segment as "<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 · executing" — dropping the milestone name and progress bar (the two biggest width costs) and collapsing narrative statuses to a single keyword. Per the #2162 approval conditions, the keyword set is the canonical vocabulary from normalizeStateStatus() in state-document.cjs (discussing/planning/executing/verifying/completed/paused) — no parallel hand-rolled list, so the vocabularies can't drift — and the canonical stuck state "paused" renders uppercase as PAUSED (no new "blocked" lifecycle state). Statuses the normalizer passes through unrecognized fall back to their first word capped at 16 chars. Lifecycle scenes preserved: active_phase wins over the body phase number, milestone completion renders "complete", idle-with-next-action renders "next <action> <phases>". Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2162): changeset fragment for PR #2175 * fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format - register statusline.state_format in the fix-1628 coercion-bypass matrix - 15/16/17-char boundary tests for the shortGsdStatus fallback cap - changeset body ends with the (#2162) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage - compact renderer gates the milestone-complete scene behind the absence of an in-flight phase id, mirroring formatGsdState's if/else precedence (Scene 1 beats Scene 3); regression test covers the non-atomic active_phase + percent=100 STATE.md shape - direct config-set accept/reject test for statusline.state_format plain strings (ENUM_KEYS matrix covers only the JSON coercion shapes) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests) Re-review Major: gating done on !phaseId held completion back for the legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100 regardless of phaseNum, so compact must too. Gate is now !s.activePhase. The phaseNum-only test now expects 'complete' and cross-checks the full renderer; a parity test feeds identical inputs to both renderers. Re-review Minor: shortGsdStatus gets fast-check property coverage (totality, canonical fixed points, separator safety, fallback shape). Golden fixtures regenerated for the hook byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
a0fafedfa0 |
feat(#2103): drive VS Code through the Embeddable Orchestration System (ADR-1239)
VS Code is a net-new EoS runtime that — unlike every prior migration — is NOT CLI-installed (Marketplace/VSIX extension). It has zero runtime==='vscode' branches in bin/install.js and stays that way (regression-guarded); it is driven entirely through the negotiated imperative Host-Integration adapter. Registry + validator (the hard part): - capabilities/vscode/capability.json (role:runtime): full hostIntegration block (imperative / palette / active vscode.lm model / engine hook bus / sandboxed-storage / mcp transport / sandboxed-web runtime; dispatch nested, maxDepth 5 per VS Code's documented subagent depth). - capability-validator.cjs extended so a role:runtime capability can legitimately declare "extension-distributed, no config directory": new configHome.kind:'none' + installSurface:'none' (+ GATE-A pairing + the parity maps), with localConfigDir and configHome.name made conditional on kind!=='none'. All 18 runtimes still validate; getDirName returns a distinct sentinel (not '.claude') for a no-config runtime. - The add-a-registry-runtime tax: NON_INSTALLABLE_RUNTIMES exemption in the runtime-flags drift guard, vscode added to global-config-home SPECIAL_CASED, EXPECTED_PROFILES.vscode='ide', and the config-adapter/derivation/pin-count guards updated. No golden-install fixture, model-catalog, or CONFIGURATION rows (vscode never enters allRuntimes). Dispatch + extension surface: - Fixed vscode/extension.js's createHub()-no-args bug (every dispatch was UnknownCommand, masked by a vacuous reachability test) — now reuses the shared dispatchGsdCommand subprocess-shim (Node/desktop); the reachability test is tightened to assert real dispatch. - Promoted the #1933 host binding to a shipped vscode/host-binding.js; activate() now composes the model/hookBus/stateIO seams through it. Corrected the model seam to VS Code's real API (vscode.lm.selectChatModels() -> model.sendRequest(); vscode.lm.sendRequest does not exist) so the binding actually composes on real desktop VS Code instead of throwing. - New vscode/browser.js Web Extension entry with ZERO Node APIs (the engine's config/capability loading is Node-bound, so the web entry registers the surface and directs full dispatch to the native MCP server — honestly documented). - UPGRADE 1: GSD skills as native Language Model Tools (contributes.languageModelTools + vscode.lm.registerTool), invoke() dispatching through the hub. - UPGRADE 2: native subagent dispatch wired onto #runSubagent / chat.subagents.allowInvocationsFromSubagents (fail-soft on API availability, maxDepth 5 enforced). - vscode/package.json: browser entry, engines.vscode ^1.105, chatParticipants + languageModelTools contributions; fixed a stale activationPoints->activationEvents manifest key. Added "vscode" to the package files array. Docs (## vscode matrix section) + changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd613566cb |
feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven hostBehaviors (byte-parity — no fold changes any install output): - 2 dead destructures dropped (uninstall, finishInstall); the dead `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable). - skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions. - legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate. - installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy (workflow-delegation target — load-bearing; local-install verified intact). - verificationStyle:"windsurf-workflows" folds the workflow-count report. - Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf). Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js, install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson (Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing .windsurf/hooks.json with two BLOCKING pre-hooks: - pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the active git worktree / into .git internals. - pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs; force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next). Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow, fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex; 4096-char cap) with the fail-closed false-positives fixed post-review. The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot #2099 faithful-subset precedent). extendedHookEvents stays []. Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8 shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash (functionally inert for non-windsurf; the established cursor pattern). No install-output change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference- windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking + allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5695522d5f |
feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads: getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix (→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite), getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity destructures (antigravity is already on the descriptor-agents path). subagentToolkit flipped undocumented→full (Context7: antigravity.google/docs/cli/features); namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical golden parity for all 16 runtimes. UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's settings.json — non-destructive, idempotent, symmetric uninstall. Added to VALID_PERMISSION_WRITERS + the FinishPermissionWriter union. UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json registering the gsd-core companion MCP server (Gemini-successor mcpServers schema, best-effort — raw schema unpublished). Both writers dispatch from finishInstall. settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable, no absolute paths) is golden-tracked → only antigravity.json changes. Tests: declarative-reference-antigravity extended (source-grep guard across 4 modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) + antigravity-upgrades (permission-writer + mcp_config live-install, idempotency, user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md + connect-gsd-mcp-server docs updated; changeset (Changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab04916682 |
feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle, reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES. Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by- name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain. UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker- delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts). GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7- confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/ SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards removed). config.toml holds absolute install paths so it's golden-excluded via an exact relative-path (.kimi/config.toml), not a basename (which would blind Codex's config.toml). New hooksSurface value added to the closed enum in capability-validator + runtime-config-adapter-registry. UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's Agent tool takes run_in_background; root agent already gets the Agent tool), so negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented' per AC (coder/explore/plan have distinct tool policies). MCP transport explicitly deferred (no installer-driven MCP for any runtime). Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others + claude-local byte-identical (kilo/zcode keep their own exclusions). Tests: kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency + backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated; changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f014ec83bd |
feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads: finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks copy), and a skills converter-name registry (the artifactLayout.converter field is now load-bearing, not decorative). frontmatterDialect stays the documented dispatch key for frontmatter (no descriptor field for it). Dead isKilo destructure bindings removed. Byte-identical golden parity for all 16 runtimes (opencode, which shares kilo's combined-family path, verified clean). UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin + extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus). UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented' per AC so dispatch degrades to 'degraded' by design. Model-catalog single-source edit ripples the shared model-catalog.json hash into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale- bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/ codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error. Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/ hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades (plugin parity+load, model-override converter, agents dispatch surface, MCP doc). Matrix + how-to + config docs updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c6ce110efa |
feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter, branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json, read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches in install-engine.cts folded to descriptor flags (claude+hermes descriptors updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity for qwen/hermes/claude(global+local). UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents artifact-layout kind + convertClaudeAgentToQwenAgent converter (name + description + tools YAML block list; color/model dropped). qwen routed onto the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES). UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the descriptor-gated hook-writer loop (activates only for qwen). Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors + no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix + how-to updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d10f9c675e | test(#2091): update closed-vocab assertions for hermes extensionEvents dialect | ||
|
|
69b309e4e0 |
feat(#1925): add ZCode (Z.ai) as a pluggable runtime descriptor
Add ZCode as a first-party runtime via a declarative capability descriptor (capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode' branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0 (ADR-1016 / ADR-1239) enables. Descriptor (all axes sourced verbatim from zcode.z.ai docs): - configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install - Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter) - hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false (foreground-only per docs); nested+maxDepth undocumented; passive model mode Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap, interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout + ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor. Revamped the brittle per-runtime golden-master tests to be count-agnostic, descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy, config-adapter-registry (intent + install-plan golden master), capability-registry, host-integration-descriptors (counts derive from curated maps). Adding a runtime descriptor now extends coverage with zero edits to those suites. Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration capability matrix (every axis cited) updated. |
||
|
|
ed79902509 |
feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually route recall/capture instead of silently behaving as `augment`: - kg_backend: the palace temporal KG is the primary knowledge-graph source; native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive. - replace: recall resolves through the palace as the source of truth; native artifacts are the fallback. Every mode stays onError:skip and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing .planning/graphs/, so no memory is lost. Cross-mode .planning/graphs/ migration remains a documented open question (PRD/ADR §17), out of scope here. Surfaces updated (instruction-only contract): recall/capture commands (+ generated skills), discuss/wave fragments, curator agent, capability.json schema. Docs: how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated capability-registry, golden install-parity fixtures (mempalace hashes only), agent-size-baseline. Added a routing-contract + cross-surface parity test. Incidental (folded per no-defer rule): removed pre-existing unused imports (spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs) that eslint flagged in/alongside the touched files. Closes #2007 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b0bd2f7a48 |
chore: move committed-generated-artifact freshness checks to lint:ci (#2000)
gsd-test's build leg runs the full 'npm run build' (which regenerates capability-registry.cjs, loop-host-contract.cjs, package-identity.cjs, etc.), so committed-freshness guards that lived in the unit suite were masked there: gsd-test passed a stale-commit that CI's shard-1/3 test then red-flagged (caught live on PR #1998). The mandated pre-push gate was green on a commit CI correctly flagged. Move the committed-state --check guards into a new 'lint:generated-sync' script wired into lint:ci (the single orchestrated entry point the lint-tests CI job already runs on a build:lib-only tree, so the committed artifacts are checked without regeneration). gsd-test no longer contains these guards, so it can no longer mask them. - package.json: add lint:generated-sync (7 generators --check); wire into lint:ci. - generate-package-identity.cjs: add --check mode (was the only generator without it); no-arg behaviour unchanged (still writes, as build expects). - Remove the committed-freshness guards from the unit suite, keeping all behavioral/structural tests: - capability-registry.test.cjs: drop the --check describe. - loop-host-contract.test.cjs: drop the committed-file staleness test (keep the normalizeLineEndings unit test). - capability-matrix-sync.test.cjs: drop --check + byte-for-byte (keep the architectural content invariants: every cap appears, security ship:pre). - issue-844-manifest-version-sync.test.cjs: drop describe D (--check). - issue-498-package-identity.test.cjs: drop the drift-check test (keep behavioral module-export tests); drop the now-unused render import and its allow-test-rule exemption (allowlist ratcheted 175 -> 174). |
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
85ed50cc4f |
test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file that owns each subject-under-test, across 52 existing suites (state, config, frontmatter, roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard, health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe wrappers; 881 subtests conserved 1:1. No new test files. Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations (intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe. Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across 8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md + ADR-0002/443/1235/3524 test-file references. lint:ci green. Part of epic #1969. Closes #1972. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b51cbf96cf |
feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) (#1946)
* feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) * fix(#1943): re-export VALID_EXTENSION_EVENTS from gen-capability-registry (test import path) * fix(#1943): regenerate capability-registry.cjs + add changeset fragment * fix(#1943): import VALID_EXTENSION_EVENTS from validator, not gen-capability-registry (golden parity) |
||
|
|
a3d3c2a445 |
refactor(#1756): derive getDirName from a documented runtime.localConfigDir descriptor axis (#1757)
ADR-1239 Phase B (parent #1679). getDirName was a hand-maintained 15-branch if-chain mapping each runtime to its local content-rewrite dot-dir. Relocate those values into a documented runtime.localConfigDir descriptor field; derive getDirName from registry.runtimes[id].runtime.localConfigDir (fallback .claude). - 16 capability.json gain runtime.localConfigDir (byte-identical values) - capability-validator.cjs requires it (non-empty dot-dir); registry regenerated - docs/reference/capability-manifest.md documents the field + the three divergent values (copilot=.github, antigravity=.agents, kimi=.kimi-code) - drift-guard test: golden value map + key-set equality both ways Byte-identical install output for all 16 runtimes (golden-parity harness #1730). Closes #1756 Co-authored-by: review-bot <review-bot@gsd> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cf2e66b39e |
feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): backgroundDispatch citations in matrix + CONTEXT note Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1708): address review findings on typed dispatch-flatten Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): backfill backgroundDispatch in role:runtime test fixtures Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): add changeset for typed dispatch-flatten Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1708): remove stray temp PR-body file Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): add issue ref to bug-853 allow-test-rule annotations ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bcc5a6d1ba |
fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command Capability hook install (applyCapabilitySharedEdits) wrote each settings.json hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a fail-closed guard could then block the whole session), and emitted a bare single-quoted script path so a .js-family hook from a git/tarball source without +x failed with Permission denied on every matching call. - Pass through an optional declared `matcher` (entry-level sibling of `hooks`); absent => omitted (match-all), so existing shipped capabilities are unchanged. - Validate `matcher` in the declaration (non-empty string, no control chars). - Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party); .sh and others keep the bare quoted path (unchanged). Root cause: the manifest hook schema (validator rule C4) was {event, script} only with no matcher, and applyCapabilitySharedEdits never read or wrote one; the command used shellSingleQuote(absScript) with no node prefix. Regression tests fail-first on both defects (matcher dropped; bare path) and pass after the fix; #1460 command assertions updated for the node prefix. * chore(#1634): backfill changeset pr:1638 * fix(#1634): resolve lint and windows CI failures - validator: replace the control-character range regex with a char-code loop. The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char codes are equally precise and lint-clean. Behavior unchanged (still rejects matchers containing ASCII control characters incl. DEL). - test: gate the executable-bit precondition on POSIX. Windows fs does not honor POSIX write modes (a 0o644 write reads back as 0o666), so the precondition is meaningless there and failed the windows-latest lane. The node-prefix assertion — the actual fix — is platform-independent and still runs everywhere. * docs(#1634): amend ADR-894 for optional lifecycle hook matcher The `role: "feature"` `hooks[]` entry now carries an optional `matcher` (settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document the field in the §2 schema table and record a Grilling-amendments entry: the install path projects a declared matcher onto the emitted settings.json hook entry (absent = match-all, so shipped capabilities are unchanged), and per-runtime matcher projection (ADR-857 D8) stays a separate concern. This amendment ships with the fix that introduced the field rather than as a follow-up. * docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a test that writes a file with a POSIX mode and then asserts statSync().mode & 0o777 === <octal> passes on macOS/Linux but fails on windows-latest (Windows fs does not honor POSIX write modes — reads back 0o666). Added as a machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/ prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward: gate the mode-bit precondition on process.platform !== 'win32' and keep the platform-independent behavioral assertion running everywhere. |
||
|
|
6d782e309d | test(#1615): update Windsurf workflow expectations | ||
|
|
08d1c57d6e |
fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
353f63d170 |
feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) Promote the registry from a frozen data file to loadRegistry({includeInstalled}), composing the first-party registry with a validated installed overlay (ADR-1244 D2): - Extract the conformance validator to a shared runtime-callable module (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it verbatim, guarded by a generative-parity test (no build-time/runtime drift). - capability-loader.cts: loadRegistry({includeInstalled}) composes first-party ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic- prefixes); full merged-set cross-capability validation; engines.gsd load-time re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path escapes rejected. - semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed. - Wire surface/state + loop to the overlay; loop injects a blocking gate for each skipped gate-kind overlay (fail-closed). - cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd) + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/ config-set call (never eager at module load, never wrong-cwd); first-party path unchanged with no cwd. - run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact); capability-validator.cjs stays linted (#551 migration coverage). Closes #1431 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1431): add changeset for runtime capability registry overlay Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52) The cwd-aware overlay config-key federation added to config-schema.cts (_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd threading) introduced mutable surface uncovered by config-schema's mutation test set, dropping its score to 39.58% (below the 52 break threshold). Add a real-overlay-fixture describe block exercising every branch (cwd guard, overlay loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker score 39.58% -> 77.08%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2421cf1b4a |
feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) Make the capability manifest versioned — the data substrate the Capability Ecosystem (ADR-1244) keys off: - capability.json gains a REQUIRED semver `version` plus the optional ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`, `provenance`); the build-time conformance validator enforces them via a new `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay). - All 32 native capabilities stamped with `version` (= package version, lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep that keeps them in sync, and the issue-844 regression guard is extended. - Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in version strings; range/integrity fields are shape-validated (satisfaction and the load-time gate are deferred to Phase 2/4). - Capability rel-paths emitted forward-slash for cross-platform git correctness. Closes #1430 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1430): add changeset for versioned capability manifest Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1f41a0ce9a |
feat(#1304): add optional activationKey capability manifest field (#1309)
Add an optional activationKey to the feature role of capability.json — the dotted config key that gates the whole capability (e.g. graphify.enabled). gen-capability-registry validates it (non-empty string, reserved-name guard, must be declared in the capability's own config slice, feature-only) and emits it per-capability in the generated registry. Declared on graphify + intel. No runtime consumption yet (resolver wiring lands in #1305). Part of #1302. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b783410815 |
refactor(#1191): inject clock/reset testability seams + handle valid-null settings (#1233)
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings - worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs). - active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting. - gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle. - install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests. Closes #1191 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1191): add changeset for valid-null settings fix (#1233) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
73b7f45140 |
feat(#1173): wire agent converters into descriptor-driven install path (#1227)
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route agents-kind entries through a converter when the descriptor carries a non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter` to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9 agent converter names, and adds a fail-first behavioral test suite (9 tests) proving the new wiring end-to-end. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b1e8a74708 |
fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no `loop render-hooks` dispatch and was absent from the conformance gate's HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post. - discuss-phase.md: add minimal discuss:pre (before analyze_phase) and discuss:post (after write_context) render-hooks dispatch steps that delegate consumption to a new shared reference (kept under the 32KB #2551 budget; no inline subagent dispatch token). - references/loop-hook-dispatch.md: new canonical, point-agnostic contract for consuming `loop render-hooks --raw` activeHooks (contribution/step/ gate) — single source for hook consumption across host loops. - gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host file) — one source of truth for the host-loop file + wired-point set. - phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES and shared scanWiredPoints (was a hand-maintained duplicate omitting discuss-phase.md + a duplicated regex). - gen-capability-registry.cjs: add validateHooksWired() gen-time guard that rejects a capability hook declared at a valid-but-unwired loop point, with a clear remediation message — failure now surfaces at gen --check/--write time instead of deep in the full conformance suite. - tests (capability-registry.test.cjs): regression + anti-pattern parity guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES; POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into the discuss-class gap again. - docs/INVENTORY*: register the new reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1196): backfill changeset PR number (#1199) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5fa4dcd78c |
fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at
|
||
|
|
aec3374bc2 | feat(#1138): make runtime descriptors authoritative (#1157) | ||
|
|
4ab5c7b3f2 |
feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities * chore(#1135): add phase 6 planning capabilities changeset * fix(#1135): satisfy lint for agent hook rendering |
||
|
|
fd01e7a12e |
feat(#1132): complete contribution hook prerequisite
Closes #1132 |
||
|
|
eb051ea696 |
feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123 Closes #1124 Refs #857 |